2h ago
New

Principal Site Reliability Engineer

lead
EngineeringDevops Engineer
2 views0 saves0 applied

Quick Summary

Overview

Define and evolve the organization's Site Reliability Engineering strategy, vision, and technical roadmap in partnership with engineering leadership and business stakeholders.

Technical Tools
EngineeringDevops Engineer
Define and evolve the organization's Site Reliability Engineering strategy, vision, and technical roadmap in partnership with engineering leadership and business stakeholders. Guide the lifecycle of services from conception to end-of-life across the organization, including establishing service design review frameworks, capacity planning methodologies, and production readiness standards that scale across teams. Establish and drive adoption of organization-wide standards and best practices related to system architecture, service delivery, reliability, automation, and operational excellence. Influence technical decisions across multiple teams and business units. Build and operate shared platforms, tooling, and frameworks that enable service and engineering teams across the organization to achieve high availability and improve incident detection and response. Drive system performance, availability, and efficiency improvements across the organization through architectural guidance, automation strategy, process refinement, and deep analysis of operational patterns in incident reviews. Partner closely with engineering and product leadership across the organization to establish reliability standards, shape technical direction, and deliver reliable services at scale. Champion a culture of operational excellence by treating operational challenges such as software engineering problems and mentoring teams on reducing toil at scale. Lead, mentor, and develop Site Reliability Engineering talent across multiple teams, establishing best practices and growing SRE capability within the organization. Partner with executive stakeholders, product leadership, and business teams to align reliability investments with business priorities and drive technical strategy that supports business outcomes. 12+ years of hands-on experience in software engineering, systems engineering, and/or cloud-based environments. 10+ years of experience working with public cloud platforms (e.g., GCP, AWS, or Azure), including designing and operating large-scale systems. 5+ years of experience designing, operating, and maintaining applications and/or systems infrastructure in large-scale, customer-facing production environments. Demonstrated experience architecting and influencing large-scale distributed systems and infrastructure solutions. Proven track record of leading cross-functional technical initiatives and mentoring engineering teams. Demonstrated understanding of observability best practices, including metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing. Experience coding in one or more higher-level programming languages (e.g., Python, Java, C# or C++). Strong working knowledge of Linux systems, including troubleshooting, performance analysis, and scripting in production environments. Experience with GitHub Actions and modern CI/CD practices. Exceptional communication and collaboration skills, with demonstrated ability to influence across teams, lead technical discussions with executives, and mentor engineers. Experience defining and communicating technical vision and strategy across large organizations. Deep expertise in distributed system design and architecture Hands-on experience with cloud-native applications and containerization technologies (e.g., Kubernetes, containers). Experience designing and managing infrastructure-as-code and configuration management strategies across organizations (e.g., Terraform, Ansible). Experience operating and optimizing production workloads at scale, including cost optimization and performance tuning. Solid grounding in at least three of the following areas: Computer Science fundamentals, Cloud Architecture, Security or Network Design Experience designing observability strategies and building metrics pipeline, operational dashboards and alerts using observability tools such as Splunk or Grafana. Experience partnering with business and product teams on technology decisions and translating technical recommendations into business value. Experience driving operational change and building consensus around new technical practices and standards

Location & Eligibility

Where is the job
—
Location terms not specified

Listing Details

Posted
October 9, 2026
First seen
October 9, 2026
Last seen
October 9, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
56%
Scored at
October 9, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Principal Site Reliability Engineer