Director, Site Reliability Engineering
Quick Summary
10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.
As Director, Site Reliability Engineering, you will lead the teams and technical strategy responsible for keeping critical infrastructure reliable, scalable, and secure for millions of users. You will work across software, systems, automation, cloud infrastructure, and operational processes to solve complex reliability challenges at scale. The role combines strategic leadership with hands-on technical depth, including troubleshooting production systems and partnering closely with software engineering teams. You will help shape the future architecture and deployment practices of a large-scale, privacy-focused technology environment. You will also drive improvements in automation, observability, incident response, and engineering efficiency. This is a remote-first leadership opportunity with significant ownership, autonomy, and impact.
-
Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems.
-
Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices.
-
Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem.
-
Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation.
-
Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks.
-
Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions.
-
Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.
-
Leverage cloud-native architectures and services to strengthen system resilience and support continued growth.
-
Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents.
-
Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture.
-
Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning.
Requirements
~2 min read-
10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.
-
4+ years of experience leading SRE or comparable engineering teams.
-
Experience participating in or managing 24/7 on-call operations for large-scale production environments.
-
Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems.
-
Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments.
-
Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem.
-
Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities.
-
Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems.
-
Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions.
-
Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.
-
Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies.
-
Experience with AI-driven software development, including the design and implementation of agentic workflows.
-
Strong ability to turn ambiguous or complex problems into practical, innovative solutions with measurable outcomes.
-
Strategic thinking and technical foresight, with the ability to anticipate future infrastructure and reliability requirements.
-
Excellent communication and collaboration skills, with the ability to work effectively across engineering teams and technical disciplines.
-
Strong sense of ownership, autonomy, and accountability in a remote-first working environment.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 23, 2026
- First seen
- September 27, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 28, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.