Senior Site Reliability Engineer
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in India.
This is a senior engineering role focused on building and operating highly available, resilient distributed systems at global scale.
You will architect reliability solutions across multi-region microservices, APIs, and authentication infrastructure running on AWS and GCP.
The role combines hands-on engineering with technical leadership across observability, disaster recovery, Kubernetes, and infrastructure automation.
You will help define reliability standards, including SLIs, SLOs, error budgets, and high-availability objectives.
You will also lead major incident response and turn production learnings into lasting systemic improvements.
Beyond reliability, you will drive cloud cost optimization, automation, and AI-assisted engineering practices.
This is an opportunity to mentor engineers, shape platform architecture, and raise the technical bar in a fast-moving, remote-first environment.
-
Architect, scale, and continuously improve the reliability, availability, and performance of multi-region microservices, APIs, and authentication infrastructure across AWS and GCP.
-
Design and maintain disaster recovery and business continuity strategies, including multi-region failover automation, recovery dashboards, and validation processes aligned with defined RTO and RPO targets.
-
Establish and enforce SLI, SLO, and error-budget frameworks across engineering teams to improve reliability and accountability.
-
Lead the observability strategy using platforms such as Datadog, with actionable Golden Signals monitoring designed to reduce MTTD, MTTR, and alert fatigue.
-
Lead on-call escalation and major incident management, ensuring rapid response to production issues and strong adherence to availability objectives.
-
Facilitate blameless post-incident reviews and drive root-cause remediation to prevent recurring failures.
-
Architect and operate production Kubernetes environments, including EKS/GKE, networking, RBAC, ingress/egress, and GitOps workflows using tools such as Argo CD and Kargo.
-
Design and maintain modular, enterprise-grade Terraform infrastructure across complex multi-account and multi-region cloud environments.
-
Build FinOps dashboards and cost-optimization initiatives covering resource utilization, right-sizing, cost allocation, and multi-cloud spend visibility.
-
Develop production-grade Python or Go tooling, automation, integrations, and platform capabilities to eliminate operational toil.
-
Champion AI-assisted engineering tools to accelerate automation, runbook creation, incident triage, and engineering productivity.
-
Create operational runbooks and architecture documentation while mentoring junior and mid-level engineers and contributing to technical standards.
Requirements
~2 min read-
8+ years of professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or software engineering supporting 24/7 mission-critical distributed systems.
-
Bachelor’s degree in Computer Science, Software Engineering, or a comparable technical discipline, or equivalent professional experience.
-
Advanced Python or Go development skills, with experience building internal SRE platforms, automation tools, and API integrations.
-
Deep hands-on Kubernetes expertise, including production EKS/GKE environments, cluster lifecycle management, networking, RBAC, ingress/egress, and GitOps tooling such as Argo CD.
-
Strong Terraform expertise, including module architecture, state management, refactoring, and infrastructure deployment across multi-account AWS environments.
-
Strong AWS and/or GCP expertise, including services and concepts such as IAM, VPC, Transit Gateway, ALB/NLB, Route53, and cloud networking.
-
Proven experience designing and testing multi-region disaster recovery architectures, automating failover, and monitoring recovery health.
-
Strong knowledge of SLI/SLO frameworks, production observability, PagerDuty, monitoring strategy, and reliability engineering practices.
-
Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy technologies such as HAProxy or NGINX.
-
Demonstrated FinOps and cloud cost-optimization experience, including right-sizing, cost allocation, workload optimization, and financial visibility.
-
Experience with technical leadership, architectural discussions, RFCs/design documentation, and mentoring engineering peers.
-
Strong troubleshooting, communication, collaboration, and problem-solving skills, with the ability to work effectively on complex distributed systems.
-
Fluency in written and spoken English and the ability to collaborate with global teams.
-
Willingness and ability to participate in an on-call rotation and respond during assigned shifts.
-
Preferred experience includes chaos engineering, secrets management tools such as Vault or AWS Secrets Manager, DevSecOps practices, automated infrastructure vulnerability remediation, and identity or IAM-focused platforms.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 5, 2026
- First seen
- October 5, 2026
- Last seen
- October 5, 2026
Posting Health
- Days active
- 0
- Repost count
- 1
- Trust Level
- 62%
- Scored at
- October 5, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.