2mo ago

Site Reliability Engineering (SRE) Lead

IndiaIndia·Hyderabadlead
OtherSite Reliability Engineering
1 views0 saves0 applied

Quick Summary

Key Responsibilities

Datadog Dynatrace Prometheus Grafana Splunk ELK Stack AppDynamics Required

Technical Tools
OtherSite Reliability Engineering
Job Title: Site Reliability Engineering (SRE) Lead Location: Hyderabad / Mumbai Employment Type: Full-Time About SID Global Solutions SID Global Solutions (SIDGS) is a leading Digital Engineering and Technology Services company specializing in Cloud, API Management, DevOps, Platform Engineering, Kubernetes, Microservices, and Digital Transformation. We are looking for an experienced Site Reliability Engineering (SRE) Lead to drive platform reliability, operational excellence, and production stability for mission-critical enterprise applications. Job Summary We are seeking a highly skilled SRE Lead with strong expertise in cloud infrastructure, Kubernetes, API Management, and production operations. The ideal candidate will be responsible for ensuring the availability, scalability, performance, and reliability of enterprise applications while leading incident management, observability, and automation initiatives. In this role, you will work closely with Development, DevOps, Infrastructure, Platform Engineering, and Application Support teams to maintain highly available production environments, reduce operational risks, and improve service reliability through automation and proactive monitoring. Key Responsibilities Reliability Engineering Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to improve platform reliability. Review application architecture and infrastructure designs to ensure scalability, resilience, and high availability. Drive reliability improvements across cloud-native applications and distributed systems. Identify opportunities to eliminate operational bottlenecks through automation and process improvements. Incident & Production Management Lead Major Incident Management (P1/P2) activities and act as the primary technical escalation point for critical production issues. Coordinate with cross-functional teams to restore services within defined SLAs. Conduct Root Cause Analysis (RCA) and drive preventive and corrective actions to reduce recurring incidents. Participate in Change Management and Release activities to ensure production stability. Prepare incident reports and communicate status updates to business and technology stakeholders. Automation & Platform Engineering Develop automation scripts and self-healing solutions to improve operational efficiency. Build reusable operational runbooks and standard operating procedures. Automate routine operational tasks using Python, Bash, or similar scripting languages. Support Infrastructure as Code (IaC) initiatives and CI/CD pipeline improvements. Monitoring & Observability Design and maintain enterprise monitoring and alerting solutions. Create dashboards and alerts to proactively monitor application health, infrastructure, APIs, and Kubernetes environments. Analyze performance trends and recommend improvements for system stability and capacity planning. Ensure effective monitoring coverage across production environments. API & Kubernetes Administration Manage and troubleshoot Google Apigee API Gateway configurations, policies, and traffic routing. Monitor Kubernetes clusters, workloads, namespaces, ingress controllers, and container health. Optimize application performance and resource utilization within Kubernetes environments. Support production deployments and post-release validation activities. Leadership & Collaboration Mentor SRE, DevOps, and Production Support engineers. Establish operational best practices, troubleshooting guidelines, and technical documentation. Collaborate with Development, QA, Infrastructure, Security, and Business teams to improve platform reliability. Drive a culture of continuous improvement, automation, and operational excellence. Required Technical Skills Strong experience with Google Cloud Platform (GCP). Hands-on expertise in Google Apigee API Management. Experience managing production Kubernetes (K8s) environments. Good understanding of NGINX, reverse proxy, and load balancing concepts. Strong knowledge of Linux/Unix Administration. Experience with Python, Bash, or Go scripting. Familiarity with CI/CD pipelines, Git, and Jenkins. Understanding of Infrastructure as Code (Terraform or equivalent is preferred). Monitoring & Observability Tools Experience with one or more of the following: Datadog Dynatrace Prometheus Grafana Splunk ELK Stack AppDynamics Required Qualifications 7–10 years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, or Production Support. Strong experience supporting enterprise production environments. Hands-on experience with Kubernetes, GCP, and Google Apigee. Good understanding of distributed systems, microservices architecture, and cloud-native applications. Experience in Incident, Problem, Change, and Release Management. Ability to troubleshoot complex production issues and coordinate cross-functional teams during critical incidents. Preferred Qualifications Experience in Banking, Financial Services, or other enterprise environments. ITIL Foundation certification. Google Cloud Professional Certification. Kubernetes Certification (CKA/CKAD) is an added advantage. Experience with Service Mesh technologies such as Istio is preferred.

Location & Eligibility

Where is the job
Hyderabad, India
On-site at the office

Listing Details

Posted
July 10, 2026
First seen
September 25, 2026
Last seen
October 6, 2026

Posting Health

Days active
11
Repost count
0
Trust Level
20%
Scored at
October 7, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Site Reliability Engineering (SRE) Lead