Site Reliability Engineer | Platform
Quick Summary
In this role, you will strengthen the SRE Platform team’s mission by advancing the foundational platforms that automate manual workflows and elevate system reliability.
In this role, you will strengthen the SRE Platform team’s mission by advancing the foundational platforms that automate manual workflows and elevate system reliability. Your work will ensure our staging environments remain stable and production-like, empowering QA and development teams to test, validate, and deploy their applications with confidence. You will also contribute to operational excellence through active participation in the weekly on-call rotation, supporting consistent and dependable infrastructure performance.
Automate and optimize operational processes
Enhance and maintain the observability stack
Oversee test/staging environments management
Develop and support critical production components
Handle and resolve production incidents
Participate in the on-call rotation
Strong teamwork and collaboration skills
Solid understanding of SRE concepts, including SLIs, SLOs, SLAs, and Error Budgets
Proficiency in Python or another scripting language
Strong grasp of software engineering principles
Hands-on experience with observability and monitoring tools such as Prometheus and Grafana
Familiarity with logging stacks (e.g., ELK, Loki) and tracing systems (e.g., Jaeger, Tempo)
Understanding of RDBMS and Redis
Experience working with Kubernetes and related tooling (e.g., Helm)
Location & Eligibility
Listing Details
- First seen
- May 6, 2026
- Last seen
- September 25, 2026
Posting Health
- Days active
- 103
- Repost count
- 0
- Trust Level
- 12%
- Scored at
- August 18, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.