Especialista de SRE
Quick Summary
Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments. Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
As a Site Reliability Engineering specialist, you will play a key role in ensuring the reliability and resilience of critical products in a large-scale technology environment.
You will serve as a technical reference for SRE practices, partnering closely with development, product, and operations teams.
The role focuses on high availability, performance, observability, automation, and secure infrastructure practices.
You will help define and monitor SLIs, SLOs, and SLAs aligned with business objectives.
Your expertise will contribute to incident prevention, capacity planning, architectural resilience, and efficient recovery from failures.
You will work with modern cloud, Kubernetes, infrastructure-as-code, CI/CD, and observability technologies.
This is an opportunity to solve complex reliability challenges while strengthening the resilience of mission-critical systems.
- Act as the technical SRE reference for Identity & Fraud products, supporting development and operations teams.
- Define, implement, and monitor SLIs, SLOs, and SLAs aligned with business goals and service expectations.
- Lead incident analysis and implement preventive and corrective actions to reduce recurring issues.
- Automate provisioning, deployment, scaling, and failure-recovery processes to improve operational efficiency and reliability.
- Design and maintain observability solutions covering logs, metrics, traces, and alerts.
- Support capacity and performance engineering to ensure systems can handle demand predictably.
- Contribute to architectural improvements focused on resilience, scalability, and security.
- Promote infrastructure-as-code, CI/CD, version control, and safe change-management practices.
- Troubleshoot and mitigate issues in real time within critical production environments.
Requirements
~1 min read- Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments.
- Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
- Advanced knowledge of automation and infrastructure as code, including Terraform and Ansible.
- Experience with monitoring and observability, particularly Datadog, along with familiarity with Prometheus, ELK, and Grafana.
- Hands-on experience with CI/CD pipelines, version control, and reliable deployment practices.
- Strong ability to analyze performance, troubleshoot complex issues, and optimize distributed systems.
- Knowledge of relational and non-relational databases.
- Ability to collaborate effectively with development, product, and operations teams.
- Strong communication, systems thinking, analytical skills, and a problem-solving mindset.
- Experience with resilience engineering in identity and fraud systems is desirable.
- Cloud certifications in AWS, OCI, Azure, or GCP are a plus.
- Experience with chaos engineering and resilience testing is desirable.
- Knowledge of application and infrastructure security is an advantage.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- September 23, 2026
- First seen
- September 27, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 46%
- Scored at
- September 29, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.