Site Reliability Engineer

EngineeringDevops Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Observability Environment Management: Design, build, and maintain our observability infrastructure, including monitoring tools, logging platforms, and distributed tracing systems (e.g., Prometheus,

Requirements Summary

Bachelors degree in computer science or a related field, or equivalent experience. 5+ years of experience as an SRE or in a similar role with a focus on observability.

Technical Tools
EngineeringDevops Engineer

We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our growing Observability team. The ideal candidate will have a strong background in building and maintaining robust observability environments, including monitoring, logging, and tracing systems. This role will focus on the design, implementation, and support of our observability infrastructure, ensuring the seamless onboarding of applications and providing critical support during incidents.

Responsibilities

~1 min read
  • →Observability Environment Management: Design, build, and maintain our observability infrastructure, including monitoring tools, logging platforms, and distributed tracing systems (e.g., Prometheus, Grafana, Elasticsearch, etc.). This includes capacity planning, performance tuning, and ensuring high availability.
  • →Application Onboarding: Work with development teams to onboard applications to our observability platform, providing guidance on instrumentation best practices and ensuring data quality. This includes creating and maintaining documentation and training materials.
  • →Incident Support: Provide timely and effective support during incidents, leveraging observability data to diagnose and resolve issues quickly. This includes contributing to post-incident reviews and implementing preventative measures.
  • →Automation: Automate repetitive tasks and processes related to observability, improving efficiency and reducing manual effort. This may involve scripting, developing tools, or integrating with CI/CD pipelines.
  • →Alerting and Monitoring: Develop and maintain effective alerting strategies, ensuring appropriate escalation procedures and minimizing noise. This includes creating dashboards and reports to visualize system health and performance.

Requirements

~1 min read
  • Bachelors degree in computer science or a related field, or equivalent experience.
  • 5+ years of experience as an SRE or in a similar role with a focus on observability.
  • Strong understanding of distributed systems and microservices architectures.
  • Experience with any monitoring, logging, and tracing tools (e.g., Prometheus, Grafana, Jaeger, Elasticsearch, Fluentd, Datadog, Dynatrace, etc.).
  • Proficiency in scripting languages such as Python, Go, or Bash.
  • Strong problem-solving and analytical skills.
  • Excellent communication and collaboration skills.

Nice to Have

~1 min read
  • Experience with cloud platforms.
  • Experience with infrastructure-as-code tools (e.g., Terraform, Ansible)

Location & Eligibility

Where is the job
Singapore
On-site within the country
Who can apply
SG

Listing Details

First seen
September 25, 2026
Last seen
September 27, 2026

Posting Health

Days active
1
Repost count
0
Trust Level
56%
Scored at
September 27, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

ntt-data-singaporeSite Reliability Engineer