Rapidai
Rapidai8h ago
New

Senior Site Reliability Engineer

IndiaIndia·BangaloreFull Timesenior
EngineeringDevops Engineer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Own the availability, performance, and incident response for Rapid's production EKS clusters Design and operate the full observability stack — metrics, logs,

Requirements Summary

10+ years in SRE, DevOps, or infrastructure engineering roles Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch,

Technical Tools
EngineeringDevops Engineer
RapidAI is the trusted leader in deep clinical AI, helping hospitals deliver faster, more informed care through intelligent imaging and integrated workflows. The Rapid Enterprise™ Platform supports disease states across the care spectrum, but it’s our clinical depth that drives the most meaningful impact — improving decision-making, patient outcomes, and health-system performance. Used by more than 2,500 hospitals in over 100 countries and backed by 700+ clinical studies, including research that helped expand national stroke-treatment guidelines, RapidAI is the most clinically validated AI platform in healthcare.

Responsibilities

~2 min read
  • Own the availability, performance, and incident response for Rapid's production EKS clusters
  • Design and operate the full observability stack — metrics, logs, traces — with
    Open Telemetry as the foundation
  • Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
  • Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
  • Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
  • Tune autoscaling, networking, and cost efficiency across AWS workloads
  • On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
  •  
    What We Looking For:

  • 10+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
    surrounding ecosystem
  • Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
  • Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
  • Strong foundation in Linux, networking, and distributed systems fundamentals
  • Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
    Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out
  • Startup mindset: you make decisions with incomplete information and iterate quickly
  • RapidAI is committed to creating an inclusive and diverse workplace. We provide equal employment opportunities to all employees and applicants and prohibit discrimination and harassment of any type in regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

    Location & Eligibility

    Where is the job
    Bangalore, India
    Hybrid — some on-site time required
    Who can apply
    IN

    Listing Details

    Posted
    August 25, 2026
    First seen
    August 25, 2026
    Last seen
    August 25, 2026

    Posting Health

    Days active
    0
    Repost count
    0
    Trust Level
    70%
    Scored at
    August 25, 2026

    Signal breakdown

    freshnesssource trustcontent trustemployer trust
    Rapidai
    Rapidai
    lever
    Employees
    125
    Founded
    2012
    View company profile
    Newsletter

    Stay ahead of the market

    Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

    A
    B
    C
    D
    Join 12,000+ marketers

    No spam. Unsubscribe at any time.

    RapidaiSenior Site Reliability Engineer