Senior Software Engineer - SRE, Retail and Pharmacy
Quick Summary
burn rate alerting, composite health signals,
catalog third-party API failure modes, shared infrastructure failure paths,
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.
Our Site Reliability Engineering team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure — spanning hybrid cloud and on-premises environments at massive fleet scale.
Our engineering philosophy is grounded in five pillars: Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX).
Our operating principle is the reliability covenant: our success is not measured by incident response volume — it is measured by the reliability capability we transfer to the engineering teams we serve. Your success in this role is measured by what the engineering teams in your domain can do independently after working with you, not by how indispensable you become to them. An SSE who has enabled a development team to detect, respond to, and learn from production failures without SRE involvement has delivered the highest-value outcome this role can produce.
We track operational toil as an engineering metric. Engineers at this level are expected to identify recurring manual work, eliminate it through automation, document the reduction, and treat toil accumulation as a reliability risk — not as a sign of operational expertise.
About the Role
~1 min readAs a Senior Software Engineer — SRE, you independently own the reliability posture of an assigned engineering domain. You are not waiting for direction — you are setting it for your domain. You design the alerting strategy, own the SLO health, lead incident command for production issues, facilitate postmortems, tune anomaly detection models, and partner directly with engineering domain owners to shift reliability left into design.
You are a technical mentor to SE-level engineers and an escalation resource during active incidents. You have the technical depth to diagnose complex distributed system failures, the data instincts to distinguish genuine anomalies from noise in ML-generated signals, and the organizational skills to drive reliability practice adoption in teams that did not necessarily ask for SRE involvement.
This role exists inside an active SRE transformation. The domain-based SRE ownership model you will operate within is in its early stages. Some of the toolchains you will work with are being built in parallel with the operational work. Engineering domain owners are simultaneously learning what SRE can offer them.
Responsibilities
~1 min readOwn SLI/SLO health for your assigned domain end-to-end; monitor error budget burn rates and drive proactive burn-down actions before incidents reach end users or pharmacy patients
Design multi-signal alerting strategies that go well beyond threshold alerts: burn rate alerting, composite health signals, and anomaly-based detection using time-series models — and validate their output against production ground truth
Lead Production Readiness Reviews (PRR) for services in your assigned domain; own the readiness gate sign-off and be accountable for what makes it into production on your watch
Design and execute fault injection experiments at service level using tools such as LitmusChaos, Chaos Toolkit, or Gremlin; validate blast radius assumptions before rollouts reach production
Partner with engineering domain owners on reliability requirements during architecture design and sprint planning — reliability is designed in, not bolted on
Own dependency risk mapping for your domain: catalog third-party API failure modes, shared infrastructure failure paths, and chain-wide blast radius scenarios
Apply a fleet operations mindset to change management: for any configuration or deployment change touching the edge fleet, assess deployment blast radius by node cohort, review the rollback procedure, and contribute to the go/no-go decision on high-risk changes
Understand progressive rollout strategies — canary cohorts, staged fleet expansion, blast radius budgets — and apply them in high-risk deployment reviews
Serve as the primary on-call Technical Incident Commander (IC) for domain incidents; drive structured bridge calls from detection to resolution using established incident command frameworks
Author and maintain P0/P1 runbooks with validated, step-by-step remediation procedures; own a quarterly review and dry-run testing cadence to ensure runbooks are accurate when they are actually needed
Lead post-incident reviews using structured root cause analysis: causal chain documentation, origin layer classification, and contributing factor identification — not just a timeline of what happened
Serve as the real-time escalation point and technical decision support for SE engineers during active incidents
Pursue Technical Incident Commander (TIC) certification; qualify as a cross-domain IC candidate available to lead major incidents beyond your assigned domain
Learning Loops at this level is systems engineering for organizational memory — not postmortem administration.
Facilitate domain-level postmortems with rigor and structure: timestamped timelines, contributing factor taxonomy (origin layer + failure pattern classification), systemic findings, and action items with owners, due dates, and measurable definitions of done
Requirements
~1 min read-
5+ years of experience in SRE, DevOps, platform engineering, or related production-systems roles
3+ years operating cloud-native distributed systems at production scale with active on-call responsibility
Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents — structured bridge leadership, not just participant involvement
Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.
Experience owning Production Readiness Reviews or service launch gates
Hands-on chaos or fault injection experience using LitmusChaos, Chaos Toolkit, or Gremlin
TIC (Technical Incident Commander) certification or equivalent structured incident command training
Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) — design or implementation experience
Experience with streaming data platforms: Apache Kafka, Redpanda, Apache Flink, or ksqlDB
Familiarity with analytical databases for observability workloads: ClickHouse, Apache Druid, or TimescaleDB
Experience with service mesh and traffic management: Istio, Envoy, Linkerd
Infrastructure-as-code proficiency at production scale: Terraform, Pulumi, or Ansible
This is a core function of the SSE role — not a soft skill addendum.
You own the SLO health and alerting strategy for your assigned domain with full independence — including the ML-based anomaly detection configuration — and the false positive rate is measurably lower than when you joined
You have led at least five P1/P2 incident bridges as Incident Commander, with structured postmortems published, action items closed, and at least two incident classes eliminated through root cause remediation rather than symptom suppression
You have facilitated your first domain-level PRR and signed off on a service launch
Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience
The typical pay range for this role is:
$92,700.00 - $203,940.00This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above.
Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.
What We Offer
~1 min readWe take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.
Additional details about available benefits are provided during the application process and on Benefits Moments.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.
Location & Eligibility
Listing Details
- First seen
- September 27, 2026
- Last seen
- September 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 56%
- Scored at
- September 27, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.