USD 159800-235000/yr

Software Engineer, Reliability Platforms

United StatesUnited States·San Francisco,Sunnyvale,New Yorkmid
Software EngineerSoftware Engineering
2 views0 saves0 applied

Quick Summary

Overview

About the Team The Reliability Platform role is a key pillar of DoorDash’s Production Lifecycle team, alongside Observability and Deploy Platform.

Technical Tools
Software EngineerSoftware Engineering

The Reliability Platform role is a key pillar of DoorDash’s Production Lifecycle team, alongside Observability and Deploy Platform. This group’s mandate is to enable users and agents to reason about the health of our services, facilitate change control safety, and provide the means to rapidly address any unexpected state. 

Ownership is fundamental in DoorDash culture, and all teams own what they build. We are not here to operate services on others’ behalf, but to provide tools that enable their success and ensure a consistently high level of quality for everything we do. We approach challenges with the pragmatic perspective of an SRE, and deliver solutions with the mindset of a SWE who detests toil and repetitive tasks. We use software and agents to “keep the lights on” and focus our energy on innovation that will level up the entire organization. 

This mission falls into three main categories.

  • Service Health – Providing SLO frameworks, analytics tools, and AI Agent enablement to extract high quality insights from our telemetry to pinpoint faults, or highlight deficiencies
  • Change Orchestration – Provide self-service provisioning orchestration, evolving from UI to Agent-driven to allow our developers to safely affect production from their IDE
  • Incident Management – Define and deliver tools/processes/policies leveraged by our peers to quickly understand and recover from any unexpected issues in the environment

This mandate implies a broad contribution across many aspects of the infrastructure, and demands equal parts software development and systems integration. Our priorities are always informed by an obsession to level up over 4,000 internal customers/peers, and obfuscate infrastructure complexity so they can focus on making the DoorDash product itself amazing!

About the Role

~3 min read

As a Software Engineer on the Reliability Platform team, you’ll help design, build, and operate services and infrastructure that deliver on the team’s broad mandate described above. This team has a unique opportunity for breadth, often in collaboration with expert peers across the Infrastructure and Product teams. Depending on need and interest, you may be working on mission-critical back-end services or pipelines, complex orchestration workflows, self-service UI, or AI Agent continuous improvements.

We have fully embraced the use of AI tools in everything we do, and believe in the incredible potential this provides while remaining pragmatic enough to ensure the critical infrastructure we maintain cannot be compromised. Our goal is to deliver innovative next generation capabilities, as well as make data in our custody available to others pursuing the same. 

A few examples of efforts the team has owned in recent years:

  • Delivering framework to capture/alert/report on SLO quality across tens of thousands of endpoints ensuring all teams are accountable for the quality of their delivered services
  • Replacement of our escalation management tools including alignment with our internal Asset/Team Catalog to allow automated alert routing and cross-brand alignment
  • Delivery of MCP back-end for Reliability Platform data/tools, as well as enabling the same for peer teams across the Core Infrastructure organization
  • Design and delivered orchestration tools to enable self-service provisioning of critical infrastructure (Kafka topics, Databases, CPU/GPU Pools, Service Scaffolding, etc)
  • PoC for internal SRE AI Agentic tooling leveraging internal MCPs and domain specific profiles to facilitate troubleshooting and Q&A capabilities replacing FAQs/Runbooks
  • Delivered per-pod realtime configuration key-value tooling enabling runtime feature flag management from a central source of truth across the fleet (100K+ pods) 

We are proud of our engineering culture, and many of our greatest successes are born from an individual with an idea spending some time hacking out a rudimentary demonstrable prototype. The mandate of this team is ripe for individuals with this creative pioneering mindset, and the ability to execute.

  • Delivery Innovative Capabilities: You don’t want to ‘turn the crank’ somewhere, but you want to contribute to some frontier thinking and help us push the industry forward
  • Build Great Infrastructure: You know great infrastructure often goes unnoticed by design. You are content knowing your efforts allow you to claim a portion of everyone’s success.
  • Balance Practical and Possible: Sometimes our pragmatic perspective is needed to maintain a high quality service; your experience will support finding the right risk balance
  • Be Custom Obsessed: We want to learn from our customers to ensure we are solving the right challenges, and also share our perspective to influence in areas of expertise
  • Automate Everything: Well… not everything… but if your first instinct is to ask how this toil could be automated or better yet avoided then you’re on the right team
  • Shape the Future of Operations: Experiment with agentic, AI-assisted workflows that can propose, validate, and safely execute production changes — moving DoorDash toward proactive, self-healing systems in step with industry first movers.
  • Platform Engineering Mindset: You think in terms of APIs, abstractions, and workflows — you enjoy building systems that other engineers depend on every day.
  • Proven Experience: You have 5+ years of experience in an infrastructure, platform, or backend engineering role, showing you can deliver and maintain complex systems.
  • Backend Development Skills: You’re fluent in Go (or a similar language) and can design and deliver mission-critical services that are scalable, resilient, performant, and efficient.
  • Cloud/Infra Expertise: You’re comfortable with AWS primitives, security best practices, containerization, and Infrastructure as Code tools like Terraform or Pulumi. 
  • SRE Experience: You understand concepts like SLOs, error budgets, and incident response though this is a platform development team, not an SRE/oncall team.
  • Flexibility: You will work on cool/fun stuff the majority of your time, but also accept that some tasks just need to get done and reflect an opportunity for automation/improvement
  • AI Alignement: You embrace the use of AI tools to be a more productive and capable engineer. This applies to coding, planning, supporting peers, and everything you do. 
  • Curiosity About the Future: You’re excited about automation and agentic, AI-assisted operations and want to help shape how engineers interact with production systems.

Requirements

~1 min read

Notice to Applicants for Jobs Located in NYC or Remote Jobs Associated With Office in NYC Only

We used Covey as part of our hiring and/or promotional process for jobs in NYC and certain features may qualify it as an AEDT in NYC. As part of the hiring and/or promotion process, we provided Covey with job requirements and candidate submitted applications. We began using Covey Scout for Inbound from August 21, 2023, through December 21, 2023.  We resumed using Covey Scout for Inbound again on June 29, 2024, and ceased using Covey Scout for Inbound on April 30, 2026.

The Covey tool has been reviewed by an independent auditor. Results of the audit may be viewed here: https://getcovey.com/nyc-local-law-144.

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
June 11, 2026
First seen
June 12, 2026
Last seen
July 28, 2026

Posting Health

Days active
20
Repost count
0
Trust Level
47%
Scored at
July 2, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Doordashusa
Doordashusa
greenhouse

Leading US food and goods on-demand delivery platform with 60%+ market share

Employees
10,000+
Founded
2013
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

DoordashusaSoftware Engineer, Reliability PlatformsUSD 159800-235000