GoDaddy
GoDaddy7h ago
New

Site Reliability Engineer

EngineeringDevops Engineer
1 views0 saves0 applied

Quick Summary

Key Responsibilities

on-call experience, incident response, and a track record of reducing toil through automation. Clear written and verbal communication — able to document systems, write post-incident reviews,

Technical Tools
EngineeringDevops Engineer

At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​

Responsibilities

~2 min read
  • Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments.
  • Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely — designing and executing migration tooling, validation, and rollback strategies that protect customer experience.
  • Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions.
  • Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed.
  • Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data.
  • Participate in on-call and incident response. Take a fair share of the on-call rotation, lead incident response when you're the responder, and drive the blameless post-incident process that turns failures into permanent fixes.
  • Raise the engineering bar. Review peers' code and designs, document systems and runbooks clearly, and mentor SRE I/II engineers.
  • Contribute to the AI/MCP initiative. Help bring AI-assisted workflows and internal MCP tooling into our operations so internal customers can resolve problems and incidents faster.
  • Participate in a shared on-call rotation (after onboarding) and help lead incident response activities when needed.
  • 5+ years in SRE, infrastructure, platform, or systems engineering roles operating production systems at scale.
  • Strong Linux systems fundamentals — networking, storage, processes, performance troubleshooting.
  • Proficiency in Python for automation and tooling (writing maintainable, tested code — not just scripts).
  • Hands-on experience operating distributed systems and diagnosing issues across service boundaries.
  • Experience with infrastructure-as-code and configuration management (Puppet, Ansible, or equivalent).
  • Comfort owning production reliability: on-call experience, incident response, and a track record of reducing toil through automation.
  • Clear written and verbal communication — able to document systems, write post-incident reviews, and collaborate across a distributed, remote team.
  • Experience operating OpenStack (Nova, Neutron, Ceph) or comparable cloud infrastructure platforms
  • Experience with Docker, Kolla, or containerised infrastructure
  • Experience supporting large distributed systems at scale
  • Experience operating Ceph or other software-defined storage at scale.
  • Exposure to OpenStack migration or cloud-migration programs.
  • Interest in applying AI/LLM tooling to operational workflows.

 

We encourage you to apply even if your experience or skillset doesn’t align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth.

Requirements

~1 min read

Location & Eligibility

Where is the job
Romania
On-site within the country
Who can apply
RO

Listing Details

Posted
August 25, 2026
First seen
August 25, 2026
Last seen
August 25, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
67%
Scored at
August 25, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
GoDaddy
GoDaddy
greenhouse

GoDaddy helps the world easily start, confidently grow, and successfully run an online presence.

Employees
5k+
Founded
1997
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

GoDaddySite Reliability Engineer