Senior Site Reliability Engineer (AWS)
Quick Summary
EKS, EC2, RDS/Aurora, Lambda, API Gateway CloudFront, WAF, ALB/NLB CloudWatch, X-Ray, IAM, Secrets Manager Lead cloud migrations and modernization initiatives, including legacy system refactoring.
At Broadridge, we've built a culture where the highest goal is to empower others to accomplish more. If you’re passionate about developing your career, while helping others along the way, come join the Broadridge team.
Responsibilities
~1 min readDesign and implement high-availability, fault-tolerant architectures across on-prem and cloud platforms (AWS).
Lead multi-region DR planning, implementation, and testing, including RTO/RPO definition and validation.
Define and enforce SLOs, SLIs, and error budgets to balance reliability with delivery velocity.
Drive self-healing automation and proactive remediation strategies.
Build and maintain infrastructure using Terraform and configuration management tools (e.g., Chef).
Develop automation to eliminate manual operational tasks (TOIL reduction).
Create reusable modules, pipelines, and guardrails for standardized deployments.
Automate certificate lifecycle management, key rotation, and security updates.
Design and implement end-to-end observability (metrics, logs, traces, synthetic monitoring).
Build dashboards, alerts, and runbooks to enable fast detection and resolution of incidents.
Improve signal-to-noise ratio in alerting to reduce operational fatigue.
Perform root cause analysis (RCA) and lead post-incident reviews with actionable follow-ups.
Engineer and operate platforms on AWS, including services such as:
EKS, EC2, RDS/Aurora, Lambda, API Gateway
CloudFront, WAF, ALB/NLB
CloudWatch, X-Ray, IAM, Secrets Manager
Lead cloud migrations and modernization initiatives, including legacy system refactoring.
Implement secure networking patterns (VPCs, private subnets, controlled egress).
Identify and resolve performance bottlenecks through testing and analysis.
Drive FinOps initiatives to optimize infrastructure cost without compromising reliability.
Implement capacity planning and autoscaling strategies.
Design and support CI/CD pipelines enabling safe, repeatable deployments.
Embed reliability practices into the SDLC (testing, rollout strategies, rollback).
Partner with development teams to improve operability of applications before production.
Partner with security and legal teams to meet regulatory and compliance requirements (e.g., data residency, GDPR-related controls).
Implement secure access controls, secrets management, and encryption best practices.
Participate in security reviews, audits, and risk assessments.
8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Systems Engineering
Strong programming experience in Python, Java, or similar languages
Deep experience with Linux/Unix systems
Hands-on expertise with AWS and cloud-native architectures
Proven experience with Terraform and Infrastructure as Code
Strong understanding of networking, security, and distributed systems
Experience operating mission-critical, high-volume platforms
Nice to Have
~1 min readExperience in financial services or highly regulated environments
Experience with EKS/Kubernetes at scale
Familiarity with Chaos Engineering and resilience testing
Experience leading cloud cost optimization (FinOps) initiatives
Prior experience transitioning traditional infrastructure teams into SRE practices
Reduced incidents and faster recovery times
Measurable reduction in operational toil
Improved system availability and performance
Clear reliability metrics aligned with business outcomes
Engineering teams empowered to ship faster with confidence
#LI-KA2
#LI-Hybrid
We are dedicated to fostering a collaborative, engaging, and inclusive environment and are committed to providing a workplace that empowers associates to be authentic and bring their best to work. We believe that associates do their best when they feel safe, understood, and valued, and we work diligently and collaboratively to ensure Broadridge is a company—and ultimately a community—that recognizes and celebrates everyone’s unique perspective.
Use of AI in Hiring
As part of the recruiting process, Broadridge may use technology, including artificial intelligence (AI)-based tools, to help review and evaluate applications. These tools are used only to support our recruiters and hiring managers, and all employment decisions include human review to ensure fairness, accuracy, and compliance with applicable laws. Please note that honesty and transparency are critical to our hiring process. Any attempt to falsify, misrepresent, or disguise information in an application, resume, assessment, or interview will result in disqualification from consideration.
Location & Eligibility
Listing Details
- First seen
- September 27, 2026
- Last seen
- September 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 56%
- Scored at
- September 27, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.