isone1mo ago
New
New
Senior Site Reliability Engineer
senior
EngineeringDevops Engineer
0 views0 saves0 applied
Quick Summary
Requirements Summary
H-1B, F-1/CPT/OPT, O-1, E-3, TN
Technical Tools
EngineeringDevops Engineer
The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions.
This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.
What we offer you:
A stable, mission-driven workplace where your impact truly matters
A highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeing
Competitive compensation with a base salary + performance bonus
Robust benefits package, including:
Enhanced 401(k) and financial planning support
Tuition reimbursement and professional development
Wellness programs, including an onsite gym
Flexible work hours
Employee Business Networks
Free coffee at our onsite café
Hybrid work environment (3 days/week onsite)
Distance-based relocation assistance available
How you will make an Impact
Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)
Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows
Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead
Develop meaningful KPIs and dashboards for business and IT service health
Engineer and implement resilience patterns including HA, DR, and automated failover
Partner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage
Conduct performance testing, capacity modeling, forecasting, and right-sizing
Participate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements
Identify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions
Design, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency
Design, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability
Identify gaps in observability coverage and drive engineering solutions to close them
Collaborate with architecture and application teams to ensure production readiness
Leverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness
Reduce repeat incidents by engineering permanent fixes and driving continuous improvement
What we are looking for
5+ years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations
Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation.
Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices
Strong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments
Demonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environment
Ability to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms
Knowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms
Familiarity with SRE principles including SLOs, error budgets, observability, and toil reduction
Ability to analyze and troubleshoot complex technical systems
Preferred Qualifications
Experience in mission-critical, highly available, or regulated environments
Experience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activities
Knowledge of ITIL processes and/or SRE best practices
Experience with performance testing, capacity planning, resilience testing, or disaster recovery validation
This employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.).
The expected salary range for this position is $134,000 - $170,000 per year, for a Senior to Lead level candidate. This role is also eligible for an annual performance bonus, comprehensive health insurance (medical, dental and vision), flexible spending and health savings accounts, a 401(k) plan with generous employer contributions and a student debt benefit, life and AD&D insurance, disability insurance, critical illness and hospital indemnity benefits, paid time off, paid leave, a wellness program, an employee assistance program and other great company perks.
#LI-HYBRID
Location & Eligibility
Where is the job
—
Location terms not specified
Listing Details
- Posted
- August 17, 2026
- First seen
- September 26, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 11%
- Scored at
- September 26, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Devops Engineer jobs
View all →Staff Cloud Engineer
$170k–$200k/yr
Site Reliability Engineer
Full-Time
S
SkyePoint DecisionsRemoteDevSecOps/Platform Engineer
Remote
Senior Site Reliability Engineer
Senior Site Reliability Engineer
full-timeRemote
Senior Cloud Engineer (GovCloud / FedRAMP)
Browse Similar Jobs
Engineering Manager1.9kFullstack Developer1.9kSoftware Architect1.8kSecurity1.7kQa Engineer1.6kMechanical Engineer1.4kBackend Developer1.4kSecurity Engineer1.3kDevOps & Infrastructure1.3kElectrical Engineer1.3kProject Engineer1.1kFrontend Developer1kDesign Engineer925Mobile Developer671Backend Engineering645Product Engineer621Process Engineer595Quality Engineer590Data Engineering570Embedded Engineer496
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.