Observability - Site Reliability Engineer
Quick Summary
This position is fully onsite in Chantilly, VA
Responsibilities
~1 min readWe are seeking Observability / Site Reliability Engineers to design, implement, and operate telemetry and reliability capabilities across enterprise infrastructure, platform services, applications, and workloads.
The engineers will implement centralized metrics, logs, traces, dashboards, alerts, and reliability practices using OpenTelemetry or equivalent industry standards and approved enterprise tooling. This role combines observability engineering, site reliability practices, production operations, automation, and performance analysis to improve system health, availability, and operational resilience.
Implement standardized collection of metrics, logs, and distributed traces across enterprise environments.
Establish OpenTelemetry-based instrumentation and telemetry collection patterns where appropriate.
Integrate cloud, Kubernetes, platform, application, and infrastructure telemetry into centralized observability services.
Maintain telemetry pipelines, retention, routing, and data-quality controls.
Develop and maintain instrumentation and telemetry configurations to support effective system monitoring and troubleshooting.
Develop dashboards, alerts, service-health views, and operational metrics to provide visibility into system performance and availability.
Define and monitor Service Level Objectives (SLOs), availability indicators, latency, error rates, saturation, and capacity measures.
Analyze system performance, capacity, utilization, and operational trends to identify potential reliability issues and corrective actions.
Apply site reliability engineering practices to improve system availability, performance, scalability, and operational resilience.
Identify opportunities to improve detection, response, and recovery across enterprise services.
Support incident diagnosis, cross-service troubleshooting, and root-cause analysis for complex reliability and observability issues.
Automate recurring monitoring, alerting, reporting, and remediation activities where appropriate.
Partner with operations and engineering teams to improve incident detection and response and reduce recurring issues.
Participate in Tier 3/4 escalation and on-call support for observability, platform, and reliability issues.
Develop and maintain operational procedures, troubleshooting documentation, and knowledge articles.
Support continuous improvement of monitoring and reliability practices based on operational data and incident trends.
Location: This position is fully onsite in Chantilly, VA
Requirements
~1 min readActive TS/SCI clearance with CI Polygraph.
Approximately 5–8 years of experience in observability, site reliability engineering, operations, platform engineering, cloud engineering, or a related technical discipline.
- 5 years experience with a BS/BA, an additional 4 years of experience may be considered in lieu of a degree.
Hands-on experience with metrics, logging, tracing, monitoring, and alerting.
Experience with Kubernetes and cloud environments.
Strong troubleshooting, systems analysis, and performance-analysis skills.
Experience operating and supporting production services and responding to technical incidents.
Experience with scripting, automation, or infrastructure-as-code practices.
Strong analytical, problem-solving, and communication skills.
Experience with one or more of the following:
OpenTelemetry.
Prometheus and Grafana.
CloudWatch, OpenSearch, or comparable cloud monitoring and observability platforms.
Distributed tracing and application performance monitoring.
SLO, SLA, and error-budget practices.
Observability for large-scale or distributed enterprise platforms.
Automated incident response or remediation.
Observability within secure or highly regulated enterprise environment
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.
Location & Eligibility
Listing Details
- Posted
- October 11, 2026
- First seen
- October 11, 2026
- Last seen
- October 11, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- October 11, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.