J
Jobgether3d ago
New

Member of Technical Staff | Observability & Reliability

United StatesUnited StatesRemoteFull-timelead
OtherMember Of Technical Staff
0 views0 saves0 applied

Quick Summary

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps.

Technical Tools
OtherMember Of Technical Staff

This is a high-impact platform engineering role focused on making distributed systems observable, reliable, and operationally resilient.
You’ll own observability across cloud and customer-hosted environments, ensuring teams can understand system health wherever workloads run.
The role spans logs, metrics, traces, alerting, SLOs, incident response, and reliability engineering.
You’ll work closely with Kubernetes-based infrastructure and software operating across environments you may not fully control.
Your work will directly support high availability, faster incident resolution, and consistent deployment health.
You’ll also help reduce telemetry costs by improving the quality and efficiency of the signals collected.
This is an autonomous, hands-on position where you’ll build, operate, and continuously improve critical platform capabilities.

  • Evolve and maintain the observability platform covering logs, metrics, traces, alerting, and system health across cloud and customer-hosted dataplanes.

  • Ensure every environment reports critical operational information, including active releases, health status, heartbeats, logs, metrics, and usage to the central control plane.

  • Implement telemetry collection within customer Kubernetes environments using outbound-only connectivity models.

  • Detect and investigate differences between desired infrastructure or deployment state and what is actually running in each environment.

  • Monitor the health and availability of deployment and runtime agents, including ephemeral workloads such as Ray clusters supporting batch inference.

  • Define and maintain Service Level Objectives (SLOs), establish actionable alerting, and contribute to error-budget practices.

  • Lead or participate in incident response and postmortems, identifying improvements that reduce recurring failures and mean time to recovery (MTTR).

  • Coordinate incident resolution across internal teams and customers when fixes involve customer-managed environments.

  • Optimize telemetry pipelines to reduce redundant data, control infrastructure costs, and improve the signal-to-noise ratio of operational information.

  • Write production-quality code, review technical changes, and take operational ownership of the systems you build.

Requirements

~1 min read
  • Deep professional experience with OpenTelemetry and modern observability platforms or backends.

  • Hands-on experience defining SLOs, working with error budgets, designing actionable alerts, and managing production incidents.

  • Strong experience with Kubernetes and infrastructure-as-code tools such as Terraform and Helm.

  • Experience operating software across distributed or customer-hosted environments where infrastructure and connectivity may not be fully under your control.

  • Strong software engineering fundamentals, with experience producing maintainable, production-ready code and conducting effective code reviews.

  • Willingness to participate in operational ownership, troubleshooting, incident response, and continuous reliability improvements.

  • Strong analytical and problem-solving abilities, with an ability to investigate complex distributed-system behavior.

  • Experience communicating clearly across engineering teams and, when required, working directly with external customers or stakeholders.

  • Experience with GCP/GKE or AWS/EKS is an advantage.

  • Familiarity with multi-node or multi-cluster ML workloads in production is a plus.

  • Experience deploying software to customer-hosted Kubernetes environments, including Helm-based deployments and outbound-only connectivity, is valuable.

  • Experience in financial services or other regulated environments is an additional advantage.

What We Offer

~2 min read
✓Fully remote role based in Brazil.
✓Full-time position within an engineering-focused environment.
✓Opportunity to own critical observability and reliability systems with direct impact on production availability.
✓Work across cloud and customer-hosted Kubernetes environments, providing broad exposure to distributed infrastructure.
✓Opportunity to work with modern observability, Kubernetes, infrastructure-as-code, and ML infrastructure technologies.
✓High degree of technical ownership, with responsibility for both building and operating the systems you develop.
✓Exposure to complex reliability challenges involving real-time services, batch workloads, customer environments, and distributed systems.
✓Opportunity to contribute to incident management, platform architecture, and long-term reliability practices.

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

Posted
September 23, 2026
First seen
September 27, 2026
Last seen
September 27, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
September 27, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

J
Member of Technical Staff | Observability & Reliability