Quick Summary
10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances,
This role focuses on the architecture and resilience of large-scale distributed systems operating under significant production load. You will look across services, queues, databases, infrastructure, and deployment patterns to identify systemic risks before they become incidents. The role begins with a high-throughput automation platform and expands across communication and emerging product systems. You will influence critical-path architecture while remaining hands-on, spending meaningful time prototyping solutions, resolving complex failures, and shipping remediation. You will help establish capacity models, consistency guarantees, isolation boundaries, and failure-handling strategies that can withstand continued growth. This is a high-ownership environment where technical rigor, proactive problem solving, and clear cross-team communication are central to the role.
- Own the architecture health of large-scale distributed systems, including failure modes, capacity constraints, consistency guarantees, resilience, and interactions across numerous services and deployments.
- Review and influence critical-path technical designs, providing rigorous architectural guidance and making well-reasoned recommendations when teams face complex technical trade-offs.
- Proactively identify systemic risks such as single points of failure, unbounded queues, missing idempotency, thundering-herd effects, capacity constraints, and potential data-loss scenarios, then drive remediation before incidents occur.
- Build and ship solutions for complex architectural problems, including prototypes, reliability improvements, production remediation, and fixes for difficult cross-team issues.
- Design resilience into systems through degradation strategies, backpressure mechanisms, isolation boundaries, capacity models, and failure-handling patterns capable of supporting sustained growth.
- Investigate and resolve the most challenging distributed-system failures by understanding interactions across application services, messaging, databases, infrastructure, and deployment environments.
- Work hands-on with Node.js/TypeScript and/or Go to prototype architectural solutions and implement critical fixes when required.
- Establish and improve engineering practices through design reviews, post-mortems, architectural patterns, documentation, and technical guidance that can be adopted across multiple teams.
- Help define how AI-assisted engineering can be used safely and effectively when developing and operating highly critical distributed systems.
- Work with technologies including GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch across large-scale production environments.
- Help engineering teams prepare systems for substantial traffic growth by identifying capacity limits, improving observability, and validating critical failure modes.
- Influence engineers across teams through technical rigor, collaboration, mentorship, and clear communication rather than relying solely on organizational authority.
Requirements
~2 min read- 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances, or systems processing billions of events is strongly preferred.
- Demonstrated experience carrying direct accountability for production systems through significant incidents, migrations, reliability challenges, and operational failures.
- Deep expertise in queueing and asynchronous architectures, including delivery semantics, ordering, backpressure, idempotency, and the practical limitations of exactly-once processing.
- Strong knowledge of multiple storage technologies across SQL and NoSQL environments, with the ability to reason about consistency models, indexing at scale, performance, and appropriate technology selection.
- Expert-level experience with Redis or comparable in-memory systems, including behavior under memory pressure, network partitions, failover scenarios, and other failure conditions.
- Strong production experience operating Kubernetes at scale, including resource limits, autoscaling, capacity planning, node failures, and workload behavior under infrastructure disruption.
- Fluent in Node.js and/or Go, with sufficient hands-on ability to prototype technical proposals and implement production fixes on critical paths.
- Exceptional technical communication skills, including the ability to produce design documents, architectural diagrams, and root-cause analyses that drive decisions across multiple engineering teams.
- Strong systems-thinking ability, with an instinct for evaluating tail behavior, failure modes, network partitions, capacity constraints, and high-load scenarios.
- Proven ability to influence engineering teams through technical depth, constructive disagreement, clear reasoning, and respect.
- Strong ownership, curiosity, judgment, and problem-solving skills, particularly when dealing with ambiguous or cross-functional technical challenges.
- Experience using AI-assisted development tools effectively on complex systems is an advantage, particularly when balancing development speed with code quality, reliability, and operational safety.
- GCP-native experience with technologies such as Pub/Sub, Cloud Tasks, GKE, and Firestore is a strong advantage.
- Previous experience in a Staff, Principal, Architect, or comparable systems-focused engineering capacity is beneficial.
- Experience taking new systems from initial architecture through production while also hardening mature systems is a plus.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 28, 2026
- First seen
- September 28, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 28, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.