Member of Technical Staff, Inference & Serving
Bay Areafull-timelead
OtherSoftware EngineerMember Of Technical StaffSoftware Engineering
1 views0 saves0 applied
Quick Summary
Overview
The Role We're looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective,
Technical Tools
OtherSoftware EngineerMember Of Technical StaffSoftware Engineering
The Role
We're looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable.
Key Responsibilities
- Build and optimize high-performance model serving systems for low-latency inference of diffusion LLMs.
- Extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving.
- Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
- Build systems for model versioning, canary deployments, and zero-downtime rollouts.
- Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.
- Collaborate with ML researchers to translate model advances (new architectures, quantization techniques, batching strategies) into production-ready serving improvements.
Qualifications
- BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
- Knowledge of ML serving frameworks (SGLang, vLLM, Triton Inference Server, TensorRT-LLM).
- Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.
- Familiarity with high-performance computing and GPU programming (CUDA).
- Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.
- Background in performance optimization and profiling of ML systems.
Preferred Skills
- Experience building and maintaining large-scale language models with tens of billions of parameters or more.
- Experience with distributed systems and cloud computing platforms (AWS/GCP/Azure).
- Experience with ML workflow orchestration tools (Kubeflow, Airflow).
- Experience with model optimization techniques (quantization, distillation, speculative decoding, continuous batching).
- Knowledge of ML-specific infrastructure challenges (checkpointing, resource scheduling, etc.).
What We Offer
~1 min readThe annual base salary range for this role is $200,000 – $350,000 USD. Final compensation is determined based on experience, skills, and qualifications. Equity and benefits are included in the total package.
Location & Eligibility
Where is the job
Bay Area
On-site at the office
Who can apply
Same as job location
Listing Details
- Posted
- March 10, 2026
- First seen
- September 26, 2026
- Last seen
- October 10, 2026
Posting Health
- Days active
- 13
- Repost count
- 0
- Trust Level
- 20%
- Scored at
- October 10, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Software Engineer jobs
View all →Remote
Senior Member of Technical Staff
$180k–$250k/yr
Full-timeRemote
Software Engineer- Farsi speaking
IT PROGRAMMER ANALYST, LEAD/ ADVANCED (FULL-TIME CONTRACTUAL)
$41.83 - $54.12/ hour, with potential growth to $65.97/ hour
Staff Software Engineer
Software Engineer, Core Systems
USD 170000-380000
Remote
Senior Software Engineer, Security Factory: Code Security
USD 139200-235200
Remote
Browse Similar Jobs
Application Developer548Integration Developer422Java Developer422Salesforce Developer234.Net Developer221Python Developer157Build Engineer141Solutions Architect135Search Engineer123Java Software Engineer116Laravel Developer115C++ Developer91Security Software Engineer83Full Stack Developer63Robotics Software Engineer61Database Developer56Php Developer55Wordpress Developer48SAP ABAP Developer47Python Software Engineer44
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.