Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
Quick Summary
About Plaud Inc. Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,500,000 users worldwide since 2023.
Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud captures, structures, and compounds the intelligence generated in conversations — so humans can think better, decide faster, and execute with clarity.
Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human–AI intelligence through a hardware–software combination. With full ISO 27001, ISO 27701, SOC 2, GDPR, EN18031, and HIPAA compliances, Plaud is committed to the highest standards of data security and privacy protection.
To learn more about Plaud, please visit https://www.Plaud.ai and follow along on Instagram, X, Facebook, LinkedIn, and YouTube
Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think.
Plaud is a bootstrapped, skyrocketing, profitable company with a $300M revenue run rate achieved in just three years.
Define the next-gen paradigm for human-AI interaction.
Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion.
Work with passionate teammates who value innovation, collaboration, and customer success.
Grow your career in a culture that champions continuous learning and fast career development.
Market-competitive compensation, global exposure, and a vibrant, creativity-fueled work atmosphere.
Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.
Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments.
Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.
Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks.
Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team.
Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs.
Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity.
Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).
Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency.
Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill.
Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.
Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- May 8, 2026
- First seen
- September 25, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 16%
- Scored at
- September 26, 2026
Signal breakdown
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.