Quick Summary
Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal.
Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.
Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.
We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.
We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
We are looking for Lead ML Engineer -LLM Inference & NLP.
- Speed up and scale LLM inference in production: SGLang, KV and prefix caching, batching, quantization, speculative decoding
- Run distributed inference for very large models (up to 1T+ parameters) across multi-GPU and multi-node setups
- Benchmark new GPU servers and hardware, bring them into production and adapt our serving code to them
- Lead the NLP and CV teams technically: review experiments, set direction, and step in early when something is heading the wrong way
- Train and fine-tune the language models, and improve the agent harnesses and chat algorithm that run on them
- Track cutting-edge research and open-source work in inference and post-training, and turn it into the ML roadmap
- Collaborate closely with the validation, content, and dataset preparation teams to design experiments and measure model quality
- Deep hands-on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT-LLM
- Experience with distributed inference or training of large models: MoE, tensor/expert/pipeline parallelism, multi-node GPU clusters
- Strong understanding of what makes inference fast: KV cache, attention kernels, batching, quantization, GPU profiling
- Experience training and fine-tuning LLMs, including post-training (RLHF, DPO, or similar)
- Proven technical leadership: you've guided engineers through reviews, mentoring and technical decisions while still writing code yourself
- Proficiency with PyTorch, transformers, and related libraries
- Experience at AI-focused startups or companies (Character AI, OpenAI, and similar is a strong plus)
- Backend engineering experience (Python, Go, C#) and knowledge of scalable deployment systems is a significant advantage
- Advanced English or Russian
Nice to Have
~1 min read- CUDA or Triton kernel development
- A computer vision background. We also welcome strong CV leads who have accelerated large generative image or video models
- Experience with multimodal LLMs
- First-author papers or notable open-source work, e.g. contributions to SGLang, vLLM, or post-training libraries
- A degree in CS, math, or physics from a strong program (MSc or PhD)
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 27, 2026
- First seen
- October 3, 2026
- Last seen
- October 3, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 36%
- Scored at
- October 3, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.