dragonfly-careers1mo ago
New
New
Senior Inference Optimization Engineer - Dragonfly Portfolio
Optimization EngineerData & AI
0 views0 saves0 applied
Quick Summary
Overview
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio,
Technical Tools
Optimization EngineerData & AI
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.
This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.
We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.
Location: Remote, USA (open to excellent candidates outside the USA)
What We’re Looking For:
- 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
- Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
- Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
- Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
- GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
- Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
- Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions
About the role:
- Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
- Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
- Optimize multivariate inference load-balancing algorithms within the inference routing system
- Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack
Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.
Process:
- We'll review your application and assess fit for this role.
- If there's a match, we'll facilitate a warm introduction to the team.
- If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.
Compensation
- Early Career: $110,000–$150,000
- Mid: $150,000–$200,000
- Senior: $180,000–$250,000
- Staff/Leadership: $230,000–$330,000
The compensation range reflects variation by seniority, location, and hiring company. Final compensation is confirmed with the specific hiring company.
Submit your information below, and we’ll reach out if there’s a potential fit.
Location & Eligibility
Where is the job
United States
Hybrid within the country
Who can apply
US
Listing Details
- Posted
- August 13, 2026
- First seen
- September 26, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 16%
- Scored at
- September 26, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Optimization Engineer jobs
View all →Radio Optimization Senior Engineer.Service Operation & Optimization
Radio Optimization Senior Engineer.Service Operation & Optimization
Radio Optimization Engineer.Service Operation & Optimization
Radio Optimization Senior Engineer.Service Operation & Optimization
Radio Optimization Senior Engineer.Service Operation & Optimization
Radio Optimization Senior Engineer.Service Operation & Optimization
Browse Similar Jobs
Data Science2.3kBusiness Analyst1.4kAnalytics903Machine Learning277Data Architect237Database Administrator228Business Intelligence163AI Engineer145Risk Analyst124Ai Research Engineer98Database Engineer91Business Intelligence Analyst87Research Scientist61Machine Learning Scientist58Master Data Specialist57Quantitative Analyst55Fraud Analyst48Data Quality Analyst36Data Modeler27Product Data Analyst26
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.