5h ago
New
$341,400 – $401,600/yr

Sr. Staff Inference Serving Engineer

New York City Areafull-timesenior
Other
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Build and optimize serving systems that handle high request volumes with low, predictable latency Apply model compilation and graph optimization techniques,

Technical Tools
Other
Mission:
GroqCore is building a high-performance software stack that converts bare metal compute into a token-producing engine This role is for a seasoned engineer with practical experience making LLMs fast, efficient, and reliable in production. This individual will work on an inference serving software stack from a trained model to a high-throughput delivery of output tokens. This will require experience on understanding/working with model internals, runtime software, and the supporting GPU & accelerator hardware underpinning it all.

Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You’ll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.

Responsibilities & opportunities in this role:
  • Build and optimize serving systems that handle high request volumes with low, predictable latency
  • Apply model compilation and graph optimization techniques, and benchmark carefully to use them only where they deliver real gains
  • Implement and tune batching, caching, scheduling, and memory management strategies for large models
  • Apply reduced-precision and quantization methods 
  • Distribute models across multiple accelerators and nodes, balancing throughput & latency
  • Profile end-to-end performance, identify bottlenecks, and fix them—ranging from low-level kernels to request routing—leveraging both lab/development environments and live telemetry from production systems.
  • Partner with infrastructure and FDE teams to bring new models into production quickly
  • Partner with vendors along the inference serving path in support of optimizing the stack
Ideal candidates have/are:
  • BS / MS / PhD in CS, CE, EE, or equivalent depth from industry
  • 5+ years shipping performance-critical products, with some portion of that serving models at scale
  • A strong applied understanding of transformer architectures
  • Fluency in at least one systems language and one high-level language
  • A rigorous, measurement-driven approach to performance work
  • Clear communication about tradeoffs to both technical and business stakeholders
Ways to stand out:
  • Experience writing or tuning custom kernels.
  • Familiarity with non-GPU or specialized inference hardware.
  • Contributions to open-source serving or compiler projects.
Why Join Us:
  • Purposeful Hiring: You’re not here by accident, and neither is anyone else. Every teammate is handpicked with intention because who we build with matters.
  • Builders Wanted: You’re not just riding the rocket ship, you’re building it. Your work directly shapes the trajectory of our company.
  • Mission-Driven Work: We’re here to make a real impact. Our mission fuels everything we do.
  • Tackling Hard Problems: If easy isn’t your thing, you’re in the right place. We solve some of the most complex and exciting challenges in our space.
  • Excellence Is The Standard: High performance isn’t just encouraged, it’s the baseline. And it’s contagious.

If this sounds like you, we’d love to hear from you!

Compensation
Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $341,400 - $401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
#LI-MS1

Location & Eligibility

Where is the job
New York City Area
On-site at the office
Who can apply
Same as job location

Listing Details

Posted
October 9, 2026
First seen
October 9, 2026
Last seen
October 9, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
52%
Scored at
October 9, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Sr. Staff Inference Serving Engineer$341k–$402k