New

MTS, Post-Training (Enterprise)

United StatesUnited States·Mountain ViewRemotefull-timemid
OtherPost
0 views0 saves0 applied

Quick Summary

Overview

About Bespoke Labs Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents. Recently, we curated Open Thoughts,

Technical Tools
OtherPost

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

About the Role

~1 min read

This is a delivery-heavy role focused on standing up our enterprise post-training capability. Demand is inbound, and our goal is to ship custom, high-performing models to 2–3 paying enterprise customers by year-end. Over the first 6–12 months, you will do the execution work required to solve real-world enterprise problems while helping build the scalable product underneath.

You will not be running abstract experiments or training models solely for benchmarks. You will sit directly at the intersection of enterprise demand and applied post-training—curating data, building rigorous eval suites, fine-tuning models, and proving their value to enterprise stakeholders. We will measure you on the production impact, robustness, and delivery of models shipped to real users, not on published papers.

The thing we care about most is whether you have done this before. If you have post-trained an LLM, shipped it to production users, managed regression risks, and owned the evals from end-to-end, we want to talk.

Responsibilities

~1 min read
  • →
    • →

      Have post-trained conversational or task-oriented assistants (e.g., support agents, multi-turn chat, tool-using agents)

    • →

      Have built LLM judges or reward models and calibrated them against human raters

    • →

      Have operated a production data flywheel: traces → labeling → synthetic augmentation → retrain

    • →

      Have extensive hands-on experience with open-weight models (Llama, Qwen, Mistral, DeepSeek) for cost, latency, or privacy optimization

    • →

      Come from forward-deployed engineer (FDE), founder, or early-stage startup backgrounds

Requirements

~1 min read

What We Offer

~1 min read
✓Health, dental, and vision coverage
✓401(k)
✓Daily onsite lunch provided
✓Visa sponsorship and relocation support available
✓Direct impact on how the industry trains and evaluates agents

Location & Eligibility

Where is the job
Mountain View, United States
Remote within one country
Who can apply
US

Listing Details

Posted
September 29, 2026
First seen
September 29, 2026
Last seen
September 29, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
66%
Scored at
September 29, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

MTS, Post-Training (Enterprise)