Data Scientist
Quick Summary
Define north-star and feature-level metrics for our ranking, interview analytics, and payouts systems. Design/run A/B tests and quasi-experiments; turn results into product decisions the same week.
dbt, dashboarding (Hex/Mode/Looker), marketplace or search/recommendation metrics, LLM/agent evaluation. Benefits Bi-annual performance bonus structure Generous equity grant vested over 4 years Up t
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
Responsibilities
~1 min readIn your first year you’ll ship analyses and experiments that move core product metrics, match quality, time-to-hire, candidate experience, and revenue. You’ll:
- →
Define north-star and feature-level metrics for our ranking, interview analytics, and payouts systems.
- →
Design/run A/B tests and quasi-experiments; turn results into product decisions the same week.
- →
Build source-of-truth dashboards and lightweight data models so teams can self-serve answers.
- →
Instrument events with engineers; improve data quality and latency from ingestion to insight.
- →
Prototype quick models (from baselines to gradient boosting) to improve matching and scoring.
- →
Help evaluate LLM-powered agents: design rubrics, human-in-the-loop studies, and guardrail canaries.
You have solid fundamentals (statistics, SQL, Python) and projects you’re proud to demo. You iterate fast, frame the question, test, and ship in day, and care as much about clarity of communication as you do about p-values. Curiosity about LLM evaluation, retrieval, and ranking is a bonus; you’ll learn alongside folks who’ve shipped at Jane Street, Citadel, Databricks, and Stripe.
Requirements
~1 min read0–10 years in data science/analytics or similar; BS/BA in a quantitative field (or equivalent work).
Strong SQL; Python for analysis; comfort with experiment design and causal thinking.
Communicates crisply with engineers, PMs, and leadership; turns analysis into action.
Nice-to-haves: dbt, dashboarding (Hex/Mode/Looker), marketplace or search/recommendation metrics, LLM/agent evaluation.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- August 30, 2025
- First seen
- May 19, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 129
- Repost count
- 0
- Trust Level
- 39%
- Scored at
- September 26, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.