Software Engineer, Agents
Quick Summary
About Mercor Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models.
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
About the Role
~1 min readWe're looking for a strong engineer who can build agentic products that scale. You will work with:
Responsibilities
~1 min read- →
Own agentic features end-to-end — from scoping with researchers/ops partners through implementation, launch, and iteration on real customer feedback.
- →
Design and ship LLM agents, harnesses, and verifiers — including the tools, prompts, and policies that make them reliable.
- →
Build the Python/FastAPI services and Temporal/Modal pipelines that orchestrate agent runs, human-in-the-loop review and iterations.
- →
Build state of the art RL environments that expand the capabilities of frontier agents, with realistic enterprise apps, simulated coworkers, and rich company data rooms that support tasks spanning hours to days.
- →
Build tooling that turns agent trajectories into insight, from statistical analysis to automated failure mode detection.
- →
Build and refine the full-stack surfaces and data infrastructure — craft Next.js/React interfaces where operators and experts work with agents, evolve data models to give agents the structured context and audit trails they need.
- →
Define agent quality and drive continuous improvement — build evals, instrument traces, analyze failure modes, and iterate on prompts, tools, and guardrails while raising the bar for reliability, cost, latency, and UX.
- →
Partner cross-functionally to shape agent autonomy — work with Product, Design, Research and Ops to draw the lines between autonomous action, propose-and-approve flows, and human-in-the-loop decisions.
Impact: Your work powers how the world’s leading AI labs train and test their models.
Learning: Get early insights into frontier model capabilities months before the market.
Growth: Work on both infrastructure and research-adjacent projects with fast paths to ownership.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- August 27, 2026
- First seen
- September 25, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 38%
- Scored at
- September 26, 2026
Signal breakdown
Similar Software Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.