Senior Software Engineer
Quick Summary
design, implementation, rollout, and production operation. Build backend services and the data pipelines that generate, validate, score, and version datasets, reproducible from a commit and a config.
Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com.
About the Role
~1 min readTuring is the world's leading research accelerator for frontier AI labs. The STEM Horizontal team produces the high-signal STEM data those labs train on, and this role builds the machinery behind it: the task-authoring tools, data pipelines, evaluation harnesses, sandboxed execution environments, and agentic workflows our researchers and domain experts work inside every day.
This sits at the seam between engineering and research. You'll own production systems end-to-end, and you'll be the person a researcher pulls in to ask why an eval is producing misleading numbers. Correctness and reproducibility are the product: a subtle bug here doesn't crash anything, it quietly poisons a training run and costs weeks. If you want to build systems that visibly shape how frontier models learn to reason about code and science, this is the seat.
Responsibilities
~1 min read- →Own services end-to-end: design, implementation, rollout, and production operation.
- →Build backend services and the data pipelines that generate, validate, score, and version datasets, reproducible from a commit and a config.
- →Ship data-dense internal interfaces: virtualised tables, server-side filtering, review and annotation UIs that stay fast at tens of thousands of rows.
- →Build and operate agent harnesses, multi-step loops with tool use, retries, and structured output, running against real repos and test suites.
- →Build sandboxed execution environments where model-generated code runs safely and deterministically at volume, and stand up the RL environments and eval harnesses on top.
- →Trace and debug agent runs end-to-end: where the loop stalled, which tool call failed, why the grader disagreed with the human.
- →Own infrastructure as code and CI/CD; debug production issues across the stack and drive the reliability work that follows.
- →Raise the bar through code review, design docs, and innovation, partnering with researchers and quality owners on what “good” means.
Requirements
~1 min read- 6+ years building and operating production software, ideally at product companies with real scale.
- Deep Python with production FastAPI and/or Django, ORM performance, migrations, async patterns.
- Strong PostgreSQL (schema design, query optimisation, indexing) and NoSQL.
- Frontend production with TypeScript / React / Next.js.
- Fluent on Linux and Docker; Terraform or comparable IaC on AWS or GCP, and ownership of the CI/CD that ships it.
- Daily use of LLM coding tools (Claude Code, Cursor, Copilot) as part of how you actually ship, with the judgement to review their output critically and know where they break down. We are an AI-forward team; this is a must-have.
- You've built something on top of LLM APIs that other people depended on, an agent, a pipeline, an eval, and you know where it was fragile.
- Comfort with non-deterministic systems: you can tell a real regression from noise, and you reach for a controlled experiment over a hunch.
- Advanced Git and real testing discipline, unit, integration, and the negative and edge cases people skip.
- Clear written communication; comfortable working without hand-holding, and comfortable saying when a spec is wrong.
Nice to Have
~1 min read- Internal tools or developer platforms used daily by a technical team.
- Agent frameworks and protocols, LangGraph, MCP, OpenAI/Anthropic tool use.
- Sandboxing and untrusted code execution: gVisor, Firecracker, seccomp.
- Workflow orchestration (Temporal, Airflow, Prefect, Dagster) and Kubernetes beyond managed defaults.
- LLM coding benchmarks (SWE-bench, Terminal-Bench), eval frameworks, or exposure to RLHF / RLVR.
Bachelor's or Master's degree with 6+ years experience in Computer Science, Software Engineering, Data Science, Machine Learning, AI, or a programming-heavy IT field.
We weight demonstrated work outcomes most heavily, a strong track record without a matching degree is fine.
- We are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value.
- We work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection
- We are AI forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity.
- Work at the frontier of AI, helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.
- Contribute to leading-edge AI research and showcase your work at top conferences such as ICLR, ICML, and NeurIPS.
- Bring frontier AI innovation to the enterprise, applying lessons learned from leading AI labs to solve real-world business challenges.
- Collaborate with and learn from exceptional colleagues with deep AI experience from Google, Meta, Amazon, and other leading technology companies.
- Move at the pace of AI innovation, with the speed, ownership, and impact of a startup.
Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace and celebrate authenticity, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.
For applicants from the European Union, please review Turing's GDPR notice here.
Location & Eligibility
Listing Details
- Posted
- September 28, 2026
- First seen
- September 28, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- September 28, 2026
Signal breakdown
Similar Software Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.