ABOUT MERCOR
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
ABOUT DEEPTUNE
Deeptune Environments is a Mercor team building the platform and tooling that AI labs use to train agents through reinforcement learning. We build the systems that let researchers turn real-world tasks into training environments at scale.
ABOUT THE ROLE
You'll build the systems our high-fidelity simulation environments run on, the APIs and tool interfaces agents interact with, the grading layer that decides whether an agent actually succeeded, the pipelines that turn raw human demonstrations into training-ready environments, and the orchestration and sandboxing that run it all at scale.
You'll do this alongside researchers who own the RL side. Your job is to make their ideas real: fast, correct, and reliable at scale. That means you should know enough about post-training, evals, and reward modeling to push back on a research spec productively. It does not mean you'll be training models.
The work is high-ownership and lightly specified. You'll set direction, drive outcomes, and stay hands-on. You'll also work directly with AI labs and enterprise partners, which means shipping against real external deadlines rather than internal ones.
WHAT YOU'LL DO
- Build the systems that build environments end-to-end: the simulated app or system, the agent-facing tool surface, the task definitions, and the verifiers that score them.
- Design and operate the backend infrastructure that runs environments at scale (containers, orchestration, queues, observability) — standard distributed-systems backend work, applied to a new domain.
- Turn messy human data into clean, reproducible training environments.
- Care about reliability and speed as much as correctness.
- Own the interface with labs and researchers: translate a research goal into a system that exists next week.
WHAT WE'RE LOOKING FOR
We care more about what you've built than how long you've been building it.
- 2+ years of full-stack / backend engineering, including at least 1 year at a startup, ideally as a founding engineer, an early engineer at a fast-growing venture, or a founder yourself
- Strong generalist with systems depth. We primarily use Python and seek engineers who are fluent in at least one language. Ideally, they should be skilled at applying agents with discernment and able to quickly adapt to different problem requirements.
- Comfortable around ML/LLM concepts. Not an ML background, but enough working knowledge of post-training, evals, and reward modeling to partner with researchers and translate their specs into systems.
- Thrives in ambiguity. You scope your own work, make pragmatic calls, and ship without a spec handed to you.
- Ability to raise the bar around you. Coaching engineers and driving execution, while staying in the code
Bonus, not required: sandboxing or virtualization, browser and computer-use automation, CI/build systems, developer tooling.
YOU'LL FIT HERE IF
- Ownership, impact, and building frontier tech are what motivate you
- Your work is a craft you want to master
- You thrive in ambiguity and like hard problems
- You appreciate diverse perspectives and uncommon ideas
- You're excited to build in person, 5 days a week, 10 am–8 pm ET, from our office at One World Trade
Mercor
AI · Series C · San Francisco, USA