About Scrimdata
Scrimdata builds frontier RL environments from real company data. We license the operational history of real companies, run it through our own anonymization pipeline, and turn it into multi-step, multi-tool training environments: tasks, rubrics, verifiers, and reference trajectories, graded against what actually happened. We're an early-stage, remote-first team selling to two sides of one market: AI labs and enterprise agent teams who train on our environments, and data partners who supply the raw material. Small team, hard problems, real customers from day one.
About the role
We're hiring a Research Scientist to prove that real-work environments make better agents than synthetic ones, and to build the methods that make each new environment worth training on. You'll design how we mine tasks out of a digital twin, calibrate difficulty so pass@k tiers mean something, and build the verifiers and reward functions that turn a messy real outcome into a reliable training signal. When we make a public claim, a benchmark result, a transfer study, an essay about why real data wins, you're the person who ran the experiment and can defend the number. This is a hands-on research role: less about writing a memo and more about running the RL post-training experiment that proves an environment is worth what we charge for it.
What you'll do
- Design and refine task-mining methods that pull realistic multi-step, multi-tool tasks out of anonymized digital twins
- Build difficulty-calibration methods so pass@k tiers reflect real task difficulty, not artifacts of how a task happened to be written
- Design, build, and validate verifiers and reward functions that grade agent trajectories against reference trajectories reliably at RL scale
- Run controlled studies on contamination and transfer: does training on a real-work environment produce gains that hold on out-of-distribution benchmarks, and does it beat a synthetic environment built to look the same
- Run RL post-training experiments using GRPO-family methods on our own environments to measure and prove environment value before it ships to a customer
- Publish essays, benchmark results, and methodology writeups that establish our research credibility in public
- Partner with the environment-construction, anonymization, and forward-deployed teams to turn what you learn into better rubric templates, verifier patterns, and difficulty heuristics across the catalog
What we're looking for
- Strong research background in machine learning, reinforcement learning, or LLM evaluation, with a track record of running real experiments, not just reading about them
- Hands-on experience with RL post-training methods, GRPO or a similar policy-optimization approach, and the infrastructure to run them
- Experience designing rubrics, reward functions, or verifiers for open-ended, multi-step, multi-tool tasks
- Comfortable with statistics and experiment design: you know how to tell a real transfer effect from noise
- Strong technical writing. You can turn an experiment into an essay or a benchmark writeup that holds up to scrutiny
- Deep curiosity about data: how task selection, difficulty calibration, and reward design change what a model actually learns
- Comfortable with ambiguity and a fast-changing environment; we're early-stage and the research agenda will shift as we learn
Nice to have
- Prior experience at an AI lab, an evals or RL-environments company, or a research team publishing on agentic benchmarks
- Familiarity with contamination studies, benchmark saturation, or memorization detection in LLM training data
- Experience with anonymization, privacy-preserving data pipelines, or working with sensitive real-world datasets
- A public research record: papers, benchmarks, or widely read technical essays
Benefits
Meaningful early equity. Remote-first, work from wherever you do your best work. Standard benefits package.
How to apply
Email careers@scrimdata.com with your resume and a short note on relevant work. No forms, no links.