Careers / Open roles

Senior Software Engineer: Agentic Coding RL Environments

Scrim Data · Remote · Full-time

About Scrimdata

Scrimdata builds frontier RL environments from real company data. We license the operational history of real companies, run it through our own anonymization pipeline, and turn it into multi-step, multi-tool training environments: tasks, rubrics, verifiers, and reference trajectories, graded against what actually happened. We're an early-stage, remote-first team selling to two sides of one market: AI labs and enterprise agent teams who train on our environments, and data partners who supply the raw material. Small team, hard problems, real customers from day one.

About the role

We're hiring a Senior Software Engineer to build the coding-agent RL environments at the center of what Scrimdata sells. You'll work inside anonymized company twins, real repositories, real ticket queues, real code review threads, and turn that engineering history into tasks a frontier model can train on: multi-step, multi-tool problems with a deterministic grader behind every one. You'll design the verifiers that score an agent's trajectory against what the engineering team actually shipped, and you'll be the one who catches an agent gaming a check instead of solving the underlying bug. This is a build role, not a research-adjacent one. Your tasksets need to run reproducibly, in containers, at rollout scale, against a lab's real post-training loop, not just look good in a demo.

What you'll do

  • Mine real engineering histories inside anonymized company twins, repos, commit history, tickets, code review threads, for tasks that reflect how software actually gets built and broken
  • Design multi-step, multi-tool coding tasks with deterministic graders and programmatic verifiers, scored against the real outcome preserved as a reference trajectory
  • Harden environments against reward hacking: find the shortcuts a model will take to pass a check without solving the underlying problem, then close them
  • Package tasksets as reproducible, containerized environments that run reliably at RL rollout scale, not just in a one-off demo
  • Calibrate difficulty across pass@k tiers so a curriculum can run from warm-up tasks to genuinely hard ones
  • Work directly with lab customers' post-training teams to understand the capability they're training for and translate it into environment specs
  • Review agent trajectories against reference trajectories to diagnose exactly where a model's behavior diverges from what a real engineer did

What we're looking for

  • Professional software engineering experience shipping real, production systems, ideally at meaningful scale
  • Strong judgment about what good code looks like and where an agent's output falls short of it
  • Comfort designing and writing tests, rubrics, and verifiers, not just application code
  • Experience with containerized environments and the discipline to make a task reproducible across runs
  • Deep fluency in at least one major stack (TypeScript, Python, Go, Rust, Java, or similar) and enough range to read code outside it
  • Independent and self-directed, comfortable owning a problem from spec to shipped environment
  • Clear written communication. You'll be documenting task specs and grading logic for people who didn't build them

Nice to have

  • Experience building or maintaining SWE-bench-style benchmarks, coding evals, or agent-training environments
  • Familiarity with RL post-training loops (GRPO and similar) and how reward signal quality affects them
  • Experience red-teaming model output or adversarially probing an eval for exploits
  • Background in data anonymization, security-sensitive data pipelines, or working with real customer data under strict handling rules

Benefits

Meaningful early equity. Remote-first, work from wherever you do your best work. Standard benefits package.

How to apply

Email careers@scrimdata.com with your resume and links to code or systems you're proud of. No forms.