About Scrimdata
Scrimdata builds frontier RL environments from real company data. We license the operational history of real companies, run it through our own anonymization pipeline, and turn it into multi-step, multi-tool training environments: tasks, rubrics, verifiers, and reference trajectories, graded against what actually happened. We're an early-stage, remote-first team selling to two sides of one market: AI labs and enterprise agent teams who train on our environments, and data partners who supply the raw material. Small team, hard problems, real customers from day one.
About the role
We're hiring a Forward Deployed Researcher to sit between our product and the labs training on it. You'll embed with customers, understand what capability they're actually trying to build, and turn that into environment specs our pipeline can produce: the tasks, the difficulty tiers, the pass criteria. When a customer's pilot stalls because a reward signal is noisy or a verifier is too strict, you're the person who digs in, finds the cause, and fixes it, in the field, on their timeline. This is a technical, customer-facing role for someone who wants their research to ship into a real training run within weeks, not get filed as a recommendation.
What you'll do
- Embed directly with AI-lab and enterprise customers to understand their training and eval objectives, and translate them into concrete environment specs
- Scope, build, and run pilots: define the task set, the rubric, the verifier logic, and the difficulty curve for a given capability
- Debug environments in production: trace a bad reward signal, a mis-scoped rubric, or an unreliable verifier back to its root cause and fix it
- Review agent trajectories against reference trajectories to diagnose where a model's behavior diverges from real outcomes
- Work across multiple customer engagements at once, moving at the pace of a live pilot rather than a quarterly roadmap
- Feed what you learn in the field back into the core product: better task-mining heuristics, better verifier patterns, better rubric templates
- Partner with the environment-construction and anonymization teams to make sure what gets built off a digital twin actually matches what the customer needs
What we're looking for
- Strong applied research or engineering background, ideally with hands-on RL, evals, or agent-training experience
- Comfortable reading and writing code well enough to debug a verifier or reshape a task pipeline yourself, not just describe the problem to someone else
- Experience working directly with technical customers: scoping ambiguous asks, running structured pilots, and being accountable for outcomes
- Sharp, fast diagnostic instincts. You can take "the agent's reward looks wrong" and find the actual cause
- Clear written and verbal communication. You'll be explaining technical tradeoffs to both researchers and non-technical stakeholders
- Comfortable with ambiguity and a fast-changing environment. We're early-stage, and the job will change under you
- Willingness to travel occasionally for on-site customer work when a pilot calls for it
Nice to have
- Prior experience at an AI lab, an evals or RL-environments company, or a data-labeling operation
- Familiarity with RL post-training loops (GRPO and similar), eval harnesses, or containerized task delivery
- Experience designing rubrics or programmatic verifiers for open-ended, multi-step tasks
- Background in a specific operational domain (finance, legal, support, sales ops, engineering) that maps to enterprise workflows
Benefits
Meaningful early equity. Remote-first, work from wherever you do your best work. Standard benefits package.
How to apply
Email careers@scrimdata.com with your resume and a short note on why this role fits. No forms, no links.