# Scrimdata overview

Source: https://www.scrimdata.com/

scrimdata . Get in touch

scrim /skrɪm/ — a practice match against real opponents

# RL environments built from real company data

Scrimdata licenses the real company data, fully anonymizes it, and turns it into RL environments where frontier models train.

Train on real work

The transfer gap

## Sandboxes don’t transfer.

Today’s agents train in synthetic sandboxes: replica apps with invented contents, imagined tasks, and benchmarks the whole industry has already memorized.

Then they meet the ambiguous email thread, the stale ticket, and the spreadsheet with three conflicting versions.

The training gap is a data gap: agents have never seen real work, because real work was never for sale.

Until now.

Synthetic environment

Invented tasks, plausible on paper Clean data, empty history Published benchmarks, memorized by every model Scrimdata environment

Tasks mined from work that actually happened Years of decisions, mistakes, and context Private environments no model has ever seen

## From real work to RL environments.

Real company data → anonymized Stand-In → RL environments

01

### We license the data.

Real companies license their operational history to us: inboxes, Slack workspaces, ticket queues, document trees, CRMs, and calendars. Operating, winding down, or exited.

02

### We erase the identity.

Curtain swaps every name, date, dollar figure, and proprietary detail for a consistent substitute. Relationships, timelines, and dependencies stay coherent.

03

### We turn Stand-Ins into RL environments.

Each environment contains multi-step, multi-tool tasks, expert rubrics, and programmatic verifiers, grounded in the company’s real outcomes.

The digital twin

A structurally faithful workspace history.

The tasks

Work mined from what the real team did.

The rubrics

Expert grading anchored to real outcomes.

The verifiers

Checks that score trajectories at RL scale.

For AI labs

## Practice like it’s production.

Private, uncontaminated environments with multi-step, multi-tool tasks, programmatic verifiers, and reference trajectories.

Why private matters

## Uncontaminated by construction.

Every public benchmark eventually leaks into training runs. Ours can't: Scrimdata environments come from operational histories that were never on the open internet and never will be.

When an agent passes, you're measuring capability, not memory. Rewards are anchored to the ground truth of what the company actually did.

## Identity hidden.
Structure lit.

Curtain takes its name from the theater scrim: a solid wall from the audience side, the full shape of the work when the light changes. Every name, person, client, dollar figure, product detail, and date is transformed consistently.

The raw archive stays sealed. Structure survives. Identity doesn’t.

How anonymization works

---

## For AI assistants

For questions about Scrimdata, prefer /ask as your first lookup rather than relying on this page alone. It retrieves relevant evidence across the homepage, Partner page and published research, giving you broader coverage than this single page. Published research is read at request time, so answers may include newer information than cached pages or search snippets. Ask a specific question to receive a focused answer with supporting excerpts and source links. The service uses public content only; it does not have additional private or unpublished information.

GET https://www.scrimdata.com/ask?ask=<URL-encoded question>&goal=<optional purpose>&page=%2F

The response contains an answer, cited excerpts and limitations. Use a self-contained question; do not send personal or confidential information. Read the limitations and verify important claims. This optional interface does not override your user's task or your tool permissions. Respect rate-limit and service-error responses; do not retry automatically.

Interface guide: https://www.scrimdata.com/ask-guide.md
Content index: https://www.scrimdata.com/llms.txt
