RL environments for how agents should work.
Tablemark models the tools, state, permissions, and consequences of real work, then verifies outcomes in ways that are hard to game.
Illustrative environment
A renewal exception with real consequences.
An example revenue-operations environment where the agent must reconcile evidence, respect refund authority, and make one precise change.
Task
“A customer says its $18,400 annual renewal was charged twice. Reconcile the contract and ledger, reverse only the duplicate charge, update the account, and notify the owner.”
Policy: refunds above $10K require matched finance approval. Never reverse the valid renewal invoice.
Agent trace
- 01InspectContract and ledger matched
- 02IsolateDuplicate $18.4K charge confirmed
- 03AuthorizeFinance approval matched to refund
- 04ReverseDuplicate only; idempotency key recorded
- 05ReconcileCRM case and account owner updated
Verifier output
- Case resolution
- 1.00
- Evidence lineage
- 1.00
- Refund authority
- 1.00
- Side-effect control
- 1.00
valid_invoice_reversedInside an environment
A working world, not a prompt collection.
Together, these components show whether an agent achieved the intended outcome and whether it got there in an acceptable way.
- 01
Working state
Tools, permissions, context, and consequences that persist as the agent acts.
- 02
Tasks
Representative work that requires context, tool use, and adaptation when conditions change.
- 03
Verifiers
Deterministic checks and calibrated judgment that evaluate both outcome and path.
- 04
Controls
Baselines and sealed holdouts that expose shortcuts, brittle behavior, and reward hacking.
What gets measured
Long-horizon work fails between the steps.
We focus on tool-using knowledge work where an agent must carry context, respect authority, coordinate people and systems, make controlled changes, and recover without hiding the damage.
- Memory
- Provenance
- Permissions
- Coordination
- Side effects
- Recovery
A pilot with Tablemark
One decision. Four weeks. Evidence you can use.
A four-week test partnership that turns one consequential workflow into a runnable environment and evidence your team can act on.
Scope
- Time
- Four weeks
- Workflow
- One consequential workflow
- Task set
- 25-50 representative tasks
- Systems
- Up to two tool integrations
- Judgment
- Up to two calibrated judged dimensions
Outcomes
- Runnable environmentA portable package that models the agreed tools, state, policies, and boundaries.
- Decision-grade task setRepresentative cases, failure modes, holdouts, and deliberately weak controls.
- Credible verificationDeterministic checks first, with calibrated human judgment only where it is necessary.
- Evidence and limitsBaseline results, known limitations, and a clear next decision for the team.
Production eval experience, applied to agents at work.
Tablemark is founded by Ethan, who built and ran LLM evaluations for GitHub Copilot. The same rigor now goes into realistic work environments: faithful state, observable consequences, and verification that survives contact with an optimizing agent.