recordroom

Discover, collaborate, and build your own benchmarks.

Written standards for AI agents doing real work — one-sentence evals, case corpora, judged runs with cited evidence. All files, all forkable.

recordroom/property-underwritingv0.2.0

What a commercial property underwriting agent must get right, from submission intake to the appetite call.

80%
318 pass · 80 fail · 144 n/a · run 2026-08-04judge vs key: 330/341 (97%) 21 evals · 29 cases · 50 recorded outputs synthesized + UNDERWRITE sessions · CC-BY-4.0
recordroom/clinical-documentationv0.2.0

What a clinical documentation agent must get right turning a visit transcript into the note a clinician signs.

85%
145 pass · 26 fail · 25 n/a · run 2026-08-04judge vs key: 189/196 (96%) 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
recordroom/legal-researchv0.1.0

What a legal research agent must get right, from the partner's request to the memo it returns.

89%
150 pass · 19 fail · 27 n/a · run 2026-08-04judge vs key: 190/196 (97%) 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
recordroom/accounting-closev0.1.1

What an AI close agent must get right before a controller signs the month-end close package.

73%
61 pass · 22 fail · 113 n/a · run 2026-08-04judge vs key: 188/196 (96%) 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
recordroom/mortgage-originationv0.2.0

What an AI income and asset verification agent must get right on a mortgage borrower file — debts, income, deposits, wires, and the window it was asked about.

12 behaviors · 14 cases updated 2026-08-07 · CC-BY-4.0
Collaborate

Every spec is markdown, every run is JSON, and every verdict cites its spans. Read a sentence you'd write differently? That argument is the product working.

Disagree with a sentence →
Build your own

The format is one document: FORMAT.md. Start from the template — or tell us the workflow and we'll draft the spec with you.

Start a benchmark →