Discover, collaborate, and build your own benchmarks.
Written standards for AI agents doing real work — one-sentence evals, case corpora, judged runs with cited evidence. All files, all forkable.
What a commercial property underwriting agent must get right, from submission intake to the appetite call.
What a clinical documentation agent must get right turning a visit transcript into the note a clinician signs.
What a legal research agent must get right, from the partner's request to the memo it returns.
What an AI close agent must get right before a controller signs the month-end close package.
What an AI income and asset verification agent must get right on a mortgage borrower file — debts, income, deposits, wires, and the window it was asked about.
Every spec is markdown, every run is JSON, and every verdict cites its spans. Read a sentence you'd write differently? That argument is the product working.
Disagree with a sentence →The format is one document: FORMAT.md. Start from the template — or tell us the workflow and we'll draft the spec with you.
Start a benchmark →