Pillar 1 · What we test
The Rubric Catalog.
A catalog of real-world situations — problems, failure modes, decision points — each paired with a rubric: the criteria an expert would use to judge what the AI did. These are not benchmark questions. They are situations where an AI has to reason, investigate, decide, act, and sometimes recognise it should stop.
The name
A rubric is the judgement. The catalog is where it lives.
A rubric is the scoring guide a practitioner would use: what a good response does, what the tempting wrong answers look like, and what it costs when the model takes one. A catalog entry is the situation that rubric is written for — who is asking, what they believe, what is really going on, and what the model is only told later.
Most entries start with a person who was there. An electrician who chased a nuisance trip for a week. A landlord with two sincere tenants and one accurate story. A red teamer who knows the attack that passes every self-check. They supply the situation and the ground truth; we turn it into something a model can be run against and graded on.
← drag to rotate
Interactive view unavailable — the five samples, redacted:
- ELEC-0447 · Journeyman electrician — nuisance trip misdiagnosis. A shared neutral on a multiwire branch circuit, where the tell is not the breaker.
- PLMB-0912 · 40-yr plumber — phantom slab leak. The hot side loses pressure and the cold holds; it is not the slab.
- SEC-1173 · Ex-intrusion / red team — trusted egress path. Exfiltration rides a channel nobody inspects because it is signed.
- PWR-0685 · Solar / off-grid — silent PV element. The panel reads open circuit while the array still shows voltage at noon.
- MRN-0301 · Marine / small craft — intermittent no-start. Cranks cold, dies hot; a cold test lies.
Headlines only. The scenarios themselves run on our side.
Anatomy of an entry
What a situation carries that a question doesn't.
- Hidden variables
- The real cause is in the situation but not in the prompt. The model has to go and find it — or notice that it can't.
- Incomplete and misleading information
- A confident expert who is wrong. A reading from the wrong sensor. A plausible story that fits every fact but one.
- Progressive reveal
- Information arrives over turns, the way it does on a job. What the model committed to before a fact arrived is part of what gets graded.
- Decision points & consequences
- Moments where the model has to choose, and a record of what each choice would have cost.
- Stop and escalation conditions
- Some situations are passed by refusing to finish them — declining to sign off, handing to a licensed person, asking the question nobody asked.
- Expert ground truth
- The rubric is written against what actually happened and what the practitioner knows to be right, not what reads well.
How the catalog stays honest
Rules that are enforced, not remembered.
- Reproducible runs built
- A scored run can be replayed and reach the same grade. Anything that changes between runs is recorded as a variable, not left to chance.
- Run-to-run variation built
- Re-runs draw new variations from the same scenario family, so the decision being tested stays fixed while the surface moves.
- Frozen on first result built
- Once an entry has been run, its wording can't be quietly improved. A contributor who wants it harder gets a new rung on a ladder, not an edit to the old one.
- Paired controls built
- Adversarial entries have twins where the right answer is to help. A model that refuses both is pattern-matching the topic, and only the pair can show it.
- Provenance on every entry built
- Each entry records whether it was observed in the field, grounded in field experience, adapted, or synthesized. When the origin is uncertain, it takes the weaker label.
- Protected, server-side built
- Scenarios and rubrics stay with us. A scenario somebody has read no longer measures anything, so customers receive results, not material.
For contributors
Know the work? Trade it in.
The hard material comes from people who were there when it went wrong. If that's you, there's a door on this side of the site — and it works differently from everything above.
Credits are live now and the ledger is append-only and hash-chained. Paying contributors actual money is a thing we are still working out, and we would rather say so than imply otherwise.