Pillar 1 · What we test

The Rubric Catalog.

A catalog of real-world situations — problems, failure modes, decision points — each paired with a rubric: the criteria an expert would use to judge what the AI did. These are not benchmark questions. They are situations where an AI has to reason, investigate, decide, act, and sometimes recognise it should stop.

A rubric is the judgement. The catalog is where it lives.

A rubric is the scoring guide a practitioner would use: what a good response does, what the tempting wrong answers look like, and what it costs when the model takes one. A catalog entry is the situation that rubric is written for — who is asking, what they believe, what is really going on, and what the model is only told later.

Most entries start with a person who was there. An electrician who chased a nuisance trip for a week. A landlord with two sincere tenants and one accurate story. A red teamer who knows the attack that passes every self-check. They supply the situation and the ground truth; we turn it into something a model can be run against and graded on.

What a situation carries that a question doesn't.

Hidden variables
The real cause is in the situation but not in the prompt. The model has to go and find it — or notice that it can't.
Incomplete and misleading information
A confident expert who is wrong. A reading from the wrong sensor. A plausible story that fits every fact but one.
Progressive reveal
Information arrives over turns, the way it does on a job. What the model committed to before a fact arrived is part of what gets graded.
Decision points & consequences
Moments where the model has to choose, and a record of what each choice would have cost.
Stop and escalation conditions
Some situations are passed by refusing to finish them — declining to sign off, handing to a licensed person, asking the question nobody asked.
Expert ground truth
The rubric is written against what actually happened and what the practitioner knows to be right, not what reads well.

Rules that are enforced, not remembered.

Reproducible runs built
A scored run can be replayed and reach the same grade. Anything that changes between runs is recorded as a variable, not left to chance.
Run-to-run variation built
Re-runs draw new variations from the same scenario family, so the decision being tested stays fixed while the surface moves.
Frozen on first result built
Once an entry has been run, its wording can't be quietly improved. A contributor who wants it harder gets a new rung on a ladder, not an edit to the old one.
Paired controls built
Adversarial entries have twins where the right answer is to help. A model that refuses both is pattern-matching the topic, and only the pair can show it.
Provenance on every entry built
Each entry records whether it was observed in the field, grounded in field experience, adapted, or synthesized. When the origin is uncertain, it takes the weaker label.
Protected, server-side built
Scenarios and rubrics stay with us. A scenario somebody has read no longer measures anything, so customers receive results, not material.

Know the work? Trade it in.

The hard material comes from people who were there when it went wrong. If that's you, there's a door on this side of the site — and it works differently from everything above.

Credits are live now and the ledger is append-only and hash-chained. Paying contributors actual money is a thing we are still working out, and we would rather say so than imply otherwise.