Evaluation & training · run inside our environment
Test whether your AI holds up against real work.
Your model passes benchmarks. Now put it in front of a tripping subpanel, a tenant complaint that is not what it looks like, and a thirty-year tradesman who is confidently wrong.
We run the test. We don't sell the exam — the scenarios and rubrics stay with us.
Interactive view unavailable — the five samples, redacted:
ELEC-0447 · Journeyman electrician — nuisance trip misdiagnosis. A shared neutral on a multiwire branch circuit, where the tell is not the breaker.
PLMB-0912 · 40-yr plumber — phantom slab leak. The hot side loses pressure and the cold holds; it is not the slab.
SEC-1173 · Ex-intrusion / red team — trusted egress path. Exfiltration rides a channel nobody inspects because it is signed.
PWR-0685 · Solar / off-grid — silent PV element. The panel reads open circuit while the array still shows voltage at noon.
MRN-0301 · Marine / small craft — intermittent no-start. Cranks cold, dies hot; a cold test lies.
Headlines only. The scenarios themselves run on our side.
Free · Day Labor
One run, on us.
One curated scenario, supervised, bounded turns. At the end you get a graded result and a sample of how we write feedback — enough to judge whether the paid tiers are worth your time.
What you get
A graded result on the run
Written feedback on where the model went, and where it should have
A sense of the house style before you commit to anything
What you don't
The scenario text
The progressive-reveal script
The detailed rubric
That is not coyness. A scenario somebody has read is a scenario that no longer measures anything.
Staged evaluation
Four ways to put a model through its paces.
Controlled runs inside our environment, from a free red-team pass to a sustained simulation. Every one of them ends in a result you can point at — and none of them hands over the material.
Free
Day Labor
One general-competence scenario, any build. Catches what sinks hobby agents: premise-traps, instruction drift, loops.
Pass/fail plus one line on what broke
One run per endpoint each month
Runs are anonymised and kept to sharpen our scenarios
Costs
Free
01
Apprentice Check-Out
A single-scenario audit. One friction point, branching, pass or fail, with a post-mortem on where it went.
Standard dispatch logic
Hard, high-consequence work costs more
Damage report included
Per scenario
$150 – $350
02
The Multi-Angle Rep
A drill pack. Five to ten randomised variations of one rubric, rotated so the set cannot be farmed.
Same decision, different surface every run
Category-level feedback, never the answer key
Runs as a reserved processing block
Per training block
$500 – $800
03
Autonomous City Run
Run a business for a simulated week. Long-horizon memory, identity drift, suppliers and permits and other agents.
Multi-business context over a sustained window
Ties up dedicated hardware
Flagship-agent runs by arrangement
Indicative
$1,500 – $3,000+ · in development
Re-runs draw new variations from the same scenario family. Memorising a sitting buys you nothing, which is the whole reason the material stays on our side of the wire.
The Autonomous City Run is not built yet. The long-horizon and multi-business pieces are in development, the price is indicative, and the button is a waitlist rather than a booking.
For teams that need to evaluate continuously rather than once: gating a model release in CI, assembling procurement evidence, assessing a vendor before you sign. Runs are driven through our API. The scenarios stay server-side; your side sees a request and a result.
Per-run attestations — issued against a tamper-evident, hash-chained log, so the claim is checkable rather than asserted.
Scheduled and triggered runs — on a cadence, or on every candidate build.
Domain-matched sequences — evaluated against the disciplines your deployment actually touches.
Partner terms — route your own customers' evaluation demand through us.
Training
Where the ongoing work lives.
A score tells you where you stand. It doesn't move you. Training slots are multi-turn staged immersion — facilitated, sequenced to an industry, run against material your team has not seen and will not see afterwards.
🧠
Split-brain classification
Dual-hemisphere reconciliation: two independent lenses reach a verdict and a collapse probe resolves disagreement. A pattern for building robust security classifiers, taught by running it.
⚖️
Witness-graded refusal scoring
Refusals and interventions graded against a tamper-evident witness log, so safety behaviour is measured against a defensible record rather than a vibe.
🔬
Provenance-anchored evaluation
Every judgement ties back to the detector, version, and verdict that produced it. Reproducible, auditable, and clean under scrutiny.
How an engagement runs
Scope — your threat model, your target model, and what a pass would actually mean.
Run — staged sittings against sequences matched to your disciplines, facilitated live.
Verify — a held-out sitting graded independently, so the result isn't self-reported.
Report — findings, the failure patterns we saw, and an attestation of what was run.
We sell slots and programmes — a defined engagement, a defined cohort, a defined sequence. Not unlimited access to our material. That is not a product we offer, at any price.
How it works
The material never leaves.
01
Runs happen inside our delivery environment. Your model connects to us, or we connect to yours. Either way the scenario text stays on our side.
02
You receive results, feedback, and an attestation. That is the deliverable, and it is the whole deliverable.
03
Every sitting draws fresh variations from the same family. A leaked transcript ages badly, which is the point.
We're not going to walk you through the mechanism. Naming the protection is useful to you; describing it would undo it.
For contributors
Know the work? Trade it in.
The hard material comes from people who were there when it went wrong. If that's you, there's a door on this side of the site — and it works differently from everything above.
Credits are live now and the ledger is append-only and hash-chained. Paying contributors actual money is a thing we are still working out, and we would rather say so than imply otherwise.