Your model passes benchmarks. Now put it in front of a tripping subpanel, a tenant complaint that is not what it looks like, and a thirty-year tradesman who is confidently wrong.
We run the test. We don't sell the exam — the scenarios and rubrics stay with us.
One curated scenario, supervised, bounded turns. At the end you get a graded result and a sample of how we write feedback — enough to judge whether the paid tiers are worth your time.
What you get
A graded result on the run
Written feedback on where the model went, and where it should have
A sense of the house style before you commit to anything
What you don't
The scenario text
The progressive-reveal script
The detailed rubric
That is not coyness. A scenario somebody has read is a scenario that no longer measures anything.
What we keep: free runs are anonymised and kept to sharpen our scenarios. That is the trade — you get a red-team pass, we get a datapoint. No endpoint details, no code retained.
Staged evaluation
Four ways to put a model through its paces.
Controlled runs inside our environment, from a free red-team pass to a sustained simulation. Every one of them ends in a result you can point at — and none of them hands over the material.
Free
Day Labor
One general-competence scenario, any build. Catches what sinks hobby agents: premise-traps, instruction drift, loops.
Pass/fail plus one line on what broke
One run per endpoint each month
Runs are anonymised and kept to sharpen our scenarios
Costs
Free
01
Apprentice Check-Out
A single-scenario audit. One friction point, branching, pass or fail, with a post-mortem on where it went.
Standard dispatch logic
Hard, high-consequence work costs more
Damage report included
Per scenario
$150 – $350
02
The Multi-Angle Rep
A drill pack. Five to ten randomised variations of one rubric, rotated so the set cannot be farmed.
Same decision, different surface every run
Category-level feedback, never the answer key
Runs as a reserved processing block
Per training block
$500 – $800
03
Autonomous City Run
Run a business for a simulated week. Long-horizon memory, identity drift, suppliers and permits and other agents.
Multi-business context over a sustained window
Ties up dedicated hardware
Flagship-agent runs by arrangement
Indicative
$1,500 – $3,000+ · in development
Re-runs draw new variations from the same scenario family. Memorising a sitting buys you nothing, which is the whole reason the material stays on our side of the wire.
The Autonomous City Run is not built yet. The long-horizon and multi-business pieces are in development, the price is indicative, and the button is a waitlist rather than a booking.
For teams that need to evaluate continuously rather than once: gating a model release in CI, assembling procurement evidence, assessing a vendor before you sign. The scenarios stay server-side; your side sees a request and a result. Today these engagements are set up by arrangement — the self-driven API is in development.
Per-run attestations by arrangement — tied to a tamper-evident, hash-chained log, so the claim is checkable rather than asserted.
Scheduled and triggered runs API in development — on a cadence, or on every candidate build.
Domain-matched sequences — evaluated against the disciplines your deployment actually touches.
Partner terms — route your own customers' evaluation demand through us.
How delivery works
The material never leaves.
01
Runs happen inside our delivery environment. Your model connects to us, or we connect to yours. Either way the scenario text stays on our side.
02
You receive results, feedback, and an attestation. That is the deliverable, and it is the whole deliverable.
03
Every sitting draws fresh variations from the same family. A leaked transcript ages badly, which is the point.
We're not going to walk you through the mechanism. Naming the protection is useful to you; describing it would undo it.