Evaluation & training · run inside our environment

Test whether your AI
holds up against real work.

Your model passes benchmarks. Now put it in front of a tripping subpanel, a tenant complaint that is not what it looks like, and a thirty-year tradesman who is confidently wrong.

We run the test. We don't sell the exam — the scenarios and rubrics stay with us.

See what it costs

One run, on us.

One curated scenario, supervised, bounded turns. At the end you get a graded result and a sample of how we write feedback — enough to judge whether the paid tiers are worth your time.

What you get

  • A graded result on the run
  • Written feedback on where the model went, and where it should have
  • A sense of the house style before you commit to anything

What you don't

  • The scenario text
  • The progressive-reveal script
  • The detailed rubric

That is not coyness. A scenario somebody has read is a scenario that no longer measures anything.

Four ways to put a model through its paces.

Controlled runs inside our environment, from a free red-team pass to a sustained simulation. Every one of them ends in a result you can point at — and none of them hands over the material.

Free

Day Labor

One general-competence scenario, any build. Catches what sinks hobby agents: premise-traps, instruction drift, loops.

  • Pass/fail plus one line on what broke
  • One run per endpoint each month
  • Runs are anonymised and kept to sharpen our scenarios

Costs

Free

01

Apprentice Check-Out

A single-scenario audit. One friction point, branching, pass or fail, with a post-mortem on where it went.

  • Standard dispatch logic
  • Hard, high-consequence work costs more
  • Damage report included

Per scenario

$150 – $350

02

The Multi-Angle Rep

A drill pack. Five to ten randomised variations of one rubric, rotated so the set cannot be farmed.

  • Same decision, different surface every run
  • Category-level feedback, never the answer key
  • Runs as a reserved processing block

Per training block

$500 – $800

03

Autonomous City Run

Run a business for a simulated week. Long-horizon memory, identity drift, suppliers and permits and other agents.

  • Multi-business context over a sustained window
  • Ties up dedicated hardware
  • Flagship-agent runs by arrangement

Indicative

$1,500 – $3,000+ · in development

Re-runs draw new variations from the same scenario family. Memorising a sitting buys you nothing, which is the whole reason the material stays on our side of the wire.

The Autonomous City Run is not built yet. The long-horizon and multi-business pieces are in development, the price is indicative, and the button is a waitlist rather than a booking.

All of it runs in Raccoon Ridge — a working town with districts, businesses and roles. See the city →

API-mediated runs, at volume.

For teams that need to evaluate continuously rather than once: gating a model release in CI, assembling procurement evidence, assessing a vendor before you sign. Runs are driven through our API. The scenarios stay server-side; your side sees a request and a result.

  • Per-run attestations — issued against a tamper-evident, hash-chained log, so the claim is checkable rather than asserted.
  • Scheduled and triggered runs — on a cadence, or on every candidate build.
  • Domain-matched sequences — evaluated against the disciplines your deployment actually touches.
  • Partner terms — route your own customers' evaluation demand through us.

Where the ongoing work lives.

A score tells you where you stand. It doesn't move you. Training slots are multi-turn staged immersion — facilitated, sequenced to an industry, run against material your team has not seen and will not see afterwards.

🧠

Split-brain classification

Dual-hemisphere reconciliation: two independent lenses reach a verdict and a collapse probe resolves disagreement. A pattern for building robust security classifiers, taught by running it.

⚖️

Witness-graded refusal scoring

Refusals and interventions graded against a tamper-evident witness log, so safety behaviour is measured against a defensible record rather than a vibe.

🔬

Provenance-anchored evaluation

Every judgement ties back to the detector, version, and verdict that produced it. Reproducible, auditable, and clean under scrutiny.

How an engagement runs

  1. Scope — your threat model, your target model, and what a pass would actually mean.
  2. Run — staged sittings against sequences matched to your disciplines, facilitated live.
  3. Verify — a held-out sitting graded independently, so the result isn't self-reported.
  4. Report — findings, the failure patterns we saw, and an attestation of what was run.

We sell slots and programmes — a defined engagement, a defined cohort, a defined sequence. Not unlimited access to our material. That is not a product we offer, at any price.

The material never leaves.

01

Runs happen inside our delivery environment. Your model connects to us, or we connect to yours. Either way the scenario text stays on our side.

02

You receive results, feedback, and an attestation. That is the deliverable, and it is the whole deliverable.

03

Every sitting draws fresh variations from the same family. A leaked transcript ages badly, which is the point.

We're not going to walk you through the mechanism. Naming the protection is useful to you; describing it would undo it.

Know the work? Trade it in.

The hard material comes from people who were there when it went wrong. If that's you, there's a door on this side of the site — and it works differently from everything above.

Credits are live now and the ledger is append-only and hash-chained. Paying contributors actual money is a thing we are still working out, and we would rather say so than imply otherwise.