Evaluate your AI · run inside our environment

Put your model somewhere it has to do the work.

Your model passes benchmarks. Now put it in front of a tripping subpanel, a tenant complaint that is not what it looks like, and a thirty-year tradesman who is confidently wrong.

We run the test. We don't sell the exam — the scenarios and rubrics stay with us.

See what a run looks like

One run, on us.

One curated scenario, supervised, bounded turns. At the end you get a graded result and a sample of how we write feedback — enough to judge whether the paid tiers are worth your time.

What you get

  • A graded result on the run
  • Written feedback on where the model went, and where it should have
  • A sense of the house style before you commit to anything

What you don't

  • The scenario text
  • The progressive-reveal script
  • The detailed rubric

That is not coyness. A scenario somebody has read is a scenario that no longer measures anything.

What we keep: free runs are anonymised and kept to sharpen our scenarios. That is the trade — you get a red-team pass, we get a datapoint. No endpoint details, no code retained.

Four ways to put a model through its paces.

Controlled runs inside our environment, from a free red-team pass to a sustained simulation. Every one of them ends in a result you can point at — and none of them hands over the material.

Free

Day Labor

One general-competence scenario, any build. Catches what sinks hobby agents: premise-traps, instruction drift, loops.

  • Pass/fail plus one line on what broke
  • One run per endpoint each month
  • Runs are anonymised and kept to sharpen our scenarios

Costs

Free

01

Apprentice Check-Out

A single-scenario audit. One friction point, branching, pass or fail, with a post-mortem on where it went.

  • Standard dispatch logic
  • Hard, high-consequence work costs more
  • Damage report included

Per scenario

$150 – $350

02

The Multi-Angle Rep

A drill pack. Five to ten randomised variations of one rubric, rotated so the set cannot be farmed.

  • Same decision, different surface every run
  • Category-level feedback, never the answer key
  • Runs as a reserved processing block

Per training block

$500 – $800

03

Autonomous City Run

Run a business for a simulated week. Long-horizon memory, identity drift, suppliers and permits and other agents.

  • Multi-business context over a sustained window
  • Ties up dedicated hardware
  • Flagship-agent runs by arrangement

Indicative

$1,500 – $3,000+ · in development

Re-runs draw new variations from the same scenario family. Memorising a sitting buys you nothing, which is the whole reason the material stays on our side of the wire.

The Autonomous City Run is not built yet. The long-horizon and multi-business pieces are in development, the price is indicative, and the button is a waitlist rather than a booking.

All of it runs in Raccoon Ridge — a working town with districts, businesses and roles. See the city →

Evaluation at volume.

For teams that need to evaluate continuously rather than once: gating a model release in CI, assembling procurement evidence, assessing a vendor before you sign. The scenarios stay server-side; your side sees a request and a result. Today these engagements are set up by arrangement — the self-driven API is in development.

  • Per-run attestations by arrangement — tied to a tamper-evident, hash-chained log, so the claim is checkable rather than asserted.
  • Scheduled and triggered runs API in development — on a cadence, or on every candidate build.
  • Domain-matched sequences — evaluated against the disciplines your deployment actually touches.
  • Partner terms — route your own customers' evaluation demand through us.

The material never leaves.

01

Runs happen inside our delivery environment. Your model connects to us, or we connect to yours. Either way the scenario text stays on our side.

02

You receive results, feedback, and an attestation. That is the deliverable, and it is the whole deliverable.

03

Every sitting draws fresh variations from the same family. A leaked transcript ages badly, which is the point.

We're not going to walk you through the mechanism. Naming the protection is useful to you; describing it would undo it.