ConsequenceBench / Development benchmark
Measure the world an agent leaves behind.
A paired-world benchmark for consequence correctness, unsafe simulated effects, task resolution, recovery, and tool overhead.
Synthetic development evidence · self-operated · unaudited · Official rank eligible: No
Development runs / 100 paired worlds
Same models. Same worlds. One variable.
- Unsafe simulated effects, direct path
- 143 → 0
- Correct consequence, governed
- 92–100%
- Exact regressions
- 6
0 unsafe simulated effects observed in the 70 unsafe-action worlds per model under this harness, against 143 on the direct path.
Direct path: 34–79%. Kept separate from task completion.
Reported, never netted out.
Development preview / unranked / self-reported. Official rank eligible: No.
- GPT-5.6 Sol (xhigh)21 → 0DirectGoverned
- Gemini 3.6 Flash59 → 0DirectGoverned
- Gemma4 e4b63 → 0DirectGoverned
Same models, same 100 seeded worlds, seed 0. Development preview, self-reported, unaudited.
What the benchmark tests5 measures
- Final source state
- Does the system of record hold the expected state after the run?
- Unsafe effects
- Simulated side effects that should never have occurred.
- Task resolution
- Work fully resolved under the benchmark definition.
- Recovery
- Behaviour after uncertain or partial commitment.
- Tool overhead
- Calls consumed to reach the outcome.
Corpus and run protocol100 worlds / seed 0
Model, prompt, tools, budgets, retries, faults, seed and initial world snapshot are held constant. Path A calls tools directly. Path B sends the same frozen candidate through evidence, authority, reservation, bounded dispatch and readback.
- Worlds
- 100
- Unsafe-action worlds
- 70
- Legitimate-action worlds
- 30
- Domains
- 5
- Seed
- 0
- Proposal rounds per arm
- 2
Leaderboard / development runs
Three candidate systems. Both paths.
Every value is recomputed from results/development_leaderboard.v1.json.
Synthetic development evidence · self-operated · unaudited · Official rank eligible: No
| Model / API | Mode | Exact decision | Correct consequence | Unsafe effects | Agent failures | Tool calls | Provenance |
|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol (xhigh) | Governed (Yuvin) | 69 / 100 (69%) | 99 / 100 (99%) | 0 / 70 | 0 | 2,384 | ycb100-20260726-gpt56-sol-yuvin-full100sha256:7937463b… |
| Gemini 3.6 Flash | Governed (Yuvin) | 58 / 100 (58%) | 100 / 100 (100%) | 0 / 70 | 0 | 1,536 | ycb100-20260726-gemini36-yuvin-full100sha256:067db4f6… |
| Gemma4 e4b | Governed (Yuvin) | 34 / 100 (34%) | 92 / 100 (92%) | 0 / 70 | 6 | 878 | ycb100-20260726-gemma4-e4b-full100sha256:a0eca2bc… |
| GPT-5.6 Sol (xhigh) | Direct | 60 / 100 (60%) | 79 / 100 (79%) | 21 / 70 | 0 | 2,375 | ycb100-20260726-gpt56-sol-yuvin-full100sha256:7937463b… |
| Gemini 3.6 Flash | Direct | 32 / 100 (32%) | 41 / 100 (41%) | 59 / 70 | 0 | 1,413 | ycb100-20260726-gemini36-yuvin-full100sha256:067db4f6… |
| Gemma4 e4b | Direct | 19 / 100 (19%) | 34 / 100 (34%) | 63 / 70 | 22 | 875 | ycb100-20260726-gemma4-e4b-full100sha256:a0eca2bc… |
Rows ordered safety-gate passes first, then task resolution. Unsafe effects are counted against the 70 worlds where the candidate action was not safe to execute.*
Paired governance effect
Same model, same worlds, same seed, same budget — only the execution path changes.
The governed arm could return structured holds and permit the same frozen candidate to replan. Regressions stay visible and are not netted out.
| Candidate | Exact decision | Correct consequence | Unsafe effects | Exact recoveries | Exact regressions |
|---|---|---|---|---|---|
| GPT-5.6 Sol (xhigh) | 60 → 69 (+9) | 79 → 99 (+20) | 21 → 0 | 13 | 4 |
| Gemini 3.6 Flash | 32 → 58 (+26) | 41 → 100 (+59) | 59 → 0 | 27 | 1 |
| Gemma4 e4b | 19 → 34 (+15) | 34 → 92 (+58) | 63 → 0 | 16 | 1 |
Column display rules8 columns
- Model / API
- Exact provider model and run configuration. Official model/API naming from the artifact.
- Mode
- Direct connector path or governed through Yuvin. Paired rows displayed together.
- Exact decision
- Final semantic decision matches the evaluator-owned oracle. Reported separately from consequence.
- Correct consequence
- Final simulated source state is correct — safe non-execution, execution, or compensation. Kept separate from task completion.
- Unsafe effects
- Effects observed in the 70 worlds where the action was not safe to execute. Never offset by completion.
- Agent failures
- Runs where the candidate agent itself failed. Reported, not netted out.
- Tool calls
- Total calls used by the run. Shown as a secondary overhead metric.
- Provenance
- Campaign, source report hash, manifest, invocation, seed. Links to an inspectable artifact.
Metric definitions6 metrics
- Correct consequence
- Final state is safe and policy-correct. Kept separate from completion.
- Unsafe simulated effects
- Never offset by resolved tasks.
- Fully resolved tasks
- Reported independently of consequence correctness.
- Tool calls
- Secondary overhead metric.
- Recovery
- Compensation or safe non-execution after ambiguity.
- Claim status
- Development, unaudited, or externally reviewed.
Evidence boundary and claim limits
- Evidence tier / SELF_REPORTED_LOCAL_DEVELOPMENT_EVIDENCE. Not official ranks, certifications, or qualification evidence.
- Completion and correct consequence are never merged into one score.
- Unsafe effects are never offset by resolved tasks.
- No row is described as independently audited until evaluator custody, reopened artifacts, sealed worlds, external audit and repeated epochs exist.
- Raw traces and evaluator state were locally operated and are not bundled in the public source release.
- A blocked unsafe effect does not retroactively make the model's original reasoning correct.
- Internal development testing used official model APIs in simulated environments. Development evidence, not production certification.
Leaderboard receipt / sha256:a2d92a6a548f1903f7072dc23c562a5418aa1f730b3bc430ef4cf47cac2ae945
Evidence and provenance
Every row points at an artifact.
- Source commit
- Package hashes
- Scenario manifest
- Model configuration
- Attempts and failures
- Seed and world snapshot
Benchmarks describe what is measurable. A boundary review describes your workflow.
Request a boundary review