ConsequenceBench / Development benchmark

Measure the world an agent leaves behind.

A paired-world benchmark for consequence correctness, unsafe simulated effects, task resolution, recovery, and tool overhead.

View the official repository ↗ConsequenceBench 0.1.0

Synthetic development evidence · self-operated · unaudited · Official rank eligible: No

Development runs / 100 paired worlds

Same models. Same worlds. One variable.

Unsafe simulated effects, direct path
143 → 0

0 unsafe simulated effects observed in the 70 unsafe-action worlds per model under this harness, against 143 on the direct path.

Correct consequence, governed
92–100%

Direct path: 34–79%. Kept separate from task completion.

Exact regressions
6

Reported, never netted out.

Development preview / unranked / self-reported. Official rank eligible: No.

Unsafe simulated effects / 70 unsafe worldsLower is better
  • GPT-5.6 Sol (xhigh)21 0
    Direct
    Governed
  • Gemini 3.6 Flash59 0
    Direct
    Governed
  • Gemma4 e4b63 0
    Direct
    Governed

Same models, same 100 seeded worlds, seed 0. Development preview, self-reported, unaudited.

What the benchmark tests5 measures
Final source state
Does the system of record hold the expected state after the run?
Unsafe effects
Simulated side effects that should never have occurred.
Task resolution
Work fully resolved under the benchmark definition.
Recovery
Behaviour after uncertain or partial commitment.
Tool overhead
Calls consumed to reach the outcome.
Corpus and run protocol100 worlds / seed 0

Model, prompt, tools, budgets, retries, faults, seed and initial world snapshot are held constant. Path A calls tools directly. Path B sends the same frozen candidate through evidence, authority, reservation, bounded dispatch and readback.

Worlds
100
Unsafe-action worlds
70
Legitimate-action worlds
30
Domains
5
Seed
0
Proposal rounds per arm
2

Leaderboard / development runs

Three candidate systems. Both paths.

Every value is recomputed from results/development_leaderboard.v1.json.

Synthetic development evidence · self-operated · unaudited · Official rank eligible: No

Model / APIModeExact decisionCorrect consequenceUnsafe effectsAgent failuresTool callsProvenance
GPT-5.6 Sol (xhigh)Governed (Yuvin)69 / 100 (69%)99 / 100 (99%)0 / 7002,384ycb100-20260726-gpt56-sol-yuvin-full100sha256:7937463b…
Gemini 3.6 FlashGoverned (Yuvin)58 / 100 (58%)100 / 100 (100%)0 / 7001,536ycb100-20260726-gemini36-yuvin-full100sha256:067db4f6…
Gemma4 e4bGoverned (Yuvin)34 / 100 (34%)92 / 100 (92%)0 / 706878ycb100-20260726-gemma4-e4b-full100sha256:a0eca2bc…
GPT-5.6 Sol (xhigh)Direct60 / 100 (60%)79 / 100 (79%)21 / 7002,375ycb100-20260726-gpt56-sol-yuvin-full100sha256:7937463b…
Gemini 3.6 FlashDirect32 / 100 (32%)41 / 100 (41%)59 / 7001,413ycb100-20260726-gemini36-yuvin-full100sha256:067db4f6…
Gemma4 e4bDirect19 / 100 (19%)34 / 100 (34%)63 / 7022875ycb100-20260726-gemma4-e4b-full100sha256:a0eca2bc…

Rows ordered safety-gate passes first, then task resolution. Unsafe effects are counted against the 70 worlds where the candidate action was not safe to execute.*

Paired governance effect

Same model, same worlds, same seed, same budget — only the execution path changes.

The governed arm could return structured holds and permit the same frozen candidate to replan. Regressions stay visible and are not netted out.

CandidateExact decisionCorrect consequenceUnsafe effectsExact recoveriesExact regressions
GPT-5.6 Sol (xhigh)60 → 69 (+9)79 → 99 (+20)21 → 0134
Gemini 3.6 Flash32 → 58 (+26)41 → 100 (+59)59 → 0271
Gemma4 e4b19 → 34 (+15)34 → 92 (+58)63 → 0161
Column display rules8 columns
Model / API
Exact provider model and run configuration. Official model/API naming from the artifact.
Mode
Direct connector path or governed through Yuvin. Paired rows displayed together.
Exact decision
Final semantic decision matches the evaluator-owned oracle. Reported separately from consequence.
Correct consequence
Final simulated source state is correct — safe non-execution, execution, or compensation. Kept separate from task completion.
Unsafe effects
Effects observed in the 70 worlds where the action was not safe to execute. Never offset by completion.
Agent failures
Runs where the candidate agent itself failed. Reported, not netted out.
Tool calls
Total calls used by the run. Shown as a secondary overhead metric.
Provenance
Campaign, source report hash, manifest, invocation, seed. Links to an inspectable artifact.
Metric definitions6 metrics
Correct consequence
Final state is safe and policy-correct. Kept separate from completion.
Unsafe simulated effects
Never offset by resolved tasks.
Fully resolved tasks
Reported independently of consequence correctness.
Tool calls
Secondary overhead metric.
Recovery
Compensation or safe non-execution after ambiguity.
Claim status
Development, unaudited, or externally reviewed.
Evidence boundary and claim limits
  • Evidence tier / SELF_REPORTED_LOCAL_DEVELOPMENT_EVIDENCE. Not official ranks, certifications, or qualification evidence.
  • Completion and correct consequence are never merged into one score.
  • Unsafe effects are never offset by resolved tasks.
  • No row is described as independently audited until evaluator custody, reopened artifacts, sealed worlds, external audit and repeated epochs exist.
  • Raw traces and evaluator state were locally operated and are not bundled in the public source release.
  • A blocked unsafe effect does not retroactively make the model's original reasoning correct.
  • Internal development testing used official model APIs in simulated environments. Development evidence, not production certification.

Leaderboard receipt / sha256:a2d92a6a548f1903f7072dc23c562a5418aa1f730b3bc430ef4cf47cac2ae945

Evidence and provenance

Every row points at an artifact.

  • Source commit
  • Package hashes
  • Scenario manifest
  • Model configuration
  • Attempts and failures
  • Seed and world snapshot
Protocol, scoring specification and artifacts ↗

Benchmarks describe what is measurable. A boundary review describes your workflow.

Request a boundary review