Results
Every domain, every model.
Verdict per model per domain, with pass rate at each scale and the mean score across all trials. Click a domain for its prompts, checks and per-trial detail.
| Domain | Claude Opus 5.5 | Claude Fable 5.1 | Claude Sonnet 5.5 |
|---|---|---|---|
| Connectivity to wire list | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Design rule checking | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Wire sizing and protection | Assist S 100% · M 67% · L 67% mean 99.3 | Ready S 100% · M 100% · L 100% mean 100 | Assist S 100% · M 100% · L 67% mean 99.9 |
| Wiring synthesis | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Assist S 100% · M 67% · L 67% mean 98.9 |
| Variants and option logic | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Assist S 100% · M 67% · L 33% mean 93.5 |
| Bundle sizing | Assist S 67% · M 100% · L 67% mean 96 | Ready S 100% · M 100% · L 100% mean 100 | Assist S 100% · M 67% · L 67% mean 98.8 |
| Wire lengths and cut list | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Terminal and seal selection | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Bill of materials | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Change impact | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Fault diagnosis | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
| Architecture allocation | Assist S 100% · M 100% · L 67% mean 99.6 | Ready S 100% · M 100% · L 100% mean 100 | Ready S 100% · M 100% · L 100% mean 100 |
By model
| Model | Ready | Assist | Not yet | Pass S | Pass M | Pass L | Mean score | Graded |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 9 | 3 | 0 | 97% | 97% | 92% | 99.6 | 108 |
| Claude Fable 5.1 | 12 | 0 | 0 | 100% | 100% | 100% | 100 | 108 |
| Claude Sonnet 5.5 | 8 | 4 | 0 | 100% | 92% | 86% | 99.3 | 108 |
About this run
Run run-2026-10-01, eval 1.0.1, tier unaided (the model alone, no tools), status complete. Started 2026-10-02 00:07 UTC on Martini; last written 2026-10-02 00:35 UTC. 3 models × 12 domains × 3 sizes × 3 reps. 324 trials graded, 0 of them gave no parseable answer (scored 0), and 0 could not run (excluded, never scored).
2026-10-02: regraded from eval 1.0.0 to 1.0.1 without re-asking any model. BOM reference fix (46f2df2): the old reference counted a cavity drawn for an abandoned wire as occupied and omitted a cavity plug the prompt's data required; every model had listed it. 324 replies regraded, 6 scores changed (claude-sonnet-5-5 bom S#1 99 -> 100; claude-opus-5-5 bom S#1 99 -> 100; claude-fable-5-1 bom S#1 99 -> 100; claude-opus-5-5 bom L#0 99.7 -> 100; claude-sonnet-5-5 bom L#0 99.7 -> 100; claude-fable-5-1 bom L#0 99.7 -> 100). Prompts unchanged: since 7fe6a0f only lib/tasks/bom.ts and the version string changed in lib/tasks, lib/rng.ts, lib/grade.ts and lib/types.ts; bom prompts sha256-identical before and after the fix for all nine seeds; all 324 regenerated instances match their stored scale. The run started at 19:07 local and 7fe6a0f was committed at 19:12, with no task-file edits between.
Raw data, every reply included: /api/runs/run-2026-10-01.