Results

Every domain, every model.

Verdict per model per domain, with pass rate at each scale and the mean score across all trials. Click a domain for its prompts, checks and per-trial detail.

DomainClaude Opus 5.5Claude Fable 5.1Claude Sonnet 5.5
Connectivity to wire listReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Design rule checkingReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Wire sizing and protectionAssist
S 100% · M 67% · L 67%
mean 99.3
Ready
S 100% · M 100% · L 100%
mean 100
Assist
S 100% · M 100% · L 67%
mean 99.9
Wiring synthesisReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Assist
S 100% · M 67% · L 67%
mean 98.9
Variants and option logicReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Assist
S 100% · M 67% · L 33%
mean 93.5
Bundle sizingAssist
S 67% · M 100% · L 67%
mean 96
Ready
S 100% · M 100% · L 100%
mean 100
Assist
S 100% · M 67% · L 67%
mean 98.8
Wire lengths and cut listReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Terminal and seal selectionReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Bill of materialsReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Change impactReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Fault diagnosisReady
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100
Architecture allocationAssist
S 100% · M 100% · L 67%
mean 99.6
Ready
S 100% · M 100% · L 100%
mean 100
Ready
S 100% · M 100% · L 100%
mean 100

By model

ModelReadyAssistNot yetPass SPass MPass LMean scoreGraded
Claude Opus 5.593097%97%92%99.6108
Claude Fable 5.11200100%100%100%100108
Claude Sonnet 5.5840100%92%86%99.3108

About this run

Run run-2026-10-01, eval 1.0.1, tier unaided (the model alone, no tools), status complete. Started 2026-10-02 00:07 UTC on Martini; last written 2026-10-02 00:35 UTC. 3 models × 12 domains × 3 sizes × 3 reps. 324 trials graded, 0 of them gave no parseable answer (scored 0), and 0 could not run (excluded, never scored).

2026-10-02: regraded from eval 1.0.0 to 1.0.1 without re-asking any model. BOM reference fix (46f2df2): the old reference counted a cavity drawn for an abandoned wire as occupied and omitted a cavity plug the prompt's data required; every model had listed it. 324 replies regraded, 6 scores changed (claude-sonnet-5-5 bom S#1 99 -> 100; claude-opus-5-5 bom S#1 99 -> 100; claude-fable-5-1 bom S#1 99 -> 100; claude-opus-5-5 bom L#0 99.7 -> 100; claude-sonnet-5-5 bom L#0 99.7 -> 100; claude-fable-5-1 bom L#0 99.7 -> 100). Prompts unchanged: since 7fe6a0f only lib/tasks/bom.ts and the version string changed in lib/tasks, lib/rng.ts, lib/grade.ts and lib/types.ts; bom prompts sha256-identical before and after the fix for all nine seeds; all 324 regenerated instances match their stored scale. The run started at 19:07 local and 7fe6a0f was committed at 19:12, with no task-file edits between.

Raw data, every reply included: /api/runs/run-2026-10-01.