Capital Eval
Can AI do Capital?
Twelve domains of electrical and wire-harness engineering, from net lists to fault diagnosis. Every problem generated fresh. Every answer graded by code. No partial credit for a harness that is almost right.
12 of 12 domains.
Ready today, with no tools at all.
Best result per domain across Claude Opus 5.5, Claude Fable 5.1, Claude Sonnet 5.5. The model alone: no calculator, no code, no Capital.
A clean sweep is a ceiling, not a finish line. At up to 140 wires this benchmark no longer separates the best model from perfect. Real harnesses run to thousands of wires, with messier data. Raising the ceiling is the next run.
The assertion
AI can already do much of what Capital does. So test it.
Not by asking a model what it knows about Capital, but by handing it the engineering Capital performs and checking the answer the way a release review would: every wire, every gauge, every fuse.
Generated, not memorised.
Each problem is built from a seed at run time. There is no answer key on the internet to recall, and a new seed makes a new problem.
Graded by code, not opinion.
A reference solver computes the right answer and a grader compares, check by check. No model ever judges another model.
Three scales, three runs.
Eight wires, forty, a hundred and forty. Each run three times. Ready means right at every scale, nine times in ten.
Twelve domains. One question each.
Each maps to a Capital module and a page of its documentation, so a score says which part of the toolset it speaks for.
Connectivity to wire list
Can it turn a logical net list into the physical wires that implement it?
Design rule checking
Can it find every rule violation in a wiring design, and nothing else?
Wire sizing and protection
Can it size a wire and its fuse against current, voltage drop and ampacity together?
Wiring synthesis
Can it route each signal along the cheapest path through a harness network with costed constraints?
Variants and option logic
Can it resolve a 150% design into the exact wire set for each buildable configuration?
Bundle sizing
Can it compute every branch diameter of a harness the way Capital Harness Designer does?
Wire lengths and cut list
Can it turn a routed harness into manufacturing cut lengths?
Terminal and seal selection
Can it pick the right terminal and seal for every cavity from a parts catalogue?
Bill of materials
Can it roll a harness design up into a correct, complete bill of materials?
Change impact
Can it work out exactly what an engineering change touches?
Fault diagnosis
Given symptoms and a few voltmeter readings, can it find the one open circuit?
Architecture allocation
Can it place functions onto controllers so every constraint holds and network traffic is minimal?
Where it breaks is scale.
Pass rate, every domain together, as the problem grows. A model that is flawless on eight wires and wrong on one wire in a hundred and forty has not replaced the tool that handles ten thousand.
Claude Opus 5.5
Claude Fable 5.1
Claude Sonnet 5.5
Do you need Capital to test this?
Not for the question asked.
The assertion is about the engineering, not the software. Sizing a wire, routing a signal or finding an open circuit has a right answer whether or not Capital computes it, so a solver and a grader settle it. Leaving Capital out also keeps the test honest: the model cannot lean on the tool it is being compared with.
Yes, for the two questions after it.
Capital earns its place twice. As a referee, it runs the same problem so the reference solver is checked against the real tool, not only against itself. And as a stage: an agent driving Capital through its API, judged by Capital’s own design rule checks. That is what replacing Capital in production would actually mean. The three tiers
What this does not show.
Capital is also a system of record, a component library, a multi-user database, drawings, formboards and the thread into Teamcenter and NX. None of that is measured here. These twelve domains test the engineering judgement and computation inside the tool, given clean data and stated rules. A Ready here is evidence about that, and only about that.
Every task, rule, solver and grader is in the repository, and every stored answer can be regraded. Read the method.