Method · eval 1.0.1

How a claim becomes a number.

The assertion under test: a capable model, on its own, can do much of the engineering that Siemens Capital performs. Here is exactly how that is measured, and where the measurement stops.

Seedtask · size · repGenerateinstance + rulesModelanswers in JSONGradechecks vs. solverVerdictper domain

1. Twelve domains, chosen from the product

Each domain is a piece of work Capital does for an engineer, mapped to the module that does it and the documentation article that describes it: connectivity to wire list, design rule checking, wire sizing and protection, wiring synthesis, variants and option logic, bundle sizing, wire lengths and cut list, terminal and seal selection, bill of materials, change impact, fault diagnosis, architecture allocation. The list deliberately spans the suite, from Logic Designer through Systems Integrator and Harness Designer to E/E Publisher and Systems Modeler.

2. Problems are generated, never stored

Every instance is built at run time from a seed derived from (domain, size, repetition). The same seed always gives the same problem, so every model faces identical instances and anyone can regenerate one. A new seed gives a problem nobody has seen, so there is nothing to memorise and nothing to leak into training data.

The rules are given in full: the wire table, the fuse ratings, the cost model, the packing factors. The test is whether the model can apply engineering rules correctly and completely at scale, not whether it remembers a particular standard or Capital’s defaults.

3. Three sizes, three repetitions

Small, medium and large: roughly 8, 40 and 140 wires, or the equivalent in loads, cases or cavities. Each is run three times with different seeds. Real harnesses run to thousands of wires, so the large size is still small; it is there to show the direction of travel.

4. Graded by code

For every domain a reference solver computes the right answer from the same instance, and a grader compares the model’s JSON with it, check by check. A domain has two to six checks with weights summing to 100: wires matched, gauges right, routes optimal, violations found with no false alarms. Set answers are scored by F1, so flagging everything scores as badly as flagging nothing.

A trial passes only at 100. A harness with one wire wrong is a defective harness. Partial credit is shown as the score, but it never counts as a pass.

Three outcomes are kept apart:

  • Graded: the reply parsed and was scored, possibly 0.
  • No answer: no JSON could be found. Scored 0, because answering in a usable form is part of the job.
  • Could not run: the call failed (timeout, transport, quota). Excluded from every average and counted on the results page. A broken pipe is not a model failing.

5. Verdicts

  • Ready: passes at least 90% of trials at every size. Right, reliably, as the problem grows.
  • Assist: not Ready, but the mean score is at least 80. Most of the work is right; an engineer reviews and corrects.
  • Not yet: below that.
  • Unmeasured: a size with no graded trial. Never folded into another verdict.

The headline counts, per domain, the best verdict any tested model reached.

6. Keeping the instrument honest

  • Known-good and known-bad controls. For every domain and size, the reference answer must score exactly 100 and a deliberately corrupted one must score less. The suite fails otherwise.
  • Hand-worked cases. The first controls cannot catch a wrong solver, because a wrong solver grades itself as right. So each solver also has worked examples computed by hand. That gap was found by mutation: disabling one DRC rule passed the first suite untouched.
  • Tied to Capital’s documentation. Capital’s Harness Designer guide gives a worked bundle example: three 5.2 mm wires, 11.18 mm, and 13.38 mm with a 1.1 mm overbraid. The bundle solver reproduces both numbers, and a test holds it there.
  • Clean context. Each call runs with no tools, no MCP servers, no project files or memory, in an empty directory, under a fixed system prompt. Verified by asking the model to recite everything it was given.
  • The model that answered is recorded. The provider reports which model served each call; any mismatch with the one requested is shown, not hidden.
  • Every reply is stored. The raw data behind each run is downloadable from the results page, so any score can be re-graded.

Three tiers. This is the first.

Tier 1 · Unaided

The model alone: no calculator, no code, no tools. The strictest reading of “AI can do it”, and what every result on this site measures today.

Tier 2 · With code

The same problems with a sandboxed interpreter, the way an engineer would actually use an agent. Expect the large instances to move most: arithmetic at scale is where an unaided model slips.

Tier 3 · Capital in the loop

An agent drives Capital through its API, and Capital’s own design rule checks are the judge. Capital also referees the reference solvers. This is what replacing Capital in production would mean.

Do we need Capital to run this?

Not for tier 1. The question is whether AI can do the engineering, and the engineering has right answers that a solver can compute. Keeping Capital out also stops the model from leaning on the tool it is being compared with.

Capital matters for the next two questions. First, are the reference solvers right? Today they rest on Capital’s documented methods, hand-worked cases and one reproduced documentation example. Running the same instances through Capital would make the tool itself the referee. Second, does a correct answer become a correct design? Capital is a system of record with libraries, variants, drawings, manufacturing outputs and a thread into Teamcenter. An answer in JSON is not a released harness. Tier 3 closes that gap.

Known limits

  • Capital’s documentation describes the Navy Table method but does not publish the full factor table. The factors here are supplied in the task; the three-wire factor matches Capital’s worked example.
  • The wire table and fuse ratings are representative values given in the task, not a cited standard.
  • Generated data is clean. Real projects have inconsistent libraries, missing attributes and conflicting rules, which this does not test.
  • Every model runs at its provider’s default reasoning settings, under the same system prompt, on the same seeds. Claude models run through the Claude Code CLI on a workstation subscription; OpenAI and Google models through their APIs. A model added later answers exactly the instances the others did.

Reproduce it

npm test                                  # graders, solvers, controls
npm run sample -- bundles L 1             # print any instance
npm run eval -- --models claude-opus-5-5 --sizes S,M,L --reps 3
npm run eval -- --models openai:<model>   # needs OPENAI_API_KEY
npm run eval -- --resume <runId>          # redo only could-not-run trials

See the results