Frontier Agents Were Asked to Design Real Parts. None Passed.
In May 2026, a group of researchers set up an experiment I have been waiting for someone to run properly.
They took the best coding agents publicly available — Codex running GPT-5.5, and Claude Code running Opus-4.7 — and gave them something closer to an actual engineering task than the usual demonstration. Not “generate a shape.” A free-form engineering brief, from which the agent had to produce a fully assembled multi-part STEP file, which was then validated by finite element analysis.
Their finding, stated plainly in the abstract:
The agents “do not produce a single strict-passing artifact in the main first-attempt sweep, with the best configuration meeting only about 20% of typed requirements on average.”
Zero strict passes. Around one requirement in five.
I think that is the single most useful number in AI-for-engineering at the moment, and not because it is discouraging. Because it is measured, on a task that resembles the job, by people who published the failure.
Why the Task Matters More Than the Score
Almost every impressive CAD-generation result you have seen was graded on the wrong thing.
The convention has been to score generated geometry by its similarity to a reference model — how closely does this shape resemble the shape a human made? The paper's authors are blunt about the problem: systems assessed that way produce designs that are visually plausible while carrying no guarantee of functional validity or load-bearing capacity.
A part can score beautifully on geometric similarity and fail under load. Similarity is not a specification.
So the interesting move here is not the score. It is the switch from grading appearance to grading admissibility — does this thing pass the checks a real part has to pass. The moment you do that honestly, the number falls off a cliff.
That fall is not a failure of the research. It is the first accurate reading we have.
What Is Actually Failing
Read the sentence again and notice what it says the agents fell short of. Not geometry. Not compilable code. Typed requirements.
That distinction is everything.
The agents can produce parametric CAD. They can write it as code, export a STEP file, and iterate on it. That part largely works. What they cannot reliably do is satisfy a full set of stated engineering constraints simultaneously — the mass target and the stiffness and the material and the fixing interface and the manufacturability and the clearance envelope — and then demonstrate that they have.
Anyone who has run a design review will recognise this immediately, because it is also the most common failure mode in human engineering. The geometry is rarely the hard part. Holding twenty constraints true at once is the hard part.
The Architecture That Is Emerging
A second paper from the same month, accepted at IJCAI-ECAI 2026, takes what I think is the correct structural position on this.
Berger and colleagues argue that rather than trying to get a model to absorb physics implicitly from data, you should embed validated knowledge-based engineering tools directly into the agent's decision loop. Design becomes a closed-loop sequential decision process, and every step is checked by explicit physical verification rather than by the model's own confidence.
In other words: stop asking whether the model understands mechanics. Assume it does not. Put a solver in the loop and let the solver be the authority.
This is not a small design choice. It is the difference between a system whose correctness depends on a model's judgement and one whose correctness depends on tools you already trust and already validate today.
The Part Practitioners Should Take From This
There is a pattern running through all of this work that is worth stating directly, because it contradicts how most people are currently building with agents.
The agent is not the pipeline. The agent is the judgement inside a pipeline you control.
In the FEA-feedback paper the split is explicit: the language model makes design decisions, while a deterministic controller handles execution, validation and feedback routing. The model decides what to try. Ordinary software decides what happens next, runs the checks, and routes the results back.
That is the shape that works, and it is a familiar one. It is what a workflow orchestrator like n8n is for — deterministic steps, explicit branches, retries, a defined stopping condition, and a log you can audit afterwards. The intelligence sits at the nodes where a judgement is genuinely required. The control flow does not.
The common failure I see in early agent projects is the opposite arrangement: hand a capable model a pile of tools and let it decide the whole sequence. It demos well. It is unreproducible, unauditable, and impossible to certify — which in an engineering context is the same as unusable.
A Bracket Pipeline That Could Be Built Today
Let me make this concrete. What follows is hypothetical — I am not describing a system in production anywhere. But every component exists, and nothing in it requires a capability beyond what the papers above demonstrate.
Take an ordinary steel bracket. Something that carries a known load between two mounting faces, in a defined envelope.
One. Encode the brief as a machine-checkable spec. Load cases, safety factor, material, mass target, fixing positions, envelope, minimum wall, draft angles, permitted processes. Not prose — fields with numbers and tolerances.
Two. Generate. The agent writes parametric CAD as code and exports a STEP file. This is the step everyone is impressed by and it is the least of your problems.
Three. Check cheaply first. Does it compile, does it stay inside the envelope, does it hit the mounting interfaces, is the mass in range. Reject in seconds, before anything expensive runs.
Four. Simulate. Mesh, apply the load cases, solve. Parse peak stress, deflection, safety margin back into structured numbers.
Five. Decide. Pass, revise, or abandon. This is the only step in the whole loop where a model's judgement earns its place — reading a solver result and deciding what to change next.
Six. Stop. A hard iteration budget, and every attempt logged with its inputs, geometry and results.
Steps one, three, four and six are ordinary automation. Step two is a solved-enough problem. Step five is the agent. Notice the ratio — the intelligence is one node in six, and the other five are what make the output trustworthy.
And notice which step is genuinely hard. It is step one.
The Real Blocker Is Not the Model
Everything downstream depends on requirements being written in a form a machine can evaluate. In most engineering organisations they are not.
They live in specification PDFs written for humans. In standards referenced by number. In the design guideline nobody has updated since 2019. In a validation engineer's memory of why that radius has to be 3mm on this family of parts.
An agent cannot satisfy a requirement it cannot read, and it cannot demonstrate compliance with a rule that was never written down. When these systems fail, the tempting conclusion is that the model was not clever enough. Very often the constraint is not on the model's side of the table.
An agent cannot check a requirement that exists only in a PDF. Most of your requirements exist only in a PDF.
Which makes the preparatory work unglamorous and quite specific: turning a body of institutional requirements into structured, checkable form. That is not an AI project. It is a knowledge-engineering project, and organisations that start it now will be able to use these tools when they mature. Organisations that wait will find the tools ready and their own requirements unusable.
What the Next Ten Years Look Like
Forecasting a decade of anything is mostly a way of being wrong in public, so let me separate what I would bet on from what I would not.
Near-certain, this decade. Validation cost keeps collapsing. Surrogate models already predict solver outcomes accurately enough to triage candidates, and the authors of one such tool have observed that as validation gets cheap, generating geometry becomes the new bottleneck — which is precisely why the agent research above exists. Expect closed-loop generate-check-revise to become normal for well-bounded components: brackets, housings, mounts. The parts where the load case is clear and the consequences of being wrong are contained.
Likely, but slower than the demos suggest. Agent-assisted design moving from single parts to assemblies and interfaces. That 20% figure will improve substantially. What will not improve on its own is the requirements problem, because that is organisational rather than technical, and no model release fixes it for you.
Genuinely uncertain. Certification. Type approval assumes a traceable chain of human engineering judgement. A design produced by a stochastic process that cannot fully explain itself does not slot into that framework, and the regulatory answer is years away. My expectation is that agents will design a great deal and sign off nothing — the human signature remains, and it remains meaningful.
What I do not expect at all. Novel vehicle architecture from agents. These systems interpolate within what they have seen. Genuinely new architecture is an extrapolation problem, and it is also where the commercial value is highest. That gap is not obviously closing.
What This Changes About the Job
If most of this lands, the centre of gravity in engineering work moves.
Less time producing candidate geometry, because that becomes cheap. More time on the two ends of the process — stating the problem precisely enough that it can be checked, and judging whether the answer that comes back deserves to be trusted.
Both of those are the parts of engineering that were always the actual work. They were simply hidden underneath the labour of drawing and meshing and rerunning.
The engineers who do well here will not be the ones with the best prompts. They will be the ones who can say exactly what “good” means for this part, in terms a machine can verify — and who can still tell when a confident answer is wrong.