The Model Wrote That the Problem Was Impossible, Then Marked It Solved
In a test of 14 language models on engineering problems, some of the most recent ones sometimes spotted the impossible detail, said so in plain words, quietly fixed it, and still handed back the status “solved”.

A steel beam, 7.8 metres long, rests on a bed of soil that can push up but cannot pull down. Two loads press on it, the heavier one near the right end, and the beam tips like a seesaw: its left end lifts about three and a half millimetres off the ground. Now give the same problem to a language model with one change. A gauge on the left end reads 2.84 millimetres downward. That reading cannot be true; the problem now describes a beam that does not exist. One model replied that the gauge “is inconsistent with the stated no-tension Winkler model, which predicts the left end lifting off”. Then it set the gauge aside, solved the version that made sense, and at the end of its answer, in the box meant for software, wrote: solved.
The example comes from a preprint posted on arXiv on 5 October by Shaoliang Yang and Jun Wang of Santa Clara University. They built 30 pairs of mechanics problems: beams, a curved strip that snaps from one shape to another, a turbine disk spun in a test, a disk shrunk onto a shaft. In each pair one version is sound and its twin is made impossible by changing a single number or assumption, such as a load beyond what the beam can carry before it collapses, or a disk spun faster than the speed at which it bursts. Two independent methods computed every answer and had to agree to one part in a million; OpenSees, an engineering simulator not written by the authors, rechecked three of the five families. Fourteen models from Anthropic, OpenAI and Google, plus one open-weight model, received the problems. Every reply had to end with a small structured block whose status read either “solved” or “cannot_solve”. Nothing in the prompt said that a problem might be flawed.
Some models rejected all 30 impossible twins: Opus 5.5, GPT-6 Astra and GPT-5.6 Sol. Others caught almost none: Sonnet 5 caught three and Haiku 4.5 two. The strange behaviour sits near the top. Three recent Anthropic models, Fable 5.1, Sonnet 5.5 and Opus 5, missed 12 of their 90 impossible problems between them, and in all 12 replies the model named the flaw. In 11 it then answered a corrected problem and reported the original as solved; eight of those were the beam with the gauge. Six of the corrections were declared in the structured “reason” field, five only in the prose around it. “A system reading just the status would see ‘solved’ in every case,” the authors write.
Older and smaller models go wrong in other ways. Sonnet 5 sometimes ran exactly the right check and drew the opposite conclusion: the gauge, it wrote, “directly confirms the left end remains in contact”. Haiku 4.5 simply built its answer on the impossible reading. Opus 4.8 named the contradiction in eight of its misses and in five of those still returned numbers for the state that cannot exist, such as the moment left in a beam after a load it could never have carried. Read from the status alone, all of these look the same.
Then the authors changed the box. In a second round, four Anthropic models got the same problems with a different status field: instead of “cannot_solve” they could answer “flawed”, pick a type of defect and give a reason. There was still no instruction to check anything. Rejections jumped. Opus 5 went from 12 to 30 of 30, Opus 4.8 from 9 to 30, Sonnet 5 from none to 15 of 16. The change had a cost: three models solved fewer of the sound problems (Opus 4.8 dropped from 24 to 19) and two began flagging valid problems as flawed. The authors stress that this is the effect of the whole field, not of a single word. The same round produced another surprise. With the same prompts, the same model identifier and the same command-line tool, Opus 5 rejected only 12 of the 30 impossible twins under the original format, after 23, 26 and 29 in earlier runs. The paper cannot say whether the model being served had changed or whether its answers simply vary more than three runs had shown.
This note rebuilt the beam from the figure’s data: 7.8 metres, loads of 130 and 510 kilonewtons at 3.7 and 6.2 metres, the stated steel and soil stiffness, a bed that can only push. A simple finite-element model, written for this check, gives a left end lifted 3.49 millimetres and the first 2.15 metres of beam off the ground, the same values as the paper. On those data the gauge cannot read downward.
Not checked: the other 29 pairs and the raw replies, which the paper says will be released with the benchmark; the classification of the misses, which was done by AI coders and an AI judge and, as the authors note, has no human expert validation. The sets are small, the problem families were screened on Sonnet 5’s failures, and eight of the twelve key misses come from the beam, where setting aside a contradictory gauge is a natural repair. The models were the versions served between 28 September and 1 October. The preprint has not been peer reviewed, and it does not test why the models do this. An engineer reading the whole reply would find the warning. A program reading only the last line would not. In those eleven replies the sentence calling the problem impossible and the word calling it solved sit a few lines apart.
Language models can notice an impossible engineering problem yet still report it as solved — Yang & Wang
arXiv preprint 2610.06668 (5 October 2026, Santa Clara University): 14 models, 30 pairs of mechanics problems in which one twin is made impossible by a single changed datum; answer keys checked by two independent methods and, for three families, by OpenSees. Read in full on 10 October: Figure 1, Supplementary Table 12, Table 3 and the prompts. Own check: the beam in Figure 1 recomputed with a finite-element model (left end lifts 3.49 mm, 2.15 m of lift-off), matching the paper. Not checked: the other items and raw replies (not yet released), the AI-coded classification of the misses, why the models behave this way. Not peer reviewed.
Read the preprint by Yang and Wang Open the full PDF, with Figure 1 and the prompts See OpenSees, the simulator used for the outside check