Glossary
June 7, 2025

Apple Researchers Argue the AI That 'Reasons' May Not Really Be Reasoning

In a paper titled 'The Illusion of Thinking,' Apple scientists reported that the step-by-step reasoning models collapse on puzzles beyond a certain difficulty — and, oddly, seem to give up and think less as problems get harder. It became the summer's sharpest note of skepticism.

Apple, which has largely stayed out of the AI race, published a research paper this week that landed hard on it. Titled "The Illusion of Thinking," it tested the reasoning models — OpenAI's o-series, DeepSeek's R1, and others that work through problems step by step — on classic logic puzzles like the Tower of Hanoi, at increasing levels of difficulty. The finding: the models did well on easy and medium puzzles, and then, past a certain complexity, collapsed completely — not degrading gradually but failing entirely. Stranger still, as the problems got harder, the models appeared to "think" less, using fewer reasoning steps, as if giving up.

The paper's argument is that what looks like reasoning may be a sophisticated form of pattern-matching that breaks down when a problem is genuinely novel rather than similar to something seen in training. In other words, the models may not be reasoning so much as recognizing.

It set off a loud debate. Critics of the paper argued the puzzles hit the models' output limits rather than their thinking — that a model failing to write out a thousand-step solution isn't the same as failing to understand it — and some showed the models could describe the correct method even when they couldn't grind out every step. Apple's defenders said the collapse was too sharp and too consistent to explain away.

The dispute matters beyond the technical details, because the whole industry's recent progress — and much of its valuation — rests on the bet that these reasoning models are a path toward genuine problem-solving. Apple, notably behind in AI and with reason to question a race it is losing, was the one to publish the doubt. Whether that makes the doubt convenient or brave, it named the question the field would rather not dwell on: is this thinking, or a very good imitation of it?

Follow the timeline
Get an email when new entries are added.
© 2026 Sugarpine