The Anatomy of Gary MarcusDissecting the Mind of a Failed AI Prophet
August 2026

The fake image
I was wrong.
The image I posted in reply to Gary Marcus was fake. I saw it in someone else's post while scrolling X, thought it was funny, and reposted it as light banter. I did not find it on Reddit, trace its source, or check whether it was authentic. It showed a fake account impersonating Marcus. He asked for an example of something he had got wrong; I answered with something that was not his. He checked it. I had not.
That failure to check is on me. It was X, not Wikipedia: I was making a joke in a feed, not presenting a researched citation. That context explains the casualness, but it does not make the attribution true. So, yeah: on the narrow question captured in that screenshot, Gary Marcus was correct and I was not.
What interested me was what happened next. Marcus treated the discovery of one fake joke as though it had settled the larger argument about his record. He had caught an error, therefore the critic was unserious; therefore the criticism could be dismissed; therefore Gary Marcus remained, somehow, correct.
For a man who has made a public career out of finding errors in artificial intelligence, it was an elegant little performance. It was also the reason I decided to stop replying with images and examine the predictions themselves.
Marcus's correction deserves to remain at the beginning because it establishes the standard I want to use. A claim should be capable of losing. When the evidence goes against it, the claim should be withdrawn or changed openly. You do not get to replace the evidence, move the deadline, or redefine the subject and still describe yourself as having been right all along.
My post loses under that standard. It attributed words to Marcus that he did not write. The proper response is not to explain that the fake post sounded like him, or that he has said similar things, or that the joke captured some deeper truth. The proper response is: it was fake, and I was wrong to use it.
Now apply the same standard to Marcus.
A definition of wrong
To me, Gary Marcus has become the definition of being wrong in a very specific way. He is often right about the defect directly in front of him and wrong about what that defect means. A model hallucinates, so the architecture cannot reason. An image generator binds the wrong color to the wrong object, so compositionality is the wall. A system requires search and verification to discover something, so the language model did not meaningfully contribute to invention.
The observation is usually real. The prophecy built on top of it is where the trouble begins.
His recurring method looks like this: find a task the current model fails; treat the failure as evidence of a missing kind of cognition; infer that the prevailing architecture cannot acquire it; then, when a later system performs the task, select a harder failure and preserve the architectural conclusion.
This is why arguing with Marcus can feel like trying to reach the edge of the horizon. Every step produces movement, but never arrival. The wall remains exactly where the newest model stops.
The prediction ledger
Before discussing intelligence, understanding, or consciousness, it is useful to separate claims that can lose from claims protected by their vagueness. Marcus has made both. His timed economic forecast can be scored. His phrase “the wall” is harder because its location changes. His predictions that hallucinations will persist are correct but set an extreme standard: one remaining error is enough to preserve them.
| Date | Claim | What followed | Verdict |
|---|---|---|---|
| 2020 | GPT-2’s failures showed the need for a different approach. | Later GPT models solved most of the showcased failures. | Wrong inference |
| 2022 | Compositionality is the wall for image generators. | Later systems passed the prompt suite used to test the claim. | Wall crossed |
| 2023 | GPT systems regurgitate ideas; they do not invent them. | LLM-based systems produced novel, verifiable mathematics. | Wrong in practice |
| 2024 | The generative-AI collapse could arrive in days or weeks. | Investment rose in 2024 and exceeded $170 billion in 2025. | Wrong forecast |
| 2022–25 | Hallucinations and agent unreliability would persist. | Both remained serious limitations. | Right |
The final row matters. Marcus is not wrong about everything. He correctly predicted that GPT-4 would remain unreliable, hallucinate, and fall short of AGI. His warnings about agents outside bounded environments were sensible. If I omitted those calls, I would be doing the same thing I am accusing him of doing: arranging the evidence so the conclusion cannot be disturbed.
But notice the asymmetry. “The model will still fail somewhere” survives almost any amount of progress. “This failure reveals the limit of the architecture” does not. Marcus is strongest when he predicts residual error and weakest when he turns residual error into a ceiling.
I graphed the pattern. The index scores the claims listed, uses the scale printed underneath, and accumulates those scores over time. The point in 2025 remains flat because Marcus's no-AGI forecast was correct. Even a wrongness index should be capable of admitting contrary data.
GPT-2’s failures diagnose the architectural limit
3/5Compositionality is the wall
4/5GPT systems regurgitate; they do not invent
4/5The generative-AI collapse is days or weeks away
5/5No AGI in 2025; agents remain unreliable
0/5Scale: 0 = correct; 1 = broadly right; 2 = too vague to score cleanly; 3 = wrong inference; 4 = contradicted by later capability; 5 = explicit, timed forecast that failed.
The self-sealing critic
A normal prediction has an exterior. Reality can approach it from outside and break it. Marcus's central claims increasingly have no exterior. If a model fails, the failure confirms the wall. If it succeeds, the task was too small, the benchmark was contaminated, the system used a scaffold, or the success was not evidence of true understanding. Both outcomes return the same verdict.
This is more than ordinary goalpost-moving. It is a self-sealing method. The critic defines himself as the person willing to notice what everyone else has missed. Contrary evidence therefore arrives already degraded: it is hype, a trick, a demo, a narrow win, or another example of the crowd misunderstanding intelligence. The possibility that the critic misunderstood the trajectory is the one hypothesis that receives no serious testing.
Catching my fake screenshot fits this structure perfectly. Marcus found an error, and the error was real. But the pleasure of correction expanded beyond its jurisdiction. One careless X reply became a miniature proof of his larger position: critics do not have quotations, people run away when challenged, and Gary Marcus is the rare adult still checking the facts.
The correction was right. The implied coronation was not. Being right about the authenticity of an image says nothing about whether compositionality was a wall, whether GPT systems can invent, or whether a financial collapse was days away. Those claims still have to face their own evidence.
This is the first organ in the anatomy of wrongness: a correction mechanism that points outward but not inward. It can identify an error in anyone else with precision. When the error belongs to the critic, it becomes context, nuance, a misunderstood definition, or proof that the original warning was directionally wise.
The test that outlived GPT-2
In 2020 Marcus published “GPT-2 and the Nature of Intelligence”. GPT-2 produced absurd continuations involving arithmetic, locations, poison, physical causality, and ordinary knowledge. Marcus was right about the outputs. GPT-2 failed the tests.
He then treated those failures as evidence that the empirical, symbol-free route to language had failed. The knowledge inside the model, he argued, was superficial and unreliable. It was time to invest in a different approach.
Then the same lineage improved. In 2022 Scott Alexander reran nine of Marcus's GPT-2 examples on a later GPT-3. GPT-2 had failed all nine; GPT-3 answered between five and seven correctly depending on grading. On another six-prompt suite that an earlier GPT-3 had failed, the later model scored 4.5 out of 6. The prompts and outputs remain available in Alexander's comparison.
Marcus and Ernest Davis then published new GPT-3 failures. After GPT-4 arrived, an evaluator ran fifteen of those later examples and reported GPT-4 succeeding on all fifteen. These were informal tests, not a controlled benchmark, and they do not prove human understanding. But the original failures were informal too. If they were good enough to diagnose architectural incapacity, their disappearance must count against the diagnosis.
| Model | Suite | Reported result |
|---|---|---|
| GPT-2 | 9 selected prompts | 0/9 |
| Later GPT-3 | Same 9 prompts | 5–7/9 |
| Earlier GPT-3 | New 6-prompt suite | 0/6 |
| Later GPT-3 | Same 6 prompts | 4.5/6 |
| GPT-4 | 15 later prompts | 15/15 |
The careful conclusion from GPT-2 was that GPT-2 failed. Marcus wanted the failure to reveal what the entire family could never become. The family refused to cooperate.
The moving wall
In April 2022, discussing image generators that could not reliably bind properties to objects, Marcus wrote: “Compositionality is the wall.”
The systems really were bad at it. Ask for a red sphere on a blue cube beside a yellow pyramid and the image might contain all the right nouns with every relationship wrong. Scott Alexander believed the defect was temporary. A commenter named Vitor believed solving it required something close to genuine semantic understanding. They created a wager with a specified prompt suite and deadline.
Marcus was not a party to that wager. This distinction matters because getting it wrong would hand him a legitimate factual objection. By 2025 the specified prompts worked, and Marcus acknowledged that Alexander had won. He then argued that passing those prompts did not establish complete mastery of compositionality.
Of course it did not. But if failure on the prompts revealed a wall, success on the prompts cannot become irrelevant merely because harder failures remain. The claim silently changes from “this inability reveals the architectural limit” to “the limit is whatever the newest system still cannot do.”
A wall that moves whenever something crosses it is not a wall. It is a person walking backward while insisting he has remained still.
Invention after regurgitation
Marcus made a cleaner claim in his 2023 essay “GPT-5 and irrational exuberance”: “GPT's regurgitate ideas; they don't invent them.” Mathematics now makes that sentence difficult to preserve through rhetoric. A new construction that improves the state of the art, survives expert checking, or compiles into a formal proof is not merely a fluent paraphrase.
In 2023, DeepMind's FunSearch paired a language model with an evaluator and evolutionary search. The system found new constructions for the cap-set problem and new bin-packing heuristics. Marcus replied that the evaluator and search scaffold did crucial work. They did. The language model was part of a system rather than a solitary chatbot answering once.
In 2025, AlphaEvolve used Gemini models inside an evolutionary coding agent and found a 48-multiplication algorithm for 4×4 complex matrix multiplication, improving a result that had stood for 56 years. Again there were evaluators, iteration, and tools.
In 2026, an OpenAI reasoning model generated a new infinite family of constructions that disproved a longstanding conjecture around the planar unit-distance problem. External mathematicians checked it; a human-verified account and OpenAI's verification history are public. The result contradicted the prevailing conjecture and used a construction absent from the literature.
On August 1, 2026, OpenAI published ten further results from Astra across mathematics and theoretical computer science, accompanied by machine-checkable Lean formalizations and a 249-page collection. The release is new, so ordinary peer review is incomplete. That is a reason for precision, not dismissal.
Humans invent with notebooks, compilers, proof assistants, instruments, collaborators, search, and correction. If an LLM-based system stops counting as inventive the moment it is allowed those things, “does not invent” has become a protected definition. It can never lose because any successful system is declared insufficiently pure after the fact.
The collapse that did not arrive
Economic forecasts are useful because dates do not move as gracefully as philosophical definitions. On March 31, 2024, Marcus wrote that if no transformative GPT-5 arrived, the generative-AI bubble might begin popping within roughly a year. On August 3 he accelerated the forecast. The collapse, he wrote, could arrive in “days or weeks” and likely before the end of 2024. Investors might stop funding the sector at previous rates; billion-dollar companies might fold or be stripped for parts.
The opposite happened. Stanford's 2025 AI Index reported $33.9 billion in global private generative-AI investment during 2024, up 18.7 percent over 2023. OpenAI raised $6.6 billion at a $157 billion valuation in October, as Reuters reported. Nvidia's fiscal 2025 revenue rose 114 percent to $130.5 billion, according to its results.
Then the 2026 AI Index reported $170.9 billion in private investment in generative-AI companies during 2025—more than five times the 2024 total.
Sources: Stanford AI Index 2025 and 2026. The 2022 and 2023 values are derived from Stanford's reported comparisons.
There may still be an AI bubble. Revenue is not profit. Capital spending can outrun demand, price competition can destroy margins, and individual firms can fail while the technology succeeds. But a future correction cannot travel backward through time and rescue a forecast of collapse within days or weeks in August 2024. That prediction was wrong.
Why the method fails
A failure case is evidence about the system that failed it. It is not an impossibility proof about every future system in the architectural family. If the current model fails an input, we have established only that the input belongs to the current model's failure set.
To make the stronger inference, Marcus needs a theory connecting the observed error to an invariant of the architecture. Instead, he usually supplies an intuition: no symbols, no world model, no grounded understanding. When descendants of the architecture later solve the supposedly diagnostic examples, confidence in the diagnosis should fall.
Instead the failure set is sampled again at the new frontier. GPT-2 fails one suite, GPT-3 passes much of it, and a new suite appears. GPT-4 passes that suite, and the remaining failures become evidence of the same unchanged incapacity. This method can make every model look permanently broken even while the set of things it can do expands dramatically.
Marcus is right that the tail matters. If an agent must complete one hundred dependent steps, small per-step error rates compound. Assuming independent steps, total success is the per-step reliability raised to the number of steps.
| Per-step reliability | End-to-end reliability |
|---|---|
| 95% | 0.59% |
| 99% | 36.6% |
| 99.5% | 60.6% |
| 99.9% | 90.5% |
This is why a model can look brilliant in a demo and fail as an autonomous worker. It is also why apparently small improvements matter. Moving from 99 percent to 99.9 percent per-step reliability changes a hundred-step task from succeeding 36.6 percent of the time to succeeding 90.5 percent of the time. The nonzero failure rate does not prove the family has stopped improving.
The pure LLM escape hatch
Marcus can answer much of this by saying modern systems are not pure LLMs. They use retrieval, reinforcement learning, search, memory, verifiers, code execution, and agent scaffolds. Progress therefore validates his call for hybrid systems rather than contradicting his skepticism.
There is a legitimate technical distinction here. Engineers should know which capabilities are learned in weights, which emerge through inference-time search, and which are supplied by external software. But that careful distinction cannot rescue broad public claims that deep learning hit a wall or that GPT systems cannot contribute to invention.
No economically relevant intelligence is pure. A mathematician with Lean is still doing mathematics. A programmer with a compiler is still programming. A scientist with an instrument is still discovering. The useful unit is the system that performs the work at a given cost and reliability, not a deliberately disabled base model forbidden from receiving feedback.
If “pure LLM” means a frozen next-token predictor answering once without tools, memory, search, or checking, Marcus may be right that it will not become a reliable general agent. He will also have won an argument about a product almost nobody serious intends to deploy.
The ego in the model
I do not know Gary Marcus privately, and I cannot inspect his motives. This is not a clinical diagnosis of a stranger from his posts. It is a reading of the public character he has built and the argumentative incentives that character creates.
Marcus is not merely an AI skeptic. He is the skeptic who warned you. That identity depends on being early, outnumbered, and eventually vindicated. Every confident product launch strengthens the role. Every embarrassing hallucination supplies another scene in which the crowd celebrates and Marcus quietly points to the crack in the stage.
There is power in that role because AI discourse is genuinely full of nonsense. Companies overstate demos. Investors repeat technical claims they cannot evaluate. Benchmarks are marketed as intelligence. A person willing to say “this does not work” can be useful. But usefulness turns into ego when the role of corrector becomes more important than whether the correction still describes reality.
The public Marcus appears to need each model's failure to mean more than the failure itself. It must confirm that the field's conceptual foundations are mistaken and that he has understood the missing piece. A temporary bug is too small. A reliability problem is too ordinary. The defect must reveal the anatomy of intelligence—and, by implication, the unusual foresight of the person diagnosing it.
This helps explain why updating is expensive. If Marcus says that a prompt suite he called diagnostic was simply solved by continued development, he has not only lost a prediction. He has weakened the public identity built around seeing the permanent boundary before the rest of us. It is cheaper to move the boundary and call everyone else confused.
Ego does not necessarily produce a lie. More often it chooses the level of abstraction at which the speaker can remain correct. “This model failed” becomes “this paradigm cannot reason.” When the paradigm reasons well enough to solve the example, “reasoning” becomes something deeper that the example was never capable of testing. The word stays. Its obligations disappear.
What Marcus gets right
Marcus's 2022 GPT-4 predictions said the model would remain hard to control, hallucinate, make unpredictable errors, fail as reliable medical advice, and fall short of general intelligence. Broadly, he was right. His prediction that AGI would not arrive in 2025 was right. His warning that agents would remain unreliable outside constrained settings was defensible.
He is also right that benchmark scores are not deployment. A model that succeeds nine times out of ten may be useless in medicine, law, or a long autonomous workflow. He is right to demand external verification, clearer reporting, and less theatrical demos. The mathematical systems that now contradict his claim about invention depend on exactly those things: evaluators, execution, formal proof, and human review.
These correct calls support a useful thesis: reliability engineering matters, and intelligence is larger than a base model. They do not support the thesis that neural language models are irrelevant to future intelligence, that scaling failed, or that every capability absent from today's model requires abandoning its lineage.
An anatomy of wrongness
The anatomy is now visible. The eye finds a genuine defect. The hand lifts that defect out of its limited context. The mouth names it a wall. The ego turns the wall into evidence that the speaker saw what the crowd could not. Then the feet move backward when the technology reaches the place where the wall was supposed to stand.
Finally, memory edits the map. The old boundary becomes a mere example, never the real claim. The successful system becomes a hybrid, never the architecture under discussion. The prediction was directionally right, even when events travelled in the opposite direction. Every organ works to preserve the same conclusion: Marcus warned us, Marcus remains correct, and the decisive failure is still one model away.
That, to me, is the definition of being wrong. It is not making a false claim; everyone does that. It is building an intellectual body whose immune system attacks correction. The wrong claim does not get removed. It is absorbed, renamed, and returned as wisdom.
This is why the accurate parts of Marcus's criticism do not rescue the pattern. In some ways they make it worse. The hallucinations are real. Reliability remains difficult. Demos are abused. Those correct observations provide permanent cover for the incorrect leap from “this system fails” to “this is where the technology ends.” He can always point to the first sentence when the second one collapses.
A claim must be allowed to die
Marcus represents something larger than Marcus. AI discussion is full of claims designed to win attention rather than encounter reality. The optimist announces AGI without defining it. The skeptic announces a wall without locating it. One sells inevitability; the other sells impossibility. Both can spend years explaining why the latest evidence did not really test what they meant.
We should demand a less theatrical form of disagreement. If there is a wall, give it a location. Name the capability, the conditions, the date, and the result that would show the wall had been crossed. If a bubble will collapse, define collapse and keep the deadline. If a system cannot invent, say whether a novel, externally verified mathematical result would change your mind before the result arrives.
Then keep a ledger. Do not count only the warnings that survived. Do not turn every loss into a more sophisticated version of the original win. The ability to update is not a humiliation added to intelligence. It is part of intelligence.
My unverified X reply lasted until someone checked it. It failed, so I am saying it failed. Marcus's prophecies last because every check is turned into another prophecy. Yeah, he was right about my screenshot. I was wrong. That is not the end of the argument. It is the standard the rest of the argument asks him to meet.
Gary Marcus does not need to stop being skeptical. He needs to permit his skepticism to be wrong. Until then, the moving wall is not merely a mistake he keeps making. It is the product he keeps selling.