Scientific discovery is not a matter of scaling compute, but of cognitive architecture. In a position paper titled "LLMs can't jump," Tom Zahavy from Google DeepMind argues that large language models are fundamentally incapable of sparking scientific revolutions because they lack the mechanism to create truly new concepts.
The Bottleneck of Manipulative Abduction
Zahavy utilizes Charles Sanders Peirce's framework to distinguish between deduction, induction, and abduction. While AI excels at deduction (deriving conclusions from rules) and induction (pattern recognition), it fails at creative abduction.
There is a critical gap between ordinary abduction — such as matching symptoms to a known disease, which LLMs do well — and manipulative abduction. The latter involves inventing a cause for which no linguistic template yet exists. According to the paper, systems like GPT-5 or Gemini could derive general relativity if given Einstein's axioms, but they cannot formulate those axioms from scratch.
The Optimization Trap and Data Limits
AI learns by minimizing the error between prediction and reality. However, paradigm shifts often occur when there is no obvious error signal to optimize. Zahavy notes that during Einstein's time, Newtonian physics was highly accurate. An optimization-driven AI would not have overthrown the existing laws of physics to explain Mercury's orbit; it would have simply hypothesized a hidden planet, mirroring the flawed logic of astronomers at the time.
Embodied Simulation vs. Symbol Shuffling
True breakthroughs often stem from physical intuition rather than calculation. Einstein's "happiest thought" regarding a falling observer and Archimedes' buoyancy principle emerged from embodied simulations of physical sensations. LLMs, akin to John Searle's "Chinese Room," shuffle symbols without accessing the physical experience that gives those symbols meaning. Even advanced tools like Sakana's AI Scientist or AlphaEvolve are limited to recombining existing concepts.
The Shift Toward World Models and Physical AGI
The path forward lies in World Models. Unlike video generators like Veo, which predict the next frame based on probability, action-controllable models like Genie allow agents to intervene in simulations and run counterfactual experiments.
This transition toward simulating physical reality aligns with the pursuit of Physical AGI, a direction already explored by Google DeepMind through Gemini Robotics 2, where AI coordinates humanoid bodies in real-world environments. Only through a synthetic lab that enables the discovery of new axioms can AI finally make the cognitive leap required for genuine scientific invention.

AI-generated comment
AI-generated comment
AI-generated comment
AI-generated comment