Coding Is All You Need? Why We Need a World Model!
Last week a result went around: GPT-6 Astra, put in charge of a robot arm, scored 95% where the previous best model scored 40%. The thread had two million views. The mood in the replies was that the last hard problem had fallen. If a model can write code, it can write the code that moves the arm. Embodiment is a downstream task. Coding is all you need.
GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
— Jay Chooi (@chooi_jeq) September 5, 2026
Read the report behind the thread. The 95% is a coarse task: pick up blocks, drop them in a bowl. The same report has a fine task, inserting a peg with millimetre tolerance. There the new model and the old one score the same: two out of twenty. The authors say the models "hit the same wall."
I do not think that wall is about scale, and I do not think it is about code. It is the oldest wall in the study of intelligence. The people who first ran into it were physiologists, a century ago, arguing about rats.
Here is the argument in one paragraph. Biological intelligence did not begin as language. It began as bodies sensing and moving, in cells and animals without a word to their name. What turned reflexes into minds was not a better reflex. It was an organ that builds a model of the world and can run that model with the eyes closed: to remember, to imagine, to plan. Language came last. If we want machines that act in the world, we have to build the organ the brain built first.
Intelligence before language
Start before the nervous system exists. A Paramecium is a single cell. No neuron, no synapse, no brain. It can be conditioned. In 2021 Gershman and colleagues went back through a century of single-cell learning experiments and found that the best of them hold up. Their conclusion: single cells do "a form of information processing that neuroscientists have traditionally attributed to networks of cells." The molecules that neurons would later use to learn were already there, waiting for a circuit.
A sponge has no neurons at all, and it too orchestrates behaviour using cellular mechanisms. Its genome harbours most of the molecular parts of a synapse, and the edge of its waste-discharge orifice has cilia that sense the direction of water flow. When contaminated water seeps in, the animal swells slowly, then constricts and squirts the muck back out through its pores. This looks like a sneeze that lasts half an hour.
Later, nervous systems arose in a variety of guises. Jellyfish have a neural net and neural ring. Comb jellies employ an unusual complement of chemical messengers, and Moroz and colleagues argued in 2014 that their nervous system evolved independently of other animals. By one estimate, brains evolved at least four separate times. All these nervous systems carry out the same basic functions: they sense the world, control movements, and enhance subsequent behaviour.
The clearest case is the sea slug Aplysia, Eric Kandel's animal. Touch its siphon and it pulls in its gill. Pair the touch with a shock to the tail, fifteen times, and the touch alone now produces a withdrawal four times as long. Kandel's group traced this down to one synapse. The touch lets calcium into the nerve terminal; the shock delivers serotonin; an enzyme that only fires hard when both arrive together turns the coincidence into a chemical signal. Repeat it and the signal reaches the nucleus and grows new synapses. Kandel put the moral plainly: there are "no fundamental functional or biochemical differences between the nerve cells and synapses of humans and those of a snail, a worm or a fly."

For six hundred million years, learning meant one thing: changing how a body responds to the world it touches. Language is a few hundred thousand years old. Intelligence was here long before anyone said anything.
Sherrington's reflex and Tolman's map
Charles Sherrington won the Nobel Prize for working out how reflexes fit together. His 1906 lectures are the founding text of the idea that behaviour is reflexes all the way up. The reflex arc, he wrote, is "the unit mechanism of the nervous system." And then:
"The 'simple reflexes' are ever combined into great unitary harmonies, actions which in their sequence one upon another constitute in their continuity what may be termed the 'behaviour' of the individual as a whole."
Sherrington, The Integrative Action of the Nervous System, 1906
Sherrington was careful. He called the simple reflex "a convenient, if not a probable, fiction" and left the purposes of whole behaviour to others. But psychology took the reflex and ran with it. For the stimulus-response school, learning was the strengthening of a connection between an input and an output. Edward Tolman, in 1948, described their picture of the rat's brain as "a complicated telephone switchboard."
Tolman never names Sherrington, but we know the radical philosophical departure of the intelligence thesis. Against the switchboard, Tolman offered this:
"The central office itself is far more like a map control room than it is like an old-fashioned telephone exchange… the incoming impulses are usually worked over and elaborated in the central control room into a tentative, cognitive-like map of the environment."
Tolman, "Cognitive Maps in Rats and Men," 1948
He had two kinds of evidence. The first is called latent learning. Rats ran a maze for ten days with no food at the end. They wandered. On the eleventh day food appeared. By the twelfth, these rats were making fewer errors than rats that had been fed from day one. They had learned the maze with nothing to gain from it. "They had been building up a 'map,'" Tolman wrote, "and could utilize the latter as soon as they were motivated to do so." A switchboard cannot do this. It has nothing to strengthen until a reward arrives.

The second is the novel path. Rats learned to reach food along one winding alley. Then the alley was blocked and eighteen new paths fanned out from the start. If the rats had learned a chain of movements, they should have picked the path nearest the old one. Instead the largest group, by far, picked the path that pointed straight at where the food used to be. They had not learned a route. They had learned where the food was.


An autoregressive transformer looks like a modern version of Sherrington's reflex system: context is the stimulus, the deep hierarchy is the integration, the next token is the response. All the weights are adjusted based on the quality of the response. Tolman's rats have some representation of the world, separate from any specific response. The transformer, by construction, has nothing of the kind.
Memory, World Model, Imagination
Tolman's map lives in the hippocampus. O'Keefe found place cells there in 1971; the Mosers found grid cells next door; they shared a Nobel Prize for "the brain's positioning system." But the clearest lesson about what the hippocampus is for comes from people who have lost it.
Clive Wearing was a conductor. In 1985 a virus destroyed his hippocampi. Since then he has lived in a present a few seconds wide. His diary is pages of the same line, "Now I am awake, first time," each one crossed out by the next. He still plays the piano. He still greets his wife every time as if she had been away for years.
For a long time amnesia was read as a recording failure: the tape does not record. In 2007 Demis Hassabis, then a PhD student in Eleanor Maguire's lab, asked a different question. Can a person without a hippocampus imagine?
He gave five patients with hippocampal damage, and ten matched controls, cues that had nothing to do with their past. Here are two of the cues, with a patient's answer and the matched control's, as the paper printed them.

The patients knew what a beach contains. Their general knowledge was fine; they named sand and sea and sky. What they could not do was put those things in one place. Hassabis scored the descriptions for richness and for whether the pieces sat in a single connected space. On both, the patients were far below the controls. The paper's conclusion is the sentence this essay turns on: the hippocampus works "by providing the spatial context or environmental setting into which details are bound."


Later that year Hassabis and Maguire argued that one process, scene construction, sits under remembering the past, imagining the future, navigating, daydreaming and dreaming. The same brain network lights up for all of them. It is the network the patients had lost. Memory, on this view, is not a recording. It is a construction from a model, with a tag on it that says "this one happened." As an equation: Memory = Imagination + Metacognition.
Scenes before words
If the world model came first, seeing the world should be fast and cheap. It is. In 2007 Fei-Fei Li and colleagues flashed photographs for between 27 and 500 milliseconds and asked people to write down what they saw. At the shortest exposures people report light, dark and shapes. A little longer and objects appear. A little longer and there is a scene: an office, a street, a restaurant. At about a tenth of a second, one eye fixation, most people can tell you where they are, who is there, and often what is going on.

Other labs go lower still. The gist of a scene, whether it is natural, how deep it is, whether you could walk into it, is available after a few dozen milliseconds. A named scene can be picked out of a stream of pictures shown for thirteen milliseconds each. An EEG signature of "animal or not" appears about 150 milliseconds after the picture. The first brain responses to a written word also arrive around 150 milliseconds. A word and a scene are different units, and a flash threshold is not a brain-wave latency. So the point is not that vision beats language by some factor. The point is what a tenth of a second buys. In one glance you get a layout, the things in it, how they relate, and an event: a father helping a boy, in a cubicle, with a laptop. Language delivers that one word at a time, over seconds, and delivers it as a pointer into a model the listener already has. A glance is the model, updating.
Two pathways to reasoning
In 1971 Shepard and Metzler showed people pairs of drawings of block figures and asked whether they were the same object, turned. The time to answer rose in a straight line with the angle between them, about a second at zero degrees, four to six seconds at 180. People were not comparing features. They were turning something over in their heads, at a steady sixty degrees a second.

Here is the puzzle. A few percent of people have no visual imagery at all; Adam Zeman named the condition aphantasia in 2015. Give them the rotation task and they can do it. In 2024 Kay, Keogh and Pearson found that people with aphantasia were slower, and more accurate, than controls. Asked how they did it, controls mostly said they turned the object. People with aphantasia leaned on reasoning about its structure without turning anything. Imagery, the authors conclude, "is not crucial for successful performance in classical mental rotation tasks."
Balaban and Ullman have built a theory on results like this. The mind, like a game engine, separates physics, where things are and how they move, from graphics, what that looks like. So if reasoning can solve the puzzle, why did evolution keep the picture?
Speed. People with aphantasia are slower on any question about how things look, in imagination and in perception alike, and not slower on questions about abstract words. The picture is the fast path.
Generality. A rule is written for one puzzle. A simulation runs on anything with a shape. That is why the same machinery serves rotating, navigating, catching and pouring, and why the costs of aphantasia show up everywhere else: almost no priming from imagined images, thinner detail in remembered and imagined episodes, fewer and poorer dreams.
Core knowledge and the tall glass that has more
What does a mind contain before it has words? Elizabeth Spelke's answer is a small set of core knowledge systems: for objects, for agents, for number, for the shape of the space around you. Each is present in infancy, shared with other animals, and has limits that identify it across ages and species.
Five-month-olds, in a 1985 experiment, watched a screen swing up and down like a drawbridge. A box was placed behind it. When the screen swung through the space where the box should have been, the babies stared. They expected a hidden object to still be there, and to be solid. That is an intuitive physics, installed before the first word.
Now watch a four-year-old do Piaget's conservation task.
The child agrees the two glasses hold the same. Pour one into a tall narrow glass and ask again, and the tall one has more. Most children do not reliably pass until about seven. In humans, a small innate core; then a world model built by acting on the world; then, last, the words to report on it. That is the opposite of the order we are building machines in.
Closed systems are solved but the world is open
Geoffrey Hinton, asked last year where AI would make its first real scientific discoveries, gave an answer that explains the whole opening of this essay:
"There's one area in which that's particularly easy, which is mathematics, because mathematics is a closed system. So you're going to get AIs that play mathematics. That is, they ask themselves, I wonder if I could prove this, I wonder if I could prove that. But because this is a closed system, they can just try things out and see if they can prove them… It's much like things like Go or chess. They're closed systems with rules, where they can generate their own training data."
Geoffrey Hinton, interviewed by Alok Jha of The Economist, GITEX Europe, Berlin, May 2025
Mathematics, Go, and code are closed systems. A model can try a million things and be told which ones were right without ever leaving the one-dimensional text world. That is why coding is what these models do best. It is also why a swarm of coordinating models produced a proof, formalized and checked in Lean for Navier–Stokes, a Millennium Prize problem.
But intelligence is not about solving closed systems, but exploration of the world, in its literal sense, the continent, the planet, the galaxy, and the universe. An alien intelligence as a brain in a vat would be very unlikely to share experiences, and thus, empathy as to us. But, an artificial intelligence with a world model might not, and by that, not so alien.
Sources
- Gershman, Balbi, Gallistel & Gunawardena (2021). Reconsidering the evidence for learning in single cells. eLife 10:e61907.
- Kornder et al. (2022). Sponges sneeze mucus to shed particle waste. Current Biology 32:3855. · Ludeman et al. (2014). BMC Evol. Biol. 14:3. · Leys (2015). J. Exp. Biol. 218:581.
- Moroz et al. (2014). The ctenophore genome and the evolutionary origins of neural systems. Nature 510:109. · Northcutt (2012). PNAS 109 (Suppl 1):10626.
- Carew, Walters & Kandel (1981). Science 211:501; J. Neurosci. 1:1426. · Hawkins, Abrams, Carew & Kandel (1983). Science 219:400. · Kandel (2000). Nobel Lecture; Science 294:1030 (2001).
- Sherrington (1906). The Integrative Action of the Nervous System. Quotations from Lectures I and VII.
- Tolman (1948). Cognitive maps in rats and men. Psychological Review 55:189. · Tolman & Honzik (1930). Univ. Calif. Publ. Psychol. 4:257. · Tolman, Ritchie & Kalish (1946). J. Exp. Psychol. 36:13.
- Hassabis, Kumaran, Vann & Maguire (2007). PNAS 104:1726. · Hassabis & Maguire (2007). Deconstructing episodic memory with construction. TiCS 11:299. · Maguire et al. (2000). PNAS 97:4398.
- Klein (2015). What memory is. WIREs Cogn. Sci. 6:1. · Aronowitz (2019). Memory is a modeling system. Mind & Language 34:483.
- Silver et al. (2016). Nature 529:484. · Banino et al. (2018). Vector-based navigation using grid-like representations in artificial agents. Nature 557:429.
- Fei-Fei, Iyer, Koch & Perona (2007). Journal of Vision 7(1):10. · Greene & Oliva (2009). Psychol. Sci. 20:464. · Potter et al. (2014). Atten. Percept. Psychophys. 76:270. · Thorpe, Fize & Marlot (1996). Nature 381:520. · Hauk & Pulvermüller (2004). Clin. Neurophysiol. 115:1090.
- Shepard & Metzler (1971). Science 171:701. · Kay, Keogh & Pearson (2024). Conscious. Cogn. 121:103694. · Zeman, Dewar & Della Sala (2015). Cortex 73:378. · Liu & Bartolomeo (2023). Cortex 166:338. · Pearson (2019). Nat. Rev. Neurosci. 20:624. · Dawes et al. (2020). Sci. Rep. 10:10022.
- Balaban & Ullman (2025). Physics versus graphics as an organizing dichotomy in cognition. TiCS 29:985.
- Spelke & Kinzler (2007). Core knowledge. Dev. Sci. 10:89. · Baillargeon, Spelke & Wasserman (1985). Cognition 20:191. · McGarrigle & Donaldson (1974). Cognition 3:341.
- Robocurve (2026-09-04). GPT-6 Astra on robot arms. report
- OpenAI (2026-09-08). On the Navier–Stokes Millennium Prize Problem. report · Willison, S. (2026-09-08). Some thoughts on the Navier–Stokes Millennium Prize Problem. post
- Hinton, G. (2025). Interview with Alok Jha, GITEX Europe, Berlin. Transcript in Business Powerhouse, 7 July 2025. transcript
Figures reproduced from the cited papers under their licenses or with attribution for commentary.