Think of a thing. Anything. The rule is that I get to ask you questions — yes/no answers only — and I have twenty of them. Anyone whose childhood happened in the nineties, or not much later, will remember Akinator and its many cousins: officially it came out a few years afterwards, but the idea it stands on is far older than that.

What few people notice is that if I play well — which really means, if every question cuts in half whatever is left — twenty answers are enough to find one object in a million. And the reason is embarrassingly simple:

$$2^{20} \approx 1{,}000{,}000$$

Which is also why people win this game routinely, so imagine how a bit of JavaScript on an old website does. There is nothing magical about the game. What is magical is that in those same few seconds spent choosing an answer, your mind ranged over something like a million possibilities and settled on one. Let’s say that speed comes to about ten bits per second (we’ll get there shortly).

Now put that number next to another one. While you were picking your answers, your eyes were pouring roughly a billion bits per second into your head. Same head, same seconds, wildly different amounts of information.

Water spilling into the Arizona spillway at the Hoover Dam, July 1983

Water spilling into the Arizona spillway, Hoover Dam, July 1983. Bureau of Reclamation, public domain.

That gap is the subject of a paper by Jieyu Zheng and Markus Meister at Caltech, published in Neuron under a title that reads almost like a complaint: The Unbearable Slowness of Being: Why do we live at 10 bits/s?1

I came across it by accident, and what struck me is that they did almost nothing new. They did something rarer instead: they added up a century of old experiments, in a single unit, and found that the sum does not add up.

One ruler for two different things

Two numbers this far apart — 10 bit/s and 10⁹ bit/s — mean nothing until we can show they were measured, let’s say, the same way. In what sense? In the sense that a thought inside your head and a retina being hit by light have almost nothing in common: one is a person finishing a sentence or picturing an object, the other is a voltage wobbling inside a cell. To compare the two you need something, a measure, that both can be converted into. The lucky part is that this something already exists, and it was invented in 1948 — not by a psychologist, incidentally.

Claude Shannon was working out how much traffic a telephone line could carry, and to answer that he first had to say what a channel is really transmitting. What happened next is that his answer turned out to hold for anything that carries messages: a wire, a neuron, a person.

Shannon’s move was to stop asking how much data is there and start asking how surprised am I. The idea being: if I can predict what you are about to do, watching you do it teaches me nothing. And on the other hand, if I cannot, it teaches me a great deal.

The simplest case to put this in is a coin toss. A normal coin has two outcomes you cannot call in advance, so the toss is worth one bit of information (assuming the coin lands heads or tails, and not on its edge). A coin with heads on both sides has an outcome you can call perfectly, so we can say the toss is worth nothing at all: no information, just theatre. Written out for a set of possible outcomes, each with its own probability:

$$H(A) = -\sum_i p(a_i) \log_2 p(a_i)$$

where \(H(A)\) is the entropy of the source — mathematically, the weighted average of the self-information of the symbols it can emit, against their probability of being emitted. Take that quantity, divide it by time, and you get so many bits per second: a unit that holds equally for a nerve fibre, a pianist and a fibre-optic cable. It is a measuring convention, clearly, not a theory of mind — the authors are careful to say so — but it is what lets you write both sides of this paradox on the same page.

The typist who is slower than her fingers

Let’s start with someone whose information out is easy to count: a professional typist, running at about ten keystrokes a second.

The way that comes to mind for scoring her is to count the keys on the keyboard, fifty or so, and conclude that every keystroke carries a fair bit of information. It is the wrong answer, and the reason the authors give is, in my opinion, genuinely brilliant.

What they say — and this reminds me of security courses and the statistical cryptanalysis part — is that English is not a random stream of letters, and neither is any other natural language. After informatio you already know what comes next. This, by the way, is the same thing that happens when you finish somebody else’s sentence (sound familiar?): it works because those words were predictable, so hearing them — at least from the point of view of Claude Shannon, and he could have been called anything, but he happens to be called Claude — taught you almost nothing.

Not content with a hunch, Shannon actually measured it in 1951, by showing people half-finished sentences and asking them to guess the next character. The result: English carries about one bit per character.2 Most of what a typist types was already implied by what came before it. Which makes her real rate:

$$I = 2 \, \frac{\text{words}}{\text{s}} \cdot 5 \, \frac{\text{characters}}{\text{word}} \cdot 1 \, \frac{\text{bit}}{\text{character}} = 10 \, \frac{\text{bits}}{\text{s}}$$

And here is the proof that this ten-bits business is real and not just a number someone liked: want to know how we can tell? Ask that same typist to type random characters and her speed collapses. Which means her fingers were never the limit. What the typist is doing is riding predictability, the way a compression algorithm does. The hands deliver ten keystrokes a second; the mind feeding them moves at ten bits a second. Okay, but… that is one case. The mind-blowing claim is this: if someone told you that this number refuses to move, would you believe them?

Measure people while they read, or listen, or play a video game at professional level, or memorise a shuffled deck of cards, and you always end up in the same place: people move at a few bits per second, a few dozen at the very best.

The prettiest confirmation comes from outside neuroscience. Linguists compared seventeen languages and found that the fast ones pack little information into each syllable and the slow ones pack a lot — and the two effects cancel out almost exactly, leaving every language delivering information at roughly the same rate as the others.3 Languages that look nothing alike settle on the same throughput, as if each had been designed, or let’s say negotiated, with the very same listener.

What is remarkable is that memory gives the same answer from a third direction. In the eighties Thomas Landauer measured how much people retain of what they see, read and hear, and found that the amount comes to roughly a couple of bits per second, whatever the material.4 Add that up across a lifetime and everything you know comes to something on the order of a gigabyte.

So we have some sense, through a few examples, of the information that comes out. You have no idea how much comes in.

Let’s count what comes in

A single light-sensitive cell in our retina is a small analogue device. The incredible thing that falls out of the arithmetic is that if you add up how many distinct signals it can send per second, you find that that one cell alone carries more information than your whole life as a sentient being does. And we are not talking about the eye here — just one cell inside it.

Now, given that we have millions of them per eye, each watching its own patch of the world, and this before counting ears, skin, or the sense that tells you where your left elbow is, let’s just say that all together the other side of the coin — the information coming into the body — comes to around a billion bits per second.

Divide one number by the other and you get what the paper calls the sifting number, a sort of measure of how much of what arrives gets thrown away:

$$Si = \frac{\text{information coming in}}{\text{behaviour going out}} \approx \frac{10^9}{10} = 10^8$$

The ratio is a hundred million to one. Zheng and Meister use a genuinely good image to give that a size: the Hoover Dam moves water about a hundred million times faster than a person drinks. Put another way, the thinking mind, listening to the senses, is the equivalent of a person standing at the foot of that dam with their mouth open.

Here is that same picture, with all the numbers from the paper along one axis.

11010010⁴10⁶10⁸10¹⁰bits per second, each step ten times the one beforeyouthinking, typing, speakingyour senseswhat the eyes alone delivereverything in between is thrown away, and nobody knows how

There is a domestic version of this same gap. Picture the home WiFi playing up, and how much that matters because the film or the show you are streaming might stall. Now picture your eyes taking in every one of those millions of bits coming off the screen, along with your ears and the rest. And yet what you are left with, after two hours of film, is who did what to whom and how you felt about it — a few sentences, maybe. We are effectively paying for enough bandwidth to cover a stadium full of people, and then feeding the equivalent of a keyhole.

Now the interesting question becomes: what if we got it wrong? What if the measurement is off, and the missing information is sitting somewhere all these experiments never looked? Let’s say there are three places it could be hiding.

The three hiding places

Photographic memory. If anyone could really take mental snapshots, they would win every memory championship on the planet. Instead the record holders sit on the same few bits per second as everybody else: just trained to spend them well.

The rest of the visual field. This one feels like an objection with no answer, because our visual world seems — only seems — sharp everywhere and all at once. You can check it while reading this sentence: hold your eyes on one word and, without moving them, try to read the line two below. The words are there. They look perfectly crisp. And yet you cannot read them, and you cannot because they are not reaching you — or rather, they are reaching you, but your mind has no way to process them.

The technical name for the gap is subjective inflation. Its most famous demonstration is the gorilla experiment: there is a video of people passing a ball around, and viewers are asked to count the passes. At the end they are asked whether they noticed a person in a gorilla suit walk into the middle of the shot and beat their chest.5 The answer is reliably no. The detail feels present because it is available: look, and you will find it there, right in the centre of the frame. Which confirms something as simple as it is invisible: available is not the same as stored.

Tor Nørretranders built a whole book on this back in 1991: in The User Illusion: Cutting Consciousness Down to Size he wrote about senses gathering millions of bits, awareness resurfacing with a single handful, and the self as a kind of user interface sitting on top of consciousness.6 His intuition tells us the idea was already in the air. What Zheng and Meister add is a single measure, one unit: every measurement forced onto the same axis.

Unconscious learning. The third hypothesis, and we can certainly say that the brain absorbs far more than it can report, filing it below the threshold of awareness. The paper uses the example of cats: raised in a room painted with vertical stripes, they grow up with a visual cortex tuned to match. Impressive — true — but how much information did that take? In the ordinary world edges arrive from every angle, let’s say a range of 180°; in the striped room, only from a narrow slice, say 40°. So by Shannon’s account, the information attached to a change of probability like that is nothing more than:

$$\log_2 \frac{180°}{40°} \approx 2 \text{ bits}$$

Two bits, for weeks of immersion. And this tells us, once again, something crucial: learning what the world is usually like is cheap, because “usually” is a simple shape. It has nothing to do with storing what you saw.

So none of the three hypotheses really holds. What a swindle. But then why does the mind run one thing at a time?

Why one thing at a time

The first suspicion is that neurons are simply poor components, and that you therefore need an enormous number of them to do anything at all. But the science and the experiments tell us the opposite: individual neurons are precise, and the brain keeps remarkably few spare copies of anything. Whatever the bottleneck is made of, we can say with near certainty that it is not made of unreliable parts.

The second suspicion is that some central resource is scarce. But the jobs we blame the bottleneck for are not demanding: a realistic model of a decision takes a few thousand neurons, and you could fit hundreds of those in a square millimetre of cortex.

There is one thing, though, that is undeniably true and worth sitting with: the input is parallel, while the processing, the control centre, is not. Our retina handles the whole scene at once. Our thinking does not: give someone two tasks and they will do one at a time. The same thing happens in a computer, by the way. Multitasking is nothing but a system emulating simultaneity, because the processor — or the processors, which do duplicate the work, but across multiple instances — is dramatically faster than every other piece of hardware, and in the time it takes for information to reach it, it effectively simulates the future, untangles chains of instructions, and works out which instruction to run next, slotting each one behind the other at enormous speed, making the path it takes to finish any job invisible. A path that is, by its nature, serial.

Another example is the cocktail party: at an event like that you can easily lock onto one voice in the room, noisy as it is — the signal processing involved is genuinely hard — but you can only do it for one voice.

These examples only tell us that the bottleneck is there, though, not why. The explanation I find hardest to argue with starts from an odd question: what was a brain for, before it was for thinking?

It was for moving. Nervous systems exist only in animals that go somewhere, and the first job of a brain was steering a body through space — not outer space, the one immediately around us. An organism working its way up a trail of scent has one question to ask, and it asks it over and over: which way, now. And that question takes exactly one answer, because the body is in exactly one place. Of five possible directions, four describe a world you will not be standing in: working them all out would be effort thrown away. A brain built for that job never had a reason to learn to do two things at once.

Thinking, on this hypothesis, is the same machine running with the muscles held still: instead of crossing a room, you cross an argument. And it carries the constraint of the original job: you can only be in one place at a time, even when the place is an idea.

Two brains at two speeds

What survives all of this is a claim about architecture, and it is the part most likely to still matter in twenty years.

The outer brain is wide and fast: millions of sensors and muscle fibres, everything happening at once, and reasonably well understood — we know why it needs so many neurons.

The inner brain is narrow and slow: a few bits per second, combining goals, memory and the present moment into the next move, with the ability to change its mind the instant the situation does.

OUTER BRAINsenses and muscles · everything at oncewide and fast99.999999% droppednobody knows howINNER BRAINone thing at a timenarrow and sloweverything you dosame cortex · similar cells · same second

The awkward part is that both halves are made of the same tissue, with similar cells and similar neuron counts. And yet the people studying vision describe a vast, detailed space, while the people studying decisions routinely boil a million neurons down to two or three numbers.

Zheng and Meister raise an unpleasant possibility: maybe the inner brain looks simple because our experiments are simple. Imagine we had only ever studied vision by showing animals a slowly rotating propeller. We would have found neurons responding, reduced the data to two dimensions, decoded which way the propeller was turning, and published: all of it true, and none of it touching what vision actually does. Receptive fields and orientation maps only showed up once someone gave the eye a rich, complicated world.

And what do we show the inner brain in the lab? A mouse choosing between the same two options, over and over. Rotating propellers, every one of them.

In Nature Neuroscience, Britton Sauerbrei and Andrew Pruszynski pointed out that the slow figure comes from tasks with countable outcomes — presses, moves, keystrokes — while most of the nervous system is busy with continuous control: standing, walking, catching a glass that slips.7 That machinery is fast, and almost entirely unconscious.

I think they are right, and I think it sharpens the claim rather than breaking it. What was measured is the throughput of the part of you that deliberates and reports, the part you can ask what it is doing. So the honest version is narrower and more interesting: the bottleneck sits at the conscious interface, not in the hardware.

Your body is broadband. The self that reads papers, picks careers and sits in meetings is a modem sitting on top of it.

What changes

Here is where the paper stops talking about neurons and starts talking about the things we build for that self to speak through.

Musk’s stated reason for Neuralink is a bandwidth problem: “you just can’t communicate through your fingers, it’s just too slow.” The paper turns it into a prediction: however many electrodes you implant, the channel will carry about ten bits per second, because the limit sits upstream of the fingers. Their deadpan observation is that a device for that data rate already exists, and it is the telephone.

The point is not that implants are useless. The point is to change the target. If you are restoring sight, do not pipe a megapixel stream into a system that will throw nearly all of it away: send the conclusion. A camera that says “your daughter is on your left and she’s smiling” delivers the bits that would have survived anyway, and it runs on a phone.

The same arithmetic resizes a question we usually ask backwards. We ask how much the brain could hold, and we get enormous numbers. The better question is a different one: how much can get in? And at ten bits a second, across a lifetime, the answer is about a gigabyte. Meanwhile we train language models on more text than a human could read in a thousand lifetimes. Which means we now know that whatever human intelligence is, it is not built on volume.

That is where the paper ends. What follows is a consequence it does not draw, but one I find hard to avoid if you build things with language models.

For a couple of years now we have all been pushing in the same direction: bigger context windows, more tokens per second in and out, longer answers, fuller reports. These are real improvements, and you see them immediately.

The problem is where they end up. At the other end of the pipe there is always a person, and that person is rated for ten bits per second. A model generating two hundred tokens a second at a human being is a dam pointed at a mouth: you can open the gates as far as you like, it does not change how much they drink.

So the scarce resource, in a system made of people and models, was never the output. It is the last ten bits, the only ones that actually make it across.

And if the narrow part is the final stretch, generating stops being the interesting problem: choosing becomes it. Pierre Lévy was saying as much in Collective Intelligence: Mankind’s Emerging World in Cyberspace.8 What is worth maximising is not how much you produce, but how much surprise you deliver — the part the reader could not have predicted on their own. Which is the definition of information we started from, coming back around from the other side.

And this is where a chat with an AI shows its weakest point: it answers very well whatever you can already put into words. But you cannot type a question about something you do not know exists — by definition, it cannot even occur to you (if you don’t know about it, how would you ask?).

what you knowwhat you don't knowyou knowit's thereyou don't knowit's therewhat you can already doyou can ask the questionsearch and chatbots live hereskills you can't explainthe question doesn't exist yetthe ten bits have to be chosenfor you, not by youthe scarce thing is the opening, not the size of the model

Three different groups have started working on that bottom-right box, arriving at it from opposite directions.

Co-STORM, out of Stanford, gets rid of the search box. Several language agents hold a conversation in front of you and ask each other the questions that would not have occurred to you, while a map keeps track of the ground covered. You listen, and step in when something catches your interest.9

A team at DeepMind started from the opposite end: they went looking inside AlphaZero for the concepts that never show up in human chess, kept the ones that could be explained, and taught them to grandmasters. Ideas nobody had found in fifteen centuries of the game.10

And a study of several hundred students gives the same thing a name on the human side: the unplanned encounter, useful information that arrives without being looked for, which turns out to predict creativity by way of mental flexibility.11

Three directions, one shape. The value is not in the machine knowing more: it is in the machine being the one that chooses what gets through the opening.

There is a second consequence, and this one is less comfortable. If you read at ten bits a second and a model crosses a codebase in an instant, the difference between you is not intelligence. It is bandwidth.

Which makes the review bottleneck permanent rather than temporary: you will never read what it wrote at the speed it writes it, and no model update will change that. Any workflow that assumes otherwise is lying… unless somebody out there is planning an upgrade to human bandwidth.

All of which sounds a little like a verdict against us, as if whoever can finally do this better than we can might one day take our place (one day? let’s say more like today, or maybe yesterday).

Ten bits, and no refunds

Ten bits a second is a brutal budget. Everything you will ever learn has to come through it. Everything you say, decide or make leaves through the same opening. We said that over a lifetime it comes to about a gigabyte.

And yet everything the species has ever made came through that straw. Every theorem, every cathedral, every piece of music, every line of code that ever ran in production. Not despite the filter — but through it.

The billion bits a second were never the valuable part. The valuable part was always the sifting: a hundred million arriving, and one of them chosen.

So the interesting question was never how to widen the pipe. It is how to get better at choosing: a question about taste, about attention, and about what you point yourself at. Which turns out to be the same question we are now asking on behalf of our machines. As Pierre Lévy was saying back in the nineties, as the internet arrived, the expert will no longer be whoever has access to information, but whoever can select it.8 Which is what we have always done by googling things — only at a much slower bandwidth, let’s say 56 Kbit/s, and still far more than our human capacity.


  1. Zheng, J. & Meister, M. “The Unbearable Slowness of Being: Why do we live at 10 bits/s?” Neuron 113(2), 192–204, 2025. PMC11758279 · arXiv:2408.10234 · DOI:10.1016/j.neuron.2024.11.008. Unless said otherwise, the figures in this post come from the paper, where the full arithmetic and sources are laid out. ↩︎

  2. Shannon, C. E. “Prediction and Entropy of Printed English.” Bell System Technical Journal 30(1), 1951. ↩︎

  3. Coupé, C., Oh, Y. M., Dediu, D. & Pellegrino, F. “Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche.” Science Advances 5(9), 2019 — seventeen languages, all landing near 39 bits/s. DOI:10.1126/sciadv.aaw2594↩︎

  4. Landauer, T. K. “How Much Do People Remember? Some Estimates of the Quantity of Learned Information in Long-Term Memory.” Cognitive Science 10(4), 1986 — about two bits per second, and roughly \(10^9\) bits over a lifetime. PDF↩︎

  5. Simons, D. J. & Chabris, C. F. “Gorillas in our midst: sustained inattentional blindness for dynamic events.” Perception 28(9), 1999. ↩︎

  6. Nørretranders, T. The User Illusion: Cutting Consciousness Down to Size. Viking, 1998 (Danish original 1991). Estimates in that literature range from 16 to 60 bits/s depending on the method. ↩︎

  7. Sauerbrei, B. A. & Pruszynski, J. A. “The brain works at more than 10 bits per second.” Nature Neuroscience 28, 1365–1366, 2025. DOI:10.1038/s41593-025-01997-0↩︎

  8. Lévy, P. L’intelligence collective. Pour une anthropologie du cyberspace. La Découverte, Paris, 1994. English edition: Collective Intelligence: Mankind’s Emerging World in Cyberspace, Plenum, 1997. Italian edition: L’intelligenza collettiva. Per un’antropologia del cyberspazio, Feltrinelli, 1996. ↩︎ ↩︎

  9. Jiang, Y., Shao, Y. et al. “Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations.” EMNLP 2024. arXiv:2408.15232↩︎

  10. Schut, L. et al. “Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero.” arXiv:2310.16410, later in PNAS, 2025. ↩︎

  11. “Serendipitous Sparks: AI Information Encounter, Cognitive Flexibility, AI Literacy, and University Student Creativity.” PMC12689981↩︎