Children learn language with less data, and no one knows why

Humans have been speaking to each other for at least a hundred thousand years, and throughout that time only one kind of learner has ever mastered a language to native level: a human child. Now there are two. Four years after the release of ChatGPT, many of us take for granted that we can hold a natural conversation with a phone or laptop. Large language models such as Claude, DeepSeek and the models from OpenAI reach a level of fluency and flexibility that lets them convincingly masquerade as humans. Yet behind the computational curtain lies a gap: teaching a computer to use human language still demands a non-human amount of data.
A large language model can easily ingest a hundred thousand times more words than a person experiences during first-language acquisition, and far more than children hear before their first birthday, when they typically begin to grasp language. “The recent progress is amazing,” says Michael C. Frank, a cognition researcher at Stanford, “but we still need to burn an entire forest and scrape all of human knowledge to reproduce a milestone that occurs in our living room over the course of a year.” This disparity between children and machines is called a “data efficiency gap,” and it poses a tantalizing question for cognition researchers and a challenge for AI model architects: how do children still outpace the most sophisticated language machines ever built?
The scale difference can only be illustrated with analogies. “Claude has seen a whole generation’s worth of raw language,” says Ethan Gotlieb Wilcox, a cognition researcher and linguist at Georgetown. If you printed on paper all the words used to train a modern model, the stack would extend beyond the International Space Station. By contrast, a hundred million words from a pre-adolescent human would pile up to only twenty meters, and children manage with far less. Meta’s Llama 3.1, released two years ago, consumed 15 trillion tokens during pre-training, the main phase before fine-tuning for a specific task such as a chatbot. Frontier models may train on ten times more data, but there is only so much internet, and by the early thirties the well of readily available data could begin to run dry.
Reverse-engineering how children learn could enable data-efficient AI models useful for everything from effective video training to chatbots that serve minority-language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and developing brains: whether we are born with a language instinct, or whether, in principle, a child could learn language purely from experience; whether our language processing is a biological quirk or reflects universal constraints on how languages can be used and learned.
Most of us realize that language is hard only when we try to learn a new one after childhood. Perfect tense, a rolling r, vowel shifts, genitive constructions, complex verbs, masculine nouns for tables and feminine nouns for spoons—these are just some of the linguistic hurdles that adults face when learning a second language.