Large language models require up to 15 trillion tokens to master human language, exposing a massive data efficiency gap when compared to human children who achieve fluency using only about 100 million words. According to cognitive scientists and researchers studying this linguistic disparity, closing that gap could democratize artificial intelligence and enable the training of smaller models for minority languages.
The Data Efficiency Gap Between Humans and AI
Human children master their native tongue by hearing roughly 100 million words by the time they reach adolescence. Stanford University cognitive scientist Michael C. Frank points out that training modern AI systems requires burning through vast computational resources to replicate a process that naturally occurs in domestic environments. According to Frank, children begin producing grammatically correct sentences after hearing 10 to 30 million words, whereas GPT-2 trained on a similar 30-million-word corpus only generates nonsense.
To illustrate the scale mismatch, Georgetown University cognitive scientist Ethan Gotlieb Wilcox notes that large language models process amounts of text equivalent to the combined lifetime output of an entire city in a single generation. If printed on paper, the text used to train a modern LLM would stretch past the International Space Station, while a child’s 100-million-word exposure stack would reach a height of only 20 meters.
Challenging Longstanding Linguistic Assumptions
For decades, cognitive science relied on MIT linguist Noam Chomsky’s 1950s theory that infants possess innate grammatical knowledge because human language is too complex to learn purely from experience. However, the success of large language models in learning syntax purely from statistical data analysis has challenged this foundational assumption. University of California, Berkeley psychologist Alison Gopnik admits she was initially surprised that AI systems could acquire syntax through statistical observation of large text samples.
Did you know? While Noam Chomsky argued that children are born with built-in grammar rules, modern large language models learn syntax simply by analyzing the statistical probabilities of word sequences across massive datasets.
BabyLM Project Imitates Child Language Acquisition
To test whether AI can learn more efficiently, University of California, San Diego linguist Alex Warstadt launched the BabyLM competition in 2022. The challenge asks researchers to train language models using developmentally plausible corpora of 100 million words for toddlers and 10 million words for children, drawn from sources like storybooks, movie subtitles, and child-directed speech transcripts.
The 2024 winning model, GPT-BERT, was trained on roughly 100 million words and outperformed Meta’s Llama 2 70B—a model trained on 15,000 times more data—on specific BabyLM benchmarks. Despite this efficiency, these smaller models still lag far behind commercial LLMs and frequently fail to generate coherent text. Boston University researcher Aaron Mueller notes that while curriculum learning—ordering training data from simple to complex—seems intuitively helpful, transformer models do not actually require that structured progression to learn effectively.
Capturing Early Childhood Experience Through Head-Mounted Cameras
Researchers are also turning to direct observation of childhood to improve machine learning architectures. Michael Frank leads the SAYCam project, which recorded three toddlers using head-mounted cameras for two hours per week between the ages of six months and two and a half years. Princeton researcher Brenden Lake used these video recordings to train models capable of identifying objects and associating them with words without relying on built-in biases.

Similarly, Princeton’s Uri Hasson recorded 17 children during their first 1,000 days of life, capturing 12 hours of video and audio daily using cameras and microphones installed throughout participating homes. Harvard researcher Elizabeth Bonawitz emphasizes that unlike passive AI models, human children actively experiment, choose their own learning experiences, and reason about the social context and knowledge of their teachers.
Future Applications for Minority Languages and Smaller Research Labs
Closing the data efficiency gap holds profound implications for low-resource languages and independent academic institutions. University of Oslo machine learning researcher David Samuel points out that minority languages like Sami often possess only tens of million available tokens—roughly equivalent to a toddler’s linguistic exposure. Developing models that function capably on such limited data remains a primary objective for preserving linguistic diversity in technology.
Alex Warstadt aims to democratize the field so that universities and smaller teams without massive corporate resources can train competitive AI models. Meanwhile, researchers increasingly utilize language models as model organisms—imperfect computational proxies that allow scientists to study human language learning much like laboratory test subjects, opening a window into how linguistic comprehension actually functions.
Frequently Asked Questions
What is the data efficiency gap in artificial intelligence?
The data efficiency gap refers to the vast difference in the amount of data required for a machine learning model to learn language compared to a human child. While children achieve fluency after hearing about 100 million words, modern LLMs require trillions of tokens.
What is the BabyLM project?
Initiated by linguist Alex Warstadt, BabyLM is an annual research competition that challenges scientists to train language models on restricted, child-sized text corpora to study how models can learn language more efficiently.

Can AI models learn language from video recordings of children?
Yes. Projects like SAYCam and infant head-mounted camera datasets record early childhood environments to train multimodal AI models in associating words with objects without relying on pre-programmed linguistic assumptions.
What are your thoughts on data-efficient AI? Do you think mimicking human childhood learning is the key to creating sustainable language models, or will brute-force scaling continue to dominate? Leave a comment below, explore our related articles on machine learning, and subscribe to our newsletter for the latest updates.
Worth a look