Large Universe Model/Reference/From Large Language Model to Large Universe Model
ReferenceFrom Large Language Model to Large Universe Model
Three generations, each widening what a machine is permitted to take in. Read in order, the sequence looks less like a series of surprises than a single movement.
1960 — the machinery arrives early
Sequential estimation — maintaining a belief about a system's state and revising it as measurements arrive — is formalised in control theory. Everything the Large Universe Model does with a single belief was mathematically available six decades before there was anything worth pointing it at.
What was missing was not the technique. It was the ability to read unstructured evidence, which had to wait for the Large Language Model.
2017 — the architecture
Attention Is All You Need introduces the transformer. Every model in all three generations still runs on it.
2018 — the first world models
World Models proposes that an agent can compress observations into a learned latent model and plan inside it. The idea that a network could hold an internal simulation, rather than a mapping from input to output, begins here.
2019 — the argument for scale
Sutton's Bitter Lesson argues that general methods which scale with computation defeat hand-engineered knowledge. It is the clearest statement of why each generation of this sequence displaces the last.
2020 — the Large Language Model era begins
Language Models are Few-Shot Learners establishes that scale alone produces general capability. The Large Language Model becomes the dominant form, and with it arrives the limitation that defines the next six years of work: the training cutoff.
2020–2024 — working around the cutoff
Retrieval-augmented generation, long context windows, tool use and fine-tuning all appear as ways of carrying fresh evidence to a model that cannot acquire it. Each is useful. None removes the underlying property.
2024 — the Large World Model
World Model on Million-Length Video and Language pushes context to a million tokens across video and text, long enough to model extended experience rather than a frame. The Large World Model is named, and the second generation begins in earnest.
2024 — the continuity problem is documented
Loss of plasticity in deep continual learning shows that networks trained continually progressively lose the ability to learn, eventually performing no better than a shallow network. This establishes that continuous ingestion is a research problem and not merely an engineering one — and it shapes how the third generation is built.
2025 — world models become interactive
Genie 3 produces navigable environments in real time without hard-coded physics. NVIDIA Cosmos makes world foundation models infrastructure. V-JEPA 2 predicts in representation space. The Large World Model matures — and remains, in every case, frozen at training time.
Throughout — continuous systems in the field
While all of this proceeds, astronomy runs alert brokers over millions of nightly detections, market surveillance maintains live models of normal behaviour, reliability engineering runs on revised baselines, and epidemiology maintains nowcasts that revise the past as well as the present.
These are Large Universe Models with domain restrictions, built by hand, one field at a time.
2026 — the Large Universe Model
The third generation is what happens when the pattern stops being bespoke: continuous ingestion and belief revision, using the language understanding of the Large Language Model and the sense of consequence of the Large World Model, applied to any stream rather than to one discipline that had no alternative.
The sequence ends here, and not for want of ambition. There is no fourth widening available. The Large Language Model took a corpus, the Large World Model took a scene, and the Large Universe Model takes everything still happening. After that, the remaining work is scale, trust and time.