LUM Large Universe Model

Large Universe Model/Reference/From Large Language Model to Large Universe Model

Reference

From Large Language Model to Large Universe Model

Three generations, each widening what a machine is permitted to take in. Read in order, the sequence looks less like a series of surprises than a single movement.

1960 — the machinery arrives early

Sequential estimation — maintaining a belief about a system's state and revising it as measurements arrive — is formalised in control theory. Everything the Large Universe Model does with a single belief was mathematically available six decades before there was anything worth pointing it at.

What was missing was not the technique. It was the ability to read unstructured evidence, which had to wait for the Large Language Model.

2017 — the architecture

Attention Is All You Need introduces the transformer. Every model in all three generations still runs on it.

2018 — the first world models

World Models proposes that an agent can compress observations into a learned latent model and plan inside it. The idea that a network could hold an internal simulation, rather than a mapping from input to output, begins here.

2019 — the argument for scale

Sutton's Bitter Lesson argues that general methods which scale with computation defeat hand-engineered knowledge. It is the clearest statement of why each generation of this sequence displaces the last.

2020 — the Large Language Model era begins

Language Models are Few-Shot Learners establishes that scale alone produces general capability. The Large Language Model becomes the dominant form, and with it arrives the limitation that defines the next six years of work: the training cutoff.

2020–2024 — working around the cutoff

Retrieval-augmented generation, long context windows, tool use and fine-tuning all appear as ways of carrying fresh evidence to a model that cannot acquire it. Each is useful. None removes the underlying property.

2024 — the Large World Model

World Model on Million-Length Video and Language pushes context to a million tokens across video and text, long enough to model extended experience rather than a frame. The Large World Model is named, and the second generation begins in earnest.

2024 — the continuity problem is documented

Loss of plasticity in deep continual learning shows that networks trained continually progressively lose the ability to learn, eventually performing no better than a shallow network. This establishes that continuous ingestion is a research problem and not merely an engineering one — and it shapes how the third generation is built.

2025 — world models become interactive

Genie 3 produces navigable environments in real time without hard-coded physics. NVIDIA Cosmos makes world foundation models infrastructure. V-JEPA 2 predicts in representation space. The Large World Model matures — and remains, in every case, frozen at training time.

Throughout — continuous systems in the field

While all of this proceeds, astronomy runs alert brokers over millions of nightly detections, market surveillance maintains live models of normal behaviour, reliability engineering runs on revised baselines, and epidemiology maintains nowcasts that revise the past as well as the present.

These are Large Universe Models with domain restrictions, built by hand, one field at a time.

2026 — the Large Universe Model

The third generation is what happens when the pattern stops being bespoke: continuous ingestion and belief revision, using the language understanding of the Large Language Model and the sense of consequence of the Large World Model, applied to any stream rather than to one discipline that had no alternative.

The sequence ends here, and not for want of ambition. There is no fourth widening available. The Large Language Model took a corpus, the Large World Model took a scene, and the Large Universe Model takes everything still happening. After that, the remaining work is scale, trust and time.