LUM Large Universe Model

Large Universe Model/Essays/Surprise is the signal

Essay

Surprise is the signal

A model that only tells you what it currently thinks is withholding the more useful half.

Two outputs, not one

Ask a Large Universe Model what it believes and you get an estimate with an interval. Useful, and available in weaker form from other systems.

The output that has no equivalent elsewhere is: this belief just moved a long way, and here is what moved it.

Neither a Large Language Model nor a Large World Model can produce it, for the same structural reason: neither holds a belief long enough for it to move. Each response is generated fresh. There is no yesterday's position to compare against today's.

Why magnitude beats threshold

Conventional alerting fires when a value crosses a line. This produces two well-known failures: alerts during expected excursions, and silence during genuinely anomalous behaviour that happens to stay inside the line.

Revision magnitude is a better trigger because it is relative to what the model already expected. A metric moving to a value the model had been anticipating for a week is not news. A metric moving a small absolute distance that the model was highly confident it would not move is news.

The quantity that matters is the movement measured against the confidence that preceded it.

This is the practical argument for the whole sequence. The Large Language Model can describe a situation. The Large World Model can simulate what might follow. Only the Large Universe Model can tell you that it changed its mind, when, and why.

Surprise as a measure of model quality

Persistent large revisions in the same belief mean the model's picture of that part of the domain is wrong — some driver is missing. Aggregated over time, the distribution of revision sizes is a running self-assessment.

A well-specified Large Universe Model should be surprised rarely and informatively. One that is surprised constantly is not tracking; it is reacting.

The failure mode

Optimising for surprise reporting produces a system that manufactures it — reporting movement because movement is what gets attention. The defence is calibration against outcomes: a revision that was not followed by anything is evidence the trigger is too sensitive, and it should feed back into the threshold.