LUM Large Universe Model

Large Universe Model/Applications/Large Universe Models for scientific literature

Application

Large Universe Models for scientific literature

Confidence in a published finding ought to fall when a replication fails. In practice it rarely does, because nobody re-reads their own bibliography.

What it ingests

PubMed and preprint servers, trial registries and their pre-registered analysis plans, replication reports and multi-lab studies, retraction notices, conference abstracts, and the citation graph.

The propagation failure

The citation graph keeps carrying results the field has quietly stopped believing. A finding is published, cited, built upon, and cited again. A replication fails three years later, is published somewhere less visible, and changes nothing about the downstream literature because no mechanism connects the two.

A Large Universe Model maintains a posterior per finding. When a replication fails, confidence in that finding falls, and confidence in every conclusion that depends on it falls in proportion.

The sequence is always the same. A Large Language Model read a corpus once and stopped. A Large World Model learned to simulate a scene it was shown. A Large Universe Model keeps watching, and revises.

Weighting evidence properly

Not every study should move a belief equally. A Large Universe Model weights by:

  • Sample size and the width of the reported interval
  • Whether the analysis was preregistered
  • Independence of the replicating group
  • Effect size relative to the original claim
  • Known incentives and publication-bias characteristics of the venue

This is standard meta-analytic practice. What is new is doing it continuously, across an entire literature, and maintaining the result rather than producing it once as a paper.

What you can ask it

What can I currently rely on in this area? — answered as of this morning, with each claim carrying a confidence and the studies behind it, and a note of which of your own citations have moved this quarter.

Limits

Quality weighting encodes judgements that researchers legitimately disagree about, and a badly calibrated weighting is worse than none. The defensible use is surfacing where the evidence base has shifted, with the reasoning exposed for argument, rather than issuing verdicts.