Collage: a generator AI chip passes a program through a Turing machine and produces bytes a learner chip has to predict; a reward arrow closes the loop and a label reads zero human data

Pretraining

Two models train each other without any human data to start from


The method starts from random initialization and improves predictably with the compute invested, according to a paper published on September 24.

September 27, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

0 examples of human data at the procedure’s starting point.

Context

The debate over the end of training data had been playing out along two paths: buying licenses and generating synthetic data from already-trained models. This paper tests a third, which doesn’t depend on a prior teacher model but on a search over the space of computable structures, with Solomonoff induction as the theoretical reference.

What’s next

Bottom line

Pretraining with no data sounds like a lab trick, and it probably will be for several years. What isn’t a trick is the curve: if it improves with compute rather than with archives, it changes the list of who can compete.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes