What happened
- A paper posted to arXiv on September 24 describes a pretraining procedure that uses no starting data. It is signed by Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman and Yoav Levine.
- The setup has two models learning in parallel from random initialization. One proposes programs that a universal Turing machine interprets to generate byte sequences; the other predicts those sequences.
- The generator is trained with reinforcement learning to produce sequences right at the edge of what the learner can solve. That builds a curriculum that adjusts itself as the learner improves.
- The authors report that performance on natural data improves predictably with the compute invested in self-play, and that in-context learning and recognizable mathematical sequences emerge during training.
- The paper presents itself as an initial proof of concept, not as an alternative to current pretraining.
Why it matters
- If the source of data stops being available text and becomes compute, the competitive advantage shifts from whoever has archives to whoever has installed capacity. For the region, that’s bad news: archives can be licensed; data centers have to be built.
- The entire debate about copyright in training assumes data comes from somewhere and that someone claims it. A procedure with no starting corpus doesn’t resolve the open lawsuits, but it defines a lane where they don’t apply.
- For anyone buying models, the practical point is different: a self-generated curriculum doesn’t carry a corpus’s bias, and it doesn’t carry its coverage either. Nobody can yet audit what gets left out.
The number
0 examples of human data at the procedure’s starting point.
Context
The debate over the end of training data had been playing out along two paths: buying licenses and generating synthetic data from already-trained models. This paper tests a third, which doesn’t depend on a prior teacher model but on a search over the space of computable structures, with Solomonoff induction as the theoretical reference.
What’s next
- The paper has been available on arXiv since September 24. No timelines announced for scaling up the procedure or for a peer-reviewed version.
Bottom line
Pretraining with no data sounds like a lab trick, and it probably will be for several years. What isn’t a trick is the curve: if it improves with compute rather than with archives, it changes the list of who can compete.
Sources
- “Self-Play Pretraining with Zero Data”, arXiv 2609.30063
- Our previous coverage: Snorkel’s $350 million round for training data, U.S. restrictions on model distillation and Mistral’s free plan that trains on conversations.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


