Tiny models never shown real data got better at predicting it
Self-Play Pretraining with Zero Data
TL;DR On web text, tiny models trained only on a second AI’s program output scored 5.34 bits per byte (prediction error: lower is better, 8 is blind guessing), against 7.75 with random programs. At small scale, it hints computing power could supply some training data.
The problem
Chatbots today learn from text that people wrote, then chose and filtered for them. The authors ask whether a model could make its own practice material instead, limited only by computing power. The first recipe they tried had a flaw. If you reward a data maker for whatever is hardest to predict, it can cheat by adding random noise. So the data must be hard enough to teach, yet regular enough to learn.
The idea
The team trains two models from scratch that train each other, a setup called self-play. A generator writes programs in a bare-bones language that can, in principle, express any computation; each run emits bytes. A learner practices predicting those bytes. The generator works like a piano teacher paid most when an exercise makes the pupil stretch the way they have been improving. Mastered pieces earn little, and so does random banging, which is hard but teaches no skill. Other learners train on random programs or hand-built grammar data for comparison. The authors argue this reward keeps practice at the edge of what the learner can pick up.
What they showed
- Chosen programs beat random ones. With a 1-million-parameter learner, web text cost 5.34 bits per byte, a measure of surprise at each next byte, against 7.75 when programs were drawn at random. As compute grew, random programs also improved far more slowly.
- It rediscovered textbook math. By training round 512 the generator was writing programs that output Fibonacci, geometric, quadratic and cubic sequences, while random sampling across more than 53,000 rounds’ worth of programs produced none of them.
- Hand-built grammars won on language. Learners trained on grammar-generated data, random rules that build nested, sentence-like patterns, did better on text and code, but self-play clearly beat them on images, music, audio and speech.
Why it matters
Researchers who build synthetic training data are the audience. The authors want training data limited by compute rather than by human knowledge. A learner warmed up by self-play learned faster when later trained on real text, images and audio. That comparison skipped self-play’s own compute, and the gap narrowed late in training. If the speedup holds at larger sizes, self-play could supplement some human-written data.
The catch
This is an early, small-scale test. Every model had under 25 million parameters, so nobody knows whether the trend holds at chatbot scale. Settings were tuned by checking scores on real web text and DNA, so a little real data steered the choices, though it never trained the models. The reward itself was picked after trying several at small scale, which may flatter it in small-scale comparisons.
Who wrote it
Seven researchers, with Aditya Cowsik (an independent researcher), Kfir Dolev (Tel Aviv University) and Michael Y. Li (Stanford University) as equal contributors, plus colleagues at those universities and LAPTh, USMB. Yoav Levine co-authored “In-Context Retrieval-Augmented Language Models” (2023).