← The Index
008
Field Note · 008
The loop tightening.
On synthetic data and what happens when models train on what models made. Spring 2026.
By a conservative 2026 estimate, more than half of the new text appearing on the public web each day is machine-generated. Image and audio shares are climbing. The training data of the next generation of models will therefore be heavily, possibly mostly, composed of the output of the previous generation. This is no longer hypothetical.
01 /
The published findings on model collapse are clear in direction, less clear in magnitude. Models trained on their own output, recursively, lose the long tail first. Rare phrasings, unusual word pairings, the specific texture of niche human writing. These are the first to go. The center holds. The edges fade.
02 /
The internet used to be a mirror of human attention. It was distorted and uneven, but originated in people. It is becoming a mirror of itself. The loop tightens by one generation a year.
03 /
The cultural consequence is downstream of the technical one. We are training the next round of models on a world that already has the previous round in it. The world we are recording is no longer fully the one we are in.
Synthetic data is not the problem. The problem is older. Every recording technology, eventually, becomes the thing it was recording. Photography did this to memory. The web is doing it now, to language.