← Perspectives/The Category

Where will new knowledge come from when the models have read everything?

Models trained increasingly on model-shaped output converge, and the research now calls the endpoint knowledge collapse. What thins out is not the average but the tail, and the tail is where anything new was always sitting.

By Hunome · 7 min read

There is a question sitting underneath the current enthusiasm that almost nobody in the industry is asking out loud. If models are trained on what humanity has written, and what humanity writes is increasingly produced with the help of models, where does the next genuinely new thing come from?

The technical literature has started to answer it, and the answer has a name: knowledge collapse. Train on a corpus, generate text, let that text become part of the next corpus, repeat, and the distribution narrows. Not immediately, and not in a way that shows up as obviously worse output. The average holds. What thins out is everything at the edges.

What actually collapses

It is worth being precise, because the popular version of this argument is wrong in an important way. Models do not get stupider. On most measures they continue to improve, and the outputs stay fluent. What degrades is variance: the range of ways a thing can be framed, the unusual formulation, the position held by very few people for reasons that turn out to be good.

Work published this year makes the mechanism concrete: epistemic diversity across a set of models is itself protective, and homogeneity across them is what creates the systemic risk. Related research on human expression finds the effect running in the other direction too: sustained exposure to model-shaped language measurably narrows how people themselves write and frame ideas. The convergence is not confined to the machines.

This is the same phenomenon that shows up inside an organisation when everyone consults the same assistant and arrives independently at the same reasonable position. Here it is running at the scale of the written record.

Structured human deliberation is the only reliable process for generating new knowledge.

The record was never the whole of what we know

There is a prior assumption worth surfacing, because the anxiety about running out of training data depends on it: that the documented is a reasonable proxy for the known.

It is not, and this was well understood long before there was anything to train. Polanyi's point was that we know more than we can tell, that competence rests on tacit knowledge which resists articulation not because nobody has got round to writing it down but because it is not the kind of thing that reduces to statement. Hayek's was that the knowledge a society runs on is irreducibly dispersed, held in particular circumstances of time and place by people who cannot transmit it upward without destroying most of what makes it useful.

Both arguments have the same consequence here. The corpus is a projection of human understanding onto the subset that happened to be written, by the people who were in a position to write, in the forms that were publishable. A model trained on all of it has read a large and systematically biased sample. What it has not read is not a remainder to be mopped up with better scraping. It is a different category.

A study shows that AI-simulated diversity always remains below the contributory diversity of human groups, so human diversity remains a valuable creative resource that current AI cannot simulate or sustain. The design of human-AI collaborative workflows determines whether it survives.

A model trained on everything ever written still has not encountered the most consequential thing about your organisation, because nobody has written it down, and nobody could.

Where new knowledge is actually made

If the frontier is not in the record, it has to be produced. And the production process is not mysterious. It is what happens when people who know differently are brought into genuine contact with each other over something that matters.

A researcher hears a practitioner describe what actually happens on the ground and revises a model that had been holding for years. Someone whose objection is a values position rather than an empirical one forces a reframe that no amount of data would have suggested. A pattern forms across contributions from people who had never spoken and would never have met. In each case what results was in nobody's opening position. It is not retrieved, recombined or synthesised from prior material. It comes into existence in the contact.

This is why every Spark on Hunome carries its knowtype: the contributor's own account of how they know what they are contributing, whether that is research, expert fact, lived experience, values or gut feel. Two people can write nearly the same sentence from completely different ground, and holding that difference is what makes the contact productive rather than merely agreeable. A SparkMap preserves the perspectives, their grounds and the connections between them, which is what allows understanding to accumulate instead of resetting.

Why this changes what deliberation is for

Deliberation has usually been justified on procedural terms: it produces buy-in, it is more legitimate, people implement what they helped shape. All true, and all somewhat beside the point now.

The stronger claim available today is that structured human deliberation is one of the few reliable processes for generating knowledge that is not already in the corpus. That reframes it from a nicer way of running a meeting into something closer to a supply line for the genuinely new, at a moment when the dominant technology is extraordinarily good at working with existing material and structurally incapable of producing more of it.

For an organisation the practical consequence is immediate. Everything your competitors can get from a model, you can also get from a model, and quickly. The only understanding that is actually yours is the understanding your people build together about circumstances nobody has documented. That was always true. What has changed is that the alternative, competing on access to the written record, has just been commoditised.

The question to hold

Models will keep improving at what they do, and what they do is valuable. The question this piece is about is not whether that continues. It is what happens to the supply of new material they depend on. What is not finite is human understanding in formation: the collective sensemaking that occurs when people who know differently work on something none of them can resolve alone. That has no operating system at scale, and it is the input everything else now runs on.