← Perspectives/The Evidence

Every AI-simulated group fell below every human group. Bad deliberation design destroys your innovation.

A 2026 study of 644 writers found that no AI-simulated pool of perspectives reached the diversity of a real human pool. It also found that using AI to start the thinking destroyed the diversity that was there, while using it to sharpen finished thinking preserved it.

By Dominique Jaurola · 7 min read

If you have sat in a product meeting in the last year, you have probably heard someone suggest simulating the panel. Generate twenty personas, give each one a background, ask them what they think. It is fast, it is cheap, and the output reads well.

Mengchen Dong and Hiromu Yakura, at the Max Planck Institute for Human Development in Berlin, ran the comparison properly. Their preregistered study, posted as a preprint in July 2026, involved 644 writers and 351 independent evaluators. It has not yet been peer reviewed, so treat the numbers as strong early evidence rather than settled fact. The design is unusually careful and the results are worth knowing now.

The writers completed metaphors. Each was given a scenario and asked to finish two sentences: it feels like something, and this is because something. Simple, open, and impossible to answer well without bringing something of your own.

Three conditions. People working alone. People given AI-generated ideas to start from. People writing first and using AI to refine what they had written.

The simulation never reached the floor

The first result is the one that should end the persona conversation.

Across three different model families, every AI-simulated pool of perspectives fell below every human pool on collective diversity. Not on average. Every one below every one.

The strongest simulation, using Claude Sonnet 4.6, recovered 75% of the diversity in the human-only baseline. The obvious fix is to turn up the randomness, and the researchers tried it. Raising the temperature did increase diversity, but only by degrading the text. At temperature 2.0, 14% of the output was gibberish.

That is worth sitting with, because it tells you what kind of thing is missing. The gap is not a tuning problem. A model asked to be more varied can only move further from the centre of what it learned, and past a certain distance further from the centre is just noise. Human variety is not distance from a centre. It is people standing in different places to begin with.

The people you already employ are the resource

The second result is considerably less comfortable.

The study included 318 writers working in their first language and 326 working in a second language, across 34 languages between them. Without any AI in the process, the second-language writers produced significantly more collectively diverse pools than the first-language writers. The difference was small in absolute terms and statistically clear.

It went further. When second-language writers were allowed to work in their own language instead, collective diversity improved again, rising steadily from first-language English speakers to second-language speakers writing in English to second-language speakers writing natively.

Read that as an organisational fact rather than a linguistic one. A multinational workforce is holding variety that an English-only working culture is already suppressing, before anyone introduces a model. The people in your organisation who think in another language are contributing less of what they have, and the loss is not in their competence. It is in the medium.

Where AI sits in the workflow decided whether diversity survived

The third result is the one to act on. The only variable was whether the machine went first.

When AI was used to generate the starting ideas, collective diversity fell for both groups, and the second-language advantage disappeared entirely. The gap between the two groups became statistically indistinguishable.

When AI was used to refine what people had already written themselves, the advantage survived.

And here is the trap. AI ideation raised individual ratings of divergent thinking for first-language writers. The individuals in the compressed condition rated their own thinking as more creative.

So the workflow that narrowed the group made the people in it feel more inventive. Nothing in the experience of using it signals the loss. That is exactly why it spreads.

To see where AI sits in a real thinking process in your organisation, talk to us.

Most enterprise deployments are the ideation condition

Look at how AI is typically introduced to a group process.

A workshop opens with a generated list of options. A strategy offsite starts from an AI-drafted set of scenarios. A research brief begins with a synthesis of the field. A brainstorm starts with fifty prompts on a board. A consultation begins by summarising responses.

Every one of those is the ideation condition. The machine goes first, everyone reads the same starting material, and thinking begins from a shared anchor. It is faster and it feels more organised. According to this study, it is also the arrangement that removes the variety you convened the group to get.

The refinement condition looks slower and is structurally different. People contribute first, from where they actually stand, before seeing anyone else's framing and before seeing a machine's. Only then does the technology do anything, and what it does is work on the material rather than produce it.

Difference has to be kept, not just collected

Getting people to contribute first is necessary and not sufficient. Once the contributions exist, something has to stop them being flattened.

Two people can write nearly the same sentence from completely different ground. One is reporting a figure from a report. The other has spent a decade with the customers the figure describes. Stored as text, those two contributions are near-duplicates and any summariser will merge them. Stored with their ground attached, they are two different pieces of evidence and the difference between them is informative.

This is why every Spark on Hunome carries its knowtype, the contributor's own account of how they know what they are contributing, whether that is research, expert fact, lived experience, values or gut feel. It is one dimension among several. Contributions are also characterised by how they relate to other contributions, by the human context they come from, by what they signal about change. That combination is what makes a SparkMap hold difference rather than average it.

The Lens is how a SparkMap’s Insights page showcases what is there: where meaning is clustering, which chains of building produced something new, where different ways of knowing are showing up across the map. AI is doing real work in that reading. It is not doing the thinking, and it is not going first.

What the finding actually licenses

The study does not show that AI reduces creativity. It shows that on this task, with these models, simulated variety stayed below human variety, and that whether real human variety survived depended on the order of operations.

That is enough. It means the diversity in your organisation is a real and finite resource that cannot be regenerated synthetically. It means the workflows you deploy are either protecting it or spending it. And it means the difference between the two is invisible to the people inside them, because the compressing workflow is the one that feels more productive.

Human difference is the input nobody else can copy. Design decides whether you still have it in a year.