The usual question about AI in a decision process is whether the model is good enough. It is the wrong question, and there is now a formal argument for why.
Yewon Byun and Bryan Wilder, at Carnegie Mellon, presented Robust Human-AI Complementarity under Uncertainty at ICML in 2026. It is a peer-reviewed conference paper and its starting point is deliberately narrow: the hard problem is not the model's accuracy, it is that the person using it does not know the model's accuracy on this case.
That sounds like a small distinction. It turns out to decide everything.
The obstacle is asymmetric information about quality
The paper opens on the problem: "Machine learning models are often intended to augment rather than replace human decision makers, by providing information that is complementary to human judgement. Yet, in practice, human decision makers routinely fail to realize such complementary gains."
The key issue is not just whether the AI is accurate, but whether it makes different mistakes from the human. A person deciding whether to go with a recommendation is not choosing between a known-good machine and their own judgement. They are choosing under uncertainty about how good the machine is right now, on this input.
Byun and Wilder show that this uncertainty makes it much harder to construct a decision rule that guarantees any benefit at all. Not harder to optimise. Harder to guarantee you have not made things worse.
This is worth pausing on, because it explains a familiar organisational experience. Teams introduce an assistant, the outputs look good, usage climbs, and nobody can demonstrate that decisions improved. That is not a measurement failure. Under uncertainty about how good the machine's output is, an arrangement can consume real effort and produce no reliable gain, and there is nothing in day-to-day use that would reveal it.
A decision maker has to decide under uncertainty. When they understand and have corroboration from all those who know better, their decision making improves. What the paper shows is that a single source of advice, however fluent, does not reduce that uncertainty in any way you can rely on.
What does reduce it is understanding the question from the dimensions that different people actually see, and being able to check where their reasoning meets and where it parts. The decision is still one person's to make. The understanding behind it does not have to be. Most organisations run it the other way round: the people who know wait for the decision to come down, and then explain what it missed. Hunome puts that knowledge in front of the decision instead of behind it.
Error correlation decides whether complementarity is possible
The paper's central result concerns the relationship between the two error patterns.
When the model's errors are negatively correlated with the human's, meaning the machine tends to fail where the person is right and the person tends to fail where the machine is right, robust strategies exist. Under uncertainty, you can construct rules that guarantee improvements in expected utility.
When the errors are positively correlated, meaning both tend to fail on the same cases, the analysis points the other way. The rational response is a movement towards automation rather than complementarity.
That second sentence deserves to be read without flinching. If your people and your systems fail on the same problems, you do not have a team that catches each other's mistakes. You have two things that agree, including when they are wrong, and the human step is adding cost and delay without adding safety.
Two judgements that fail together are not two judgements.
The empirical finding is the uncomfortable one
Byun and Wilder tested this on real-world forecasting tasks. They found that current language models often make errors that are positively correlated with human errors.
They also tried the obvious remedy. Alternative prompting strategies did not fix it. These results suggest that AI systems meant to support human judgement should be trained and evaluated not only for accuracy, but for complementarity: providing information that humans are likely to miss.
Be careful with how far that generalises. It is a finding about the forecasting benchmarks used in that paper, not a universal law about every task in every organisation. But forecasting is not a marginal case. It is what strategy, planning, risk and demand work all consist of. And there is no particular reason to expect a favourable answer elsewhere by default.
Why correlation is rising rather than falling
The mechanism that produces positive correlation is not mysterious, and organisations are actively strengthening it.
Errors correlate when two judges share information and share reasoning. The more a person's picture of a situation comes from the same public material the model was trained on, and the more their reasoning follows the same well-worn path, the more their failures line up.
Now consider what has happened in most organisations over the last two years. Everyone has access to the same class of assistant. People consult it early, before forming a view. The framing they end up with is the model's framing. Positions arrived at independently start converging, and the convergence feels like corroboration.
The result is an organisation raising the correlation between its own judgement and the machine's, while believing it is building a check.
If you want to find out whether your people are currently adding a different view or a confirming one, talk to us.
Negative correlation has to be built on purpose
The paper is clear about where favourable correlation comes from. Different information, or different reasoning processes. That is the condition, and it is not something you get by hiring well or by writing a policy.
Different information means the person holds something the system does not. Not a different opinion about the same material. Actual different material: what happened in a meeting nobody minuted, what a customer said that was never logged, why a previous attempt failed in a way the post-mortem did not name, what a regulator's tone signalled.
Different reasoning means the person arrived at a view through a route the model does not take. Reasoning from values. Reasoning from a stake in the outcome. Reasoning from a decade of pattern recognition in a domain that has never been written down properly.
Both of those are perishable. Unique information stops being unique the moment the person reads the model's summary first. Independent reasoning stops being independent the moment the answer arrives before the thinking. So the sequence is not a preference. It is the whole mechanism.
What an organisation would actually have to do
Take the finding seriously and a specific set of requirements follows.
People have to form and record a position before they see a machine's framing. Not to prove a point about human worth, but because a position formed afterwards is correlated with the thing it was supposed to check.
The ground of each position has to be captured. Does this person know it from research, from expert fact, from lived experience, from values or from gut-feel? That is exactly what you need in order to work out whether their errors would line up with the model's. Two people agreeing from different grounds is meaningful. Two people agreeing because both read the same summary is not.
Positions have to be able to act on each other, so that a genuinely different view can change someone else's rather than sitting in a list of comments.
Disagreement has to be preserved rather than resolved. The cases where the human and the machine diverge are the only places complementarity can be created, and a process that smooths them into consensus has deleted its own value.
And the whole thing has to accumulate, so that the organisation can see over time where its people have been right against the machine and where they have not.
This is what Hunome is built to run. Each Spark carries characterisations, for example its knowtype, the contributor's own account of how they know what they are contributing. A SparkMap holds the perspectives, their grounds and the connections between them, so independent contribution stays visible as independent rather than being averaged into the consensus. The Lens is how a SparkMap’s insight page reads what is there: where meaning clusters, which chains of building produced something new, where different ways of knowing appear across the map. What decision makers receive is deliberative intelligence, including the places where human understanding and machine output do not agree.
The honest version of the choice
Most organisations are talking about human-AI teams as though the arrangement produces value automatically. This work says it does not.
Where errors correlate positively, keeping a person in the loop is ceremony. It costs time, it produces a feeling of oversight, and it does not improve the outcome. It builds a weak, wrong base for robust ‘what next’ decisions. The intellectually honest options in that regime are to automate the task or to change the conditions.
Changing the conditions means deliberately building and keeping the context that is not in the corpus, and protecting the independence of the reasoning that runs on it. That is not a side benefit of collective sensemaking. It is the specific thing that moves an organisation from the regime where the machine should decide alone to the regime where people and machines together are genuinely better than either.
The difference between those two regimes is not how good your models are. It is where your people know something that is not written and understood already.
