AI & Work

Working With AI Was Better Than Working Alone, but Not Best

Split-screen illustration showing a man working successfully with an AI assistant and manager beside a celebrating collaborative team.
AI can improve individual work, but the strongest results often come from people working together.

Human-AI teamwork improved performance compared with people working alone, but it still failed to beat the stronger partner on average. That distinction, drawn from a preregistered meta-analysis of 106 experiments, challenges the assumption that combining human judgement with artificial intelligence automatically produces the best result.

The peer-reviewed study in Nature Human Behaviour analysed 370 effect sizes from experiments that measured three conditions: a human working alone, an AI system working alone, and the two working together. The combined systems helped humans substantially, yet they performed significantly worse than whichever of the human or AI was already better at the task.

Human-AI Teamwork Passed One Test but Failed Another

The researchers made an important distinction between augmentation and synergy. Augmentation means the human-AI combination outperforms the human alone. Synergy sets a higher standard: the combination must outperform both the human and the AI working separately.

Across the studies, the average augmentation effect was positive and substantial, with a Hedges’ g of 0.64. The average synergy effect was negative, at -0.23. Put simply, AI assistance usually made people better, but joining the two partners did not usually create the strongest available performer.

This is not a semantic difference. A business comparing an AI-assisted employee with an unassisted employee may conclude that its deployment has succeeded. If the AI system could have completed the same bounded task more accurately on its own, the combined workflow may still be leaving performance on the table.

The analysis included experiments published between January 2020 and June 2023. To qualify, a study had to report quantitative performance for all three conditions. That requirement is what allowed the authors to test the stronger claim of synergy rather than the more common question of whether AI helps a person.

The paper’s supplementary methods show that 1,903 database records were screened before forward, backward and review-based searches expanded the pool. Seventy-four papers containing 106 qualifying experiments were ultimately included.

The Stronger Partner Changed the Direction of the Result

The clearest dividing line was which partner performed better alone. When humans were stronger than the AI, the combined system outperformed both, producing a positive synergy effect of 0.46. When the AI was stronger, adding human input created a negative effect of -0.54 relative to the AI alone.

That reversal suggests people were not equally good at adding value in both situations. When humans possessed the better overall judgement, they may also have been better placed to decide when to accept the machine’s recommendation and when to rely on their own view. When the AI was better, the human often became a source of error.

More than 95 per cent of the combined systems left the final decision with a human after showing them an AI output. That familiar design sounds responsible, but it can produce a weak division of labour. The person may override a correct recommendation, defer to an incorrect one, or blend two answers without knowing which partner has the relevant advantage.

The finding does not make human oversight pointless. In high-stakes work, overall accuracy is only one consideration. A human may be needed for legal accountability, values, consent, rare harms or escalation. The study instead shows that adding a final human decision is not itself evidence of better performance.

Decision Tasks Were Harder Than Creative Work

Task type also mattered. For decisions with a finite set of possible answers, the combined systems showed a significantly negative synergy effect of -0.27. Creation tasks, where participants produced open-ended content, had a positive point estimate of 0.19, although that result was not statistically distinguishable from zero.

The difference between the two task groups was statistically significant. Creative work may offer more room for complementary contributions because quality can emerge from iteration, variety and revision. A person can redirect a draft, combine suggestions or recognise an unusual idea even when neither partner would have produced the finished result alone.

Decision work creates a sharper coordination problem. If the task has one correct classification or choice, someone must decide which answer to trust. Explanations and confidence scores do not automatically solve that problem. The meta-analysis found no statistically significant average moderation from AI explanations or reported confidence.

This helps place LiveAIWire’s report on an AI coding trial that added time in a wider context. An assistant can generate useful material while the surrounding process of checking, rejecting and integrating its output consumes the expected gain. Performance belongs to the whole workflow, not the most impressive moment inside it.

What This Means for Your Work

For workers, the practical question is not simply whether AI helps. It is which parts of a task each partner performs better, and whether the workflow sends those parts to the right place. Using AI for a first draft, search or comparison may be valuable even when handing it the final decision would not be.

For managers, adoption metrics are especially weak evidence. A high number of prompts, licences or active users does not reveal whether the combined system beats the best available alternative. Teams need comparable tests of human-only, AI-only and combined performance on representative work, including quality, time, rework and serious errors.

That approach may expose uncomfortable results. Some tasks will still require a human even when the AI scores better, because accountability or customer trust cannot be delegated. Others may not benefit from a human reviewing every routine output. The useful outcome is not maximum automation, but an intentional explanation for where each partner enters and why.

LiveAIWire’s analysis of the AI workplace divide showed how contextual judgement is becoming more valuable as routine processing is automated. The meta-analysis adds a qualification: human judgement must be relevant to the task and placed where it can improve the result. A human included only as a ceremonial approver may add delay without adding insight.

Why More Oversight Can Produce Less Control

Organisations often respond to AI risk by adding review stages. That can be sensible, but oversight becomes fragile when reviewers lack time, expertise or a clear reason to distrust the system. The process creates the appearance of control while encouraging people to click through recommendations they cannot meaningfully assess.

It can also generate cognitive load. A worker who must inspect a constant stream of plausible outputs makes repeated acceptance and rejection decisions. LiveAIWire’s coverage of AI workflow fatigue describes how supervising multiple tools can turn employees into intermediaries between machine output and usable work.

The alternative is capability-aware delegation. A team can identify subtasks where humans have richer context, where AI has reliable speed or pattern recognition, and where uncertainty requires escalation. Only three of the experiments in the meta-analysis tested a predetermined division of separate subtasks, and their small combined evidence base was inconclusive. That makes workflow design a research priority rather than a solved formula.

Better interfaces may help by showing when the model is likely to fail, not merely displaying a general confidence score. Training can focus on recognising those boundaries. Evaluation should also track whether people learn to calibrate reliance over time instead of assuming that a short onboarding session creates effective collaboration.

The Result Is Strong but Not Timeless

The authors reported very high heterogeneity. The experiments varied in tasks, populations, systems and measures, so the pooled average is not a prediction for every workplace. It is evidence against a universal claim that human plus AI is necessarily best.

The search also stopped in June 2023. Current generative models, agentic tools and interfaces can behave differently from the systems covered. More recent tools may produce better collaboration, but they may also introduce new coordination costs. The finding should therefore guide measurement, not freeze expectations around an older generation of technology.

Selection rules create another limitation. Studies that did not report human-only, AI-only and combined conditions were excluded, even if they offered useful evidence about collaboration. The authors also noted possible publication bias and differences in study design. Their public article record links the data and code, but the wide variation still demands caution when moving from an average to a specific deployment.

It is equally important not to interpret the negative synergy result as proof that AI should work without human control. The performance metric may omit rare but severe mistakes, fairness, explanation, legal duties and the cost of a wrong decision. The paper explicitly argues for more robust measures that incorporate completion time, financial cost and practical consequences.

Design the Team Around Strengths, Not Assumptions

The meta-analysis replaces a comforting slogan with a more useful design rule. Human-AI teamwork creates value when the partners contribute different strengths and the process knows how to use them. Putting a person and a model in the same loop is only the beginning.

That principle also applies at an individual level. Someone using an AI assistant should be able to explain what the tool is good at, what they themselves know that it does not, and what evidence would cause them to reject its answer. If none of those boundaries is clear, the collaboration is operating on habit rather than judgement.

The rapid spread of AI tools makes this distinction more urgent. As LiveAIWire has reported, productivity tools can consume the attention they promise to save. A system that improves a narrow output but adds supervision, interruptions and rework may still reduce the performance of the larger team.

Working with AI was better than working alone in the average experiment. It was not best. The next stage of workplace AI should therefore focus less on adding assistants everywhere and more on proving, task by task, that the partnership beats the strongest partner without hiding new costs elsewhere in the process.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.