AI & Work

AI Closed Three Quarters of a Workplace Performance Gap

LiveAIWire AI News and Insights
LiveAIWire covers the latest AI news, developments and insights in artificial intelligence.

AI helped lower-education workers catch up on a workplace task

A randomized experiment offers evidence for a less familiar side of the AI productivity debate: generative AI may narrow performance gaps while people are using it. In the study, 1,174 adults aged 25 to 45 completed a workplace-style problem-solving task with or without an AI assistant. The Stanford SCALE summary of the randomized experiment reports that everyone improved with AI, but participants with less education gained more.

Without AI, the performance gap between higher- and lower-education participants was 0.548 standard deviations. With AI, it fell to 0.139, closing about three quarters of the initial difference. That does not mean education stopped mattering. The gap re-emerged once the tool was removed, although the lower-education group retained some of its improvement.

The strongest result is about assisted performance, not permanent equality

This is where the finding becomes more interesting than a simple claim that AI makes everyone equal. The researchers found that intensive use could raise assisted performance even when users did not exert as much of their own effort. But performance after AI was removed improved most when heavy AI use was combined with sustained human effort. In other words, delegating more work can lift the immediate score, while learning still depends on participation.

That distinction echoes concerns in student dependence and empowerment, where students reported both empowerment and dependence when using AI. It also fits the broader question raised by AI assistance and weaker human skills: assistance can increase output, but people still need opportunities to practise the underlying skill if they are expected to perform without the tool.

Why this matters outside the experiment

Many employers are interested in AI because it can raise the floor of performance. A less experienced employee can get help structuring a document, checking a calculation or exploring alternatives before asking a colleague. The new experiment suggests that this effect may be particularly valuable for people who start with fewer formal educational advantages.

But the evidence does not justify replacing training with a chatbot subscription. The task was online and workplace-like rather than a long period of real employment, and the study followed people through an unassisted module rather than months of skill development. The result is better read as proof that AI can compress an immediate performance gap under controlled conditions.

The design of the work still decides who benefits

A second Stanford review of AI in education stresses that rigorous causal evidence remains much thinner than the volume of commentary around the subject. Stanford review of AI education evidence That is a useful brake on sweeping conclusions. The most promising question is not whether AI is good or bad for skills in general, but which forms of use create scaffolding and which forms create dependency.

Employers can apply the same logic. If the AI produces a final answer that no one inspects, the user may gain speed without understanding. If it produces a draft that the user must critique, revise and defend, the same tool can become part of learning. The difference is managerial design rather than model capability.

A productivity tool can also be a mobility tool

The most optimistic interpretation is that generative AI could reduce some advantages that come from knowing how to frame a task, structure a response or navigate specialist conventions. That would matter in workplaces where formal education still acts as a proxy for readiness even when the real job depends on practical problem-solving.

The caution is that higher-education participants in the experiment also used AI more effectively. Tool skill therefore becomes another form of human capital. Closing a gap in one task does not automatically close differences in judgement, domain knowledge or the ability to decide when the model is wrong. LiveAIWire’s coverage of human-AI teamwork is relevant here because the best results often come from complementary human and AI strengths.

The real test is what happens after the assistant disappears

The follow-up module is the part of the experiment that deserves attention. Lower-education participants did not simply collapse once AI was removed, but a sizeable gap returned. That suggests some learning or transfer, but not enough to claim that the original difference had vanished.

For organisations, that points towards a useful measurement habit. Do not only ask whether employees are faster with AI. Test whether they can explain the result, detect a bad answer and perform key tasks without assistance when necessary. If AI narrows access to competent performance while preserving those abilities, the productivity gain becomes much more valuable.

The important point is that the result should not be read as a universal forecast. The evidence describes a particular setting, population or technical system, and the strongest conclusion is about what happened under those conditions. That distinction matters because AI stories often travel faster than their limitations. A useful reading keeps the headline finding intact while separating it from broader claims that the source did not test.

There is also a practical reason to watch this development. AI products are moving from isolated demonstrations into ordinary workflows, which means small design choices can have large effects once they are repeated across millions of interactions. The next phase will be less about whether a system can perform a task at all and more about reliability, human control, cost, access and what happens when the technology meets messy real-world behaviour.

For readers, the safest takeaway is neither enthusiasm nor dismissal. The evidence is strongest when it is used to identify a real change and weakest when it is stretched into a prediction about everyone. What matters next is replication, wider deployment data and whether the same effect survives outside the original conditions. Those are the tests that turn an interesting result into something people can reasonably use.

The wider pattern across AI is becoming clearer: capability alone is not the whole story. Context determines whether a tool helps, distracts, saves time, shifts power or simply moves effort somewhere else. That is why seemingly narrow findings can matter. They expose the conditions under which AI changes behaviour, and those conditions are often more useful than a single benchmark score or product claim.

The important point is that the result should not be read as a universal forecast. The evidence describes a particular setting, population or technical system, and the strongest conclusion is about what happened under those conditions. That distinction matters because AI stories often travel faster than their limitations. A useful reading keeps the headline finding intact while separating it from broader claims that the source did not test.

There is also a practical reason to watch this development. AI products are moving from isolated demonstrations into ordinary workflows, which means small design choices can have large effects once they are repeated across millions of interactions. The next phase will be less about whether a system can perform a task at all and more about reliability, human control, cost, access and what happens when the technology meets messy real-world behaviour.

For readers, the safest takeaway is neither enthusiasm nor dismissal. The evidence is strongest when it is used to identify a real change and weakest when it is stretched into a prediction about everyone. What matters next is replication, wider deployment data and whether the same effect survives outside the original conditions. Those are the tests that turn an interesting result into something people can reasonably use.

The wider pattern across AI is becoming clearer: capability alone is not the whole story. Context determines whether a tool helps, distracts, saves time, shifts power or simply moves effort somewhere else. That is why seemingly narrow findings can matter. They expose the conditions under which AI changes behaviour, and those conditions are often more useful than a single benchmark score or product claim.

The important point is that the result should not be read as a universal forecast. The evidence describes a particular setting, population or technical system, and the strongest conclusion is about what happened under those conditions. That distinction matters because AI stories often travel faster than their limitations. A useful reading keeps the headline finding intact while separating it from broader claims that the source did not test.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.