AI & Work

AI Assistance Raised Performance but Hid Weaker Human Skills

A friendly white AI robot assists monkeys using laptops in a cluttered office.
AI can improve a team’s output while masking the underlying skills people may still need to build.

AI assistance helped people perform better while it was available, but it also made their later unaided ability look stronger than it really was. In a controlled logic-puzzle experiment, participants could ask a perfectly accurate simulated AI to reveal part of an answer. The help improved progress in the assisted phase, yet the resulting performance became a less reliable guide to how well those people could solve new problems alone. The clearest predictor of skill growth was not how rarely someone called the AI. It was how much of the reasoning they still did independently.

That distinction matters anywhere performance is judged while an AI tool is switched on. A student can submit better work without mastering the underlying method. An employee can finish more tasks without becoming more capable of handling the next one unaided. A manager can see stronger output while losing sight of which skills belong to the person and which are being supplied by the system.

How the AI assistance experiment worked

The study by researchers at the University of California, Irvine used ordering puzzles designed to test logical reasoning. Each problem presented six objects and five constraints. Participants had to place every object in the correct order, with some recurring visual features offering a hidden structure that could be learned across problems.

The researchers recruited 150 English-speaking adults in the United States through Prolific. After exclusions, the final sample contained 124 people, all of whom reported at least an undergraduate degree. The experiment had three phases. Participants first worked without AI, then entered a longer learning phase under one of three conditions, and finally completed another unaided phase.

One group received no AI during the learning phase. Two groups could request help from a simulated assistant that revealed the correct position of one randomly selected object. The assistant was deliberately made 100 per cent accurate, although participants were not told that. Requests carried either a low or a high points cost, allowing the researchers to test whether friction changed usage.

The low-cost group averaged 6.67 requests, compared with 3.33 in the high-cost group. That difference provided some evidence that even a small penalty could change reliance. More importantly, the design let the researchers compare what people achieved with help against what they could later do when the help disappeared.

AI-assisted success overstated later human performance

Performance during the learning phase did not mean the same thing for everyone. For participants who used AI, their assisted results tended to overpredict their unaided reward rate in the final phase by 0.22 units. For those who did not use it, the learning-phase measure underpredicted later performance by 0.15 units.

In practical terms, the tool made assisted users look more capable than their next solo performance justified. It was not simply adding speed to a stable level of human skill. It was changing the relationship between visible output and the ability underneath it.

The final accuracy scores did not differ significantly across the three experimental conditions. The more revealing difference appeared in time. People in the low-cost condition took an average of 111.81 seconds per final-phase problem, compared with 86.78 seconds in the high-cost condition and 92.21 seconds in the no-AI condition. That suggests the group given the easiest access to help reached broadly similar answers but needed longer when returned to independent work.

This is a more precise warning than the claim that AI automatically destroys skills. A previous study of 26,000 students found that AI could improve homework performance while weakening exam results. The new experiment isolates a related measurement problem: assisted output can conceal how much learning has actually taken place.

The strongest signal was independent reasoning

The researchers calculated each participant’s “solo share”, the proportion of learning-phase problem-solving time spent reasoning before or without an AI request. After accounting for initial ability, a higher solo share was positively associated with greater latent skill growth.

In the authors’ model, one additional minute of independent reasoning was linked to a 1.5 per cent increase in predicted final-phase accuracy and a 2.5 per cent reduction in predicted response time. The estimate does not prove that every extra minute caused those gains, because solo share was observed rather than randomly assigned. It does show that preserved human effort explained later unaided performance better than a simple count of AI requests.

Participants who never requested help improved their average reward rate from 2.03 to 3.86, an increase of 90.2 per cent. Across the assigned conditions, the average rise from the first to the final phase was 1.95 in the no-AI group, 1.66 in the high-cost group and 1.32 in the low-cost group. Evidence for some of those differences was weak, so the ordering should not be treated as a definitive ranking. The consistent message is narrower: the amount of independent reasoning remained important even when the assistant was perfectly reliable.

That helps explain why human-AI teams can outperform people working alone without becoming the best-performing arrangement in every setting. A tool can improve the immediate product while changing the practice through which a person would otherwise build skill.

Request frequency was not the whole problem

The number of AI requests was not significantly associated with latent skill change after the researchers controlled for starting ability and independent effort. That is an important result for schools and employers tempted to set crude usage limits.

Two people might each ask for five hints but use them very differently. One may struggle with the problem, form a hypothesis and request a targeted clue. The other may ask immediately and let the system remove the difficult part. A request counter treats those behaviours as identical even though the learning experience is not.

The study therefore points towards better measures of AI-supported work. Organisations need to examine when help arrives, what cognitive work remains with the person and whether users can explain or reproduce the result. The relevant question is not only “How much AI did you use?” It is “Which parts could you still perform if the tool were unavailable?”

This also changes how productivity claims should be read. In one field trial, experienced developers believed AI had accelerated their coding even when measured completion times increased. Both studies show why impressions and visible output can be poor substitutes for carefully separated measures of human and machine contribution.

What this means for learning and work

AI systems do not need to be banned from training environments. They need to be placed so that assistance does not erase the useful struggle that builds transferable knowledge. A tutor could require a learner to attempt a solution before revealing a hint. A workplace system could ask for a short plan or diagnosis before generating the finished answer. Assessments could include an unaided stage so that managers and teachers can distinguish supported performance from retained capability.

The same principle applies to the design of entry-level work. Junior tasks can appear inefficient once an AI can complete them quickly, but some of those tasks are how professionals learn to notice exceptions, test assumptions and build judgement. LiveAIWire has previously examined how automation can threaten the work through which professional expertise is formed. Removing every routine step may lift short-term throughput while narrowing the route to senior competence.

Employers should also separate production metrics from development metrics. Output, speed and quality describe what a person-tool system delivered. They do not automatically describe what the person learned. If promotion, staffing or safety decisions depend on independent competence, that competence needs its own test.

Why the study does not prove permanent deskilling

The experiment was short, controlled and based on one type of logic puzzle. Its AI was a simulated hint system that never made mistakes, unlike real assistants that can be inconsistent or confidently wrong. The sample was relatively small, highly educated and limited to adults in the United States. The research measured immediate skill change, not whether differences persisted for weeks or transferred to other tasks.

The study’s Bayesian analysis also identified an association between independent effort and skill growth, not a causal effect created by randomising solo share. Some comparisons between the low-cost and no-AI groups met only a relaxed one-sided threshold. Those details make the findings useful but not universal.

The authors describe the work as a conference-accepted manuscript, and its public record provides the version history and submission details. The publisher DOI identifies the accepted HCOMP 2026 paper. Further studies will need to test longer learning periods, fallible generative systems and real educational or professional tasks.

AI should support the struggle, not erase it

The most useful lesson is not that every AI request is harmful. It is that performance with assistance and skill without assistance are different outcomes. When an organisation measures only the first, it can reward apparent competence while overlooking a weakening human foundation.

Good AI support should leave users with more than a finished answer. It should preserve enough diagnosis, recall, judgement and practice for the person to improve. The new research suggests that the human share of the reasoning is not wasted time. It may be the part that makes today’s help become tomorrow’s skill.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.