Companies can buy the same generative AI tools, give staff similar training and still end up with radically different results. A new MIT Sloan study argues that the difference may lie in the hidden work behind AI implementation: the repeated testing, checking, persuading and redesign that employees do after the software arrives.
The research compared AI innovation inside an academic medical centre and a corporate law firm over two years. Both organisations gave employees secure environments for experimenting with generative AI and encouraged them to build useful applications. One ultimately had 141 organisation-wide AI solutions in use. At the other, more than 80% of the domain experts involved in AI innovation dropped out, leaving only three organisation-wide solutions.
AI implementation is not finished when staff get access
The familiar enterprise AI story begins with procurement. A company chooses a model, negotiates security requirements, makes the tool available and trains employees to use it. That looks like deployment. The MIT researchers argue that for organisation-wide innovation it is closer to the starting line.
Employees then have to discover which tasks the system can reliably perform, test failure cases, refine prompts or workflows, review outputs with colleagues, integrate tools with existing processes and keep adapting as the underlying models change. The researchers describe this as collective experimentalist work. Much of it happens alongside people’s ordinary jobs.
The MIT Sloan working-paper catalogue identifies the research as “Experimentalist Intensification Governance,” a qualitative field study by Arvind Karunakaran, Katherine C. Kellogg and Batia Mishan Wiesenfeld. Because it is a working paper based on two organisations, the results should not be treated as a universal formula for every employer. Its value is in showing what the implementation process can look like at close range.
The hidden work behind AI can become a second job
At the law firm studied, employees were initially willing to experiment. The problem was persistence. As the extra burden accumulated, participants gradually stopped doing the trial-and-error work needed to turn individual experiments into reliable shared tools. The study describes withdrawal rather than open resistance.
That distinction matters for managers. A failed AI programme may not produce a dramatic revolt or a clear technical incident. It can simply lose momentum. Fewer experts test ideas. Fewer people review outputs. The most knowledgeable users decide they have more urgent work. Eventually the organisation still owns the software, but the learning process around it has thinned out.
LiveAIWire has reported on workers spending their own money on AI tools. That behaviour shows how strongly employees can value useful AI. The MIT research adds a different warning: enthusiasm does not mean people have unlimited capacity to perform unpaid or unrecognised implementation work on top of their normal responsibilities.
This may help explain why adoption numbers can be misleading. A company can count licences, logins or experimental projects and conclude that AI is spreading. Those metrics do not reveal whether the experiments are being converted into stable processes that colleagues trust enough to use.
The healthcare organisation built support around experimentation
The academic medical centre in the study followed a different path. Researchers say it created continuing education, shared evaluation practices, technical support for higher-priority projects and formal risk screening. It also made experimentation visible in job responsibilities, performance evaluation and career recognition.
The result was not merely that more people tried AI. More people stayed involved long enough for ideas to be reviewed, improved and turned into organisation-wide applications. The researchers report 141 solutions in use, with more under development, compared with three at the law firm.
Those numbers are striking, but they are not a controlled experiment in which every other organisational feature was held constant. The researchers say they did not identify meaningful differences in technology access, AI readiness, regulatory constraints or the suitability of the use cases that explained the gap. Even so, a two-case comparison cannot prove that the support structures alone caused all of the difference.
What it can do is make the missing labour visible. Testing an AI-generated discharge summary with real clinical requirements or validating legal research across colleagues requires domain expertise. If that effort is treated as a hobby for keen staff, the organisation may consume its most valuable human input without budgeting for it.
AI does not remove coordination work, it can create more of it
The popular image of automation is that software removes tasks. In practice, new systems often create new coordination around the tasks they automate. Somebody has to decide what good output looks like, who is accountable for errors, when human review is mandatory and how a changed model should be retested.
Generative AI makes that especially demanding because the system is flexible. The same model can be used for dozens of jobs, and its output can vary with instructions, context and model updates. That flexibility is valuable precisely because it makes the implementation less like installing a fixed machine and more like developing an evolving process.
LiveAIWire has examined evidence that human-AI teams can outperform people or AI working alone in some settings. The hidden-work study helps explain one cost of that partnership: someone has to design and maintain the collaboration, and that effort does not disappear once the first successful demonstration is complete.
The same problem appears when people use AI to save time. A tool may reduce the minutes spent drafting a document while increasing the need for review, exception handling or cross-team agreement. Productivity gains are real only if the whole workflow becomes easier, not merely the visible generation step.
Managers may need to budget for experimentation like real work
The practical implication is uncomfortable because it challenges the idea that employee experimentation is free. Organisations often encourage staff to “play with AI” in spare moments, hoping useful applications will naturally emerge. That can work for discovering possibilities. Scaling those possibilities requires a more deliberate investment.
Dedicated technical help can reduce the burden on domain experts. Shared evaluation rubrics can stop every team reinventing its own standards. Documentation can preserve lessons when employees move roles. Recognition can signal that experimentation is part of the job rather than an invisible extra. None guarantees success, but each changes the cost of staying engaged.
There is also a governance benefit. When experimentation is formalised, risk review can happen before a tool quietly spreads through an organisation. A supported team is more likely to document what data a model receives, where outputs are checked and what happens when the provider changes the model.
That is especially important because AI systems can weaken skills if workers hand over too much of a task. LiveAIWire has covered research on AI assistance and weaker human skills. Good implementation therefore needs to ask not only whether a workflow is faster, but which human expertise must remain strong enough to supervise it.
The quiet failure mode may be the most common one
The MIT study is memorable because its failure case was not a spectacular technical collapse. People simply stopped contributing. That is a plausible risk in many organisations where innovation depends on a relatively small group of motivated experts doing extra work that leaders cannot easily see.
If the paper’s interpretation holds more broadly, buying better AI models will not solve that organisational problem. Faster, cheaper systems may actually produce more experiments and therefore more demand for evaluation, integration and governance. The hidden workload could grow as the technology improves.
Enterprise AI is often discussed as a choice between adopting quickly and falling behind. The more useful question may be whether an organisation has created the conditions for employees to keep doing the difficult work after the excitement of adoption fades. A licence can put AI on a desk. It takes sustained human effort to turn it into infrastructure.
There is a useful accounting implication too. If organisations want to know whether AI pays, they should count the human implementation time as part of the investment. Hours spent testing, documenting, reviewing and coordinating are not incidental just because they do not appear on a software invoice. Treating them as free can make an AI project look cheaper during the pilot and surprisingly expensive when it has to scale.
That also gives finance teams a better way to compare pilots. A cheap tool that consumes large amounts of expert review may be less attractive than a more expensive system that fits existing controls and reduces coordination. The full workflow, rather than the subscription price, is the unit that matters.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
