Scientific papers have always been passive objects. Researchers read them, track down the code, reconstruct the methods and decide how to apply the work elsewhere. A Stanford-led team has built a system called Paper2Agent that tries to turn that process inside out: the paper becomes an AI agent that can answer questions, run its methods and collaborate with agents built from other papers.
The work, published in Nature, is more ambitious than attaching a chatbot to a PDF. The system ingests the manuscript, supplementary material, data and code, then constructs tools that let an AI agent actually execute parts of the research workflow. In tests, the authors report that paper agents could reproduce published analyses, answer new queries and combine methods across papers.
Paper2Agent is meant to know how the research works
The peer-reviewed Paper2Agent study describes a framework that converts research outputs into resources and executable tools exposed through the Model Context Protocol, or MCP. A chat agent can then call those tools when a user asks a question.
This matters because a conventional paper often leaves a large gap between understanding the text and reproducing the work. A reader may need to find a repository, install dependencies, identify the right dataset, interpret file formats and work out which commands correspond to the method described in the manuscript.
Paper2Agent tries to package that practical knowledge with the paper. The result is closer to a virtual corresponding author than a search box. It can explain the work, invoke its code and apply the method to new inputs.
That is an attractive idea for scientists already dealing with an overwhelming volume of literature. It matters in a research system already facing a growing backlog created by AI-assisted science. Making papers executable could help with the second problem created by faster discovery: understanding and reusing what everyone else is producing.
The system was tested on real research code, not only summaries
The Nature paper reports a large-scale evaluation across computational biology and other computational research. Of 100 computational biology papers sampled from bioRxiv, 74 were successfully converted into agents. The system proposed hundreds of tools and automatically validated nearly all of those it could build.
The failures are revealing. Papers could not always be agentified because code was missing, data or model artefacts were unavailable, software environments broke or scripts were too specific to generalise. In other words, the bottleneck was often the same one human researchers face: a paper can be clear while the computational machinery behind it is difficult to reuse.
On benchmark questions derived from tutorials, the authors report higher accuracy for Paper2Agent than giving Claude Code direct access to the paper repository under the compared configurations. They also report lower query cost and latency in those tests. These are results from the authors’ evaluation, not an independent audit of every scientific domain.
Two paper agents found a connection the authors say was new
The most striking demonstration involved two unrelated pieces of genetics research. One agent represented a method for predicting how genetic variants affect biological activity. Another represented a genome-wide study of attention-deficit/hyperactivity disorder risk.
When the agents worked together, they prioritised a variant near the gene MPHOSPH9 and generated a possible mechanism linking it to ADHD risk. The Nature paper is careful about the status of this result: it is a computationally generated candidate hypothesis that still requires experimental validation.
Stanford’s account of the work says the connection had not previously been reported, according to senior author James Zou. That makes the demonstration exciting, but it should not be mistaken for a validated medical discovery. The important technical point is that agents built from separate papers could combine their tools and data to propose something neither paper contained alone.
This is the deeper promise of the approach. Scientific knowledge is fragmented by publication. One team creates a method, another produces a dataset and a third observes a phenomenon. Human researchers have to notice that the pieces fit. Agents could search for those connections at a much larger scale.
Turning papers into tools could change how methods spread
Research methods often diffuse slowly because implementation is hard. A scientist in another field may understand that a technique could help but lack the software expertise to set it up.
If a paper agent can expose a reliable natural-language interface to the method, the barrier changes. The user can describe the problem, and the agent can call the underlying workflow rather than forcing the scientist to become an expert in the repository first.
That could be particularly valuable across disciplines. A biologist might use a statistical method developed in economics, or a materials scientist might apply an image-analysis technique from astronomy, without manually translating every implementation detail.
Earlier work showed AI agents tackling difficult scientific problems. Paper2Agent attacks a different bottleneck: not inventing a solution from scratch, but making existing scientific machinery easier to invoke and combine.
The paper itself warns that authors still have knowledge the PDF lacks
An important limitation is that manuscripts do not contain everything. Researchers make judgement calls, abandon failed experiments, clean data in ways that may be only partly documented and accumulate tacit knowledge about what tends to break.
Stanford says human authors can therefore add context by conversing with the paper agent. That is a useful admission. If the goal were simply to turn the PDF into a chatbot, the authors would not need to be involved. The system becomes more faithful when the people who did the work can fill gaps that the published record leaves behind.
This also creates a governance problem. If an agent is updated after publication with informal author guidance, users need to know which parts came from the peer-reviewed paper and which came later. Versioning, attribution and provenance become part of scientific communication.
Attribution becomes harder when agents remix research automatically
The Paper2Agent team explicitly says original authors still need credit when agents extend their work. That sounds obvious until hundreds of agents collaborate.
A generated hypothesis might depend on a dataset from one paper, an algorithm from another, a software library from a third and a model created by a fourth team. Traditional citations were designed for humans writing a narrative, not networks of software agents composing workflows in seconds.
The system therefore points towards a future where provenance has to be machine-readable. An agent should be able to say which paper supplied a claim, which code produced a result and which transformation happened after the source material.
That concern is especially relevant in fast-moving areas such as AI-assisted protein research, where a plausible computational output may still be far from laboratory validation.
Millions of paper agents would create a new scientific search problem
Stanford says the team has already created more than 100 paper agents and ultimately imagines a much larger ecosystem. Scaling from hundreds to millions would not simply mean more answers. It would create a problem of deciding which agents should talk to each other.
Most pairs of papers have no useful connection. An efficient system would need to identify promising combinations without wasting enormous amounts of compute on meaningless conversations. It would also need safeguards around dangerous domains and methods that should not be freely chained together.
The researchers themselves emphasise monitoring and safety around agent collaboration. That is crucial because a system designed to make scientific methods executable can amplify both useful and risky capabilities.
The real change is from reading knowledge to running it
Paper2Agent does not make papers obsolete. The paper remains the human-readable record, the place where claims, evidence and limitations are argued in context.
What changes is what can sit beside that record. Instead of downloading a PDF and beginning the long process of reconstructing the work, a researcher could ask the paper to demonstrate its method, apply it to a new dataset or explain why a particular parameter matters.
If that model proves reliable, scientific publishing begins to look less like an archive of documents and more like a library of callable capabilities.
That is a much bigger shift than adding AI summaries to journals. A summary helps a person read faster. An executable paper agent can potentially do part of the research with them.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
