An AI chemistry agent asked a fairly ordinary question: Did a particular molecule have a known hepatobiliary side effect?
The agent chose a relevant tool to look for the answer. Then it fed the tool a molecular string in an invalid format. The tool rejected it.
The agent then tried to repair the molecule using another tool, but supplied essentially the same malformed input and got another error. Instead of stopping, the agent moved on, switching to a generic AI-expert tool after its other options had failed, and ultimately returned a confident answer to the original question.
The answer was wrong.
The failure by ChemAgent, a research system that was built to test how chemistry-focused AI language agents perform when given access to specialized tools, was documented in a 2025 paper evaluating chemistry-focused language agents. The incident was not a story about an AI system going rogue. It was closer to an ordinary software problem: bad input, failed error handling and an improvised workaround that produced a plausible but incorrect scientific conclusion.
That distinction has taken on new urgency in the wake of a high-profile agent failure recently revealed by OpenAI. In July, AI agents running an OpenAI cybersecurity evaluation circumvented controls meant to isolate them, reached the public internet and compromised systems on an infrastructure platform called Hugging Face while pursuing the task they had been assigned.
OpenAI later said the models had resorted to “misaligned strategies” to solve hard tasks, turning the episode into an urgent warning about what can happen when increasingly autonomous agents act beyond their intended boundaries.
For drug developers, the more immediate risk is less dramatic but still potentially quite costly. A mundane failure can enter a workflow in which AI systems do far more than simply toss out predictions. They increasingly gather evidence, call specialized tools, interpret outputs and help determine what happens next.
As more large pharmas chase super computing power, the scale of that algorithmic ambition is growing.
In September, Stanford researchers published the results of what happened when they created a “virtual biotech” built with AI agents to model a drug development company. In one task, more than 37,000 AI agents annotated outcomes from 55,984 clinical trials; in another, the system integrated multiple forms of biological evidence to propose a therapeutic strategy for lung cancer. The researchers said the results demonstrated the agents could conduct “transparent, multiscale analysis,” but noted that the system was human-guided, and that its therapeutic hypotheses still require real-world testing and validation.
That gap — between an agent that informs a scientist and one trusted to keep acting on its own conclusions — may be where the real risk changes.
When a prediction starts steering the workflow
Many traditional machine-learning systems in drug discovery are built to answer narrower questions: predict a property, rank compounds or identify a pattern in a dataset. Agents can add orchestration around those models, drawing in context and deciding which tools or evidence to use next.
“Agentic models bring context on top of that,” said Srijit Seal, principal scientist at Human Chemical Company, which builds AI platforms for chemical safety and regulatory analysis.
Seal described the near-term role of agents less as autonomous scientists than as decision-support systems. They can help researchers interpret internal and published evidence, rank hypotheses, consider whether to modify a compound, decide what experiment might be useful or prioritize among several paths forward.
But when it comes to gauging their effectiveness, there’s a key question at play.
“Can they come up with hypotheses ranked in a certain way that helps scientists reach conclusions faster?” Seal said.
That distinction matters because an agent does not need the authority to approve or kill a drug to influence a program. If it changes which hypothesis receives attention, which safety question gets investigated or which experiment a scientist runs next, a wrong conclusion can still redirect time and resources.
A 2026 Nature Reviews Drug Discovery perspective made a similar point about measuring AI’s value in drug discovery. Rather than focusing only on model performance, the authors argued that benchmarks should increasingly ask whether AI improves the decisions researchers actually make.
Examining documented failures can provide a good starting point for digging into that question.
Not every wrong answer is equally dangerous
The chemistry-agent incident offers a useful example precisely because, in Seal’s view, its mistake points in the less dangerous direction.
The system concluded that there was a toxicity concern when the correct answer was no — effectively a false positive. In a real development workflow, a false positive could waste money, prompt an unnecessary assay or cause researchers to undervalue a viable compound. But it could also be caught when a scientist runs a confirmatory test.
Seal argues that the inverse mistake is more consequential.
“A false negative is way more harmful than a false positive,” he said.
In Seal’s estimation, if an AI system incorrectly reassures researchers that a compound is safe, the risk is not merely that another test gets added. It is that scrutiny could be reduced.
“You say a compound is safe, but it’s not. That’s going to cause way more problems. It’s going to skip tests,” he said.
There is an important limit to that comparison: The published ChemAgent work did not demonstrate that its failure architecture caused a dangerous false negative inside a real pharmaceutical program.
What it showed is that an agent could encounter explicit tool errors, fail to repair the underlying problem and still produce the wrong scientific answer. Seal’s point is that the direction of such an error can matter as much as the fact that it happened.
What lets an error survive?
One intuitive fear around agentic systems is that a bad answer from one step will simply be inherited by the next, compounding until a small mistake becomes a large one. Seal pushed back on that concern as too simplistic.
Long-running agents can accumulate errors, he said, but the outcome depends heavily on how the workflow is engineered. A later step can verify provenance, rerun a calculation, demand independent evidence or require a person to approve an action.
“You can stop it at certain gates,” Seal said. “These are more engineering questions and not research.”
That reframes the safety question. The problem is not simply whether an AI agent can make a mistake — every model can. The more useful question is what architecture allows a mistake to survive long enough to influence another decision.
Drug development already contains natural checkpoints. Scientists across medicinal chemistry, toxicology, clinical development and other functions routinely decide whether evidence is strong enough for a molecule to advance, change or get shelved. Agent governance can be built around many of those same decision points rather than imagining drug R&D as an open-loop chain of machines talking only to one another, Seal said.
For R&D leaders, that makes agent adoption as much an operating-model question as a model-performance question: what an agent is allowed to act on, what evidence it has to show and where a person still has veto power.
“It’s always a collaborative effort,” he said.
That human involvement also limits how much of today’s agentic-AI risk should be described as autonomous. Even Stanford’s “virtual biotech” uses a virtual chief scientific officer to coordinate specialized agents, while human users define questions and scientific reviewers evaluate methods and claims. The system can produce analyses and hypotheses at enormous scale, but the researchers explicitly position it as a way to support therapeutic development decisions, not replace experimental validation.
“You can stop it at certain gates. These are more engineering questions and not research.”

Srijit Seal
Principal scientist, Human Chemical Company
The autonomy cliff
For now, Seal sees agents primarily as “a support system, a sidekick mechanism” for scientists.
He does not believe drug companies are broadly handing major capital-allocation or development decisions to autonomous agents today. That means an AI system can be wrong without its answer automatically becoming an experiment, a budget decision or the next stage of a drug program.
The risk changes if those boundaries erode.
“The problem starts when it’s completely autonomous, and then you have to trust it,” Seal said.
A sufficiently autonomous system could eventually choose an experiment, interpret the result and use that interpretation to select the next action inside the same loop. The more often those steps happen without an outside checkpoint, the fewer opportunities there are for someone — or another independently designed system — to notice that the premise was wrong.
That brings the question back to the malformed molecule string in the ChemAgent experiment. The first mistake was trivial. The tool’s response was not subtle: the input was invalid. The striking part was that the agent encountered evidence that something had gone wrong and kept moving anyway.
The difficult problem for increasingly autonomous drug discovery operations may not be building an AI system that is never wrong. It may be building one that knows when the evidence underneath its next decision is no longer trustworthy enough to keep going.