Ai2 conference exposed five barriers to autonomous scientific research agents

Artificial-intelligence systems can already retrieve literature, analyse data, generate code and propose hypotheses, but building agents that researchers can steer in real time as projects evolve remains a hard problem. That was the central conclusion of an event held by Ai2 on 27 August to mark the expansion of its partnership with the Paul Allen Center for Cancer Research at the Providence Cancer Institute in Sweden. The gathering brought together physician-researchers, computer scientists and public-health professionals to map current gaps in AI for science, and five challenges surfaced repeatedly across all talks and the panel discussion.
The collaboration with the Paul Allen Center served as the starting point for a broader discussion. Ai2 researchers presented AutoDiscovery, a system that generates hypotheses through statistical analysis. They observed that many of the system’s suggestions were numerically surprising but lacked biological or clinical plausibility without expert context. Only after researchers injected domain knowledge did the system become useful. The implication was that the shortfall lies not in computational power but in the interface between the machine and human expertise.
The first challenge, described as “scientific taste” by Bodhisattwa Prasad Majumder, a senior researcher at Ai2, is the judgment that lets a researcher distinguish an interesting result from one that is trivial, implausible, already known or hopeless. Existing systems cannot make this distinction on their own. The goal is not to transfer the judgment entirely to AI, but to give researchers better ways to feed expertise, priorities and evolving problem understanding into the system so that AI assists without independently deciding which directions merit investigation.
The second challenge concerns maintaining steerability over time. Scientific research rarely follows a fixed plan: an unexpected experimental outcome, a new paper, or a decision to add a dataset, switch tools or redirect an agent can all occur. Majumder argued that current agents are difficult to steer in long-term investigations. Scientists need to update instructions, context and tools as new results appear, without restarting from scratch or retraining the base model each time. For scientific agents, the ability to be guided and to incorporate continuous knowledge updates are interdependent; a useful agent cannot rely solely on an initial brief but must remain adaptive as the research progresses.
The third challenge involves deciding what to delegate to an agent. Hoifung Poon, chief AI officer at Recursion, distinguished two categories of contribution: productivity gains from tasks that people already know how to do but find tedious or time-consuming—such as searching records, structuring information or synthesising literature—and creativity gains. The former category is relatively easy to define and, crucially, to validate. While the source material cuts off before detailing the remaining challenges, Poon’s distinction marks the practical boundary where automation ends and genuine collaborative research begins.