Fleming Initiative and Google DeepMind launch AI evaluation infrastructure for antimicrobial resistance

The Fleming Initiative has launched a three-year programme funded by Google DeepMind to develop evaluation methods and standards for AI systems in antimicrobial resistance (AMR). The goal: let researchers, clinicians and policymakers know when a model is truly accurate, reliable and ready for clinical use — not just when it looks good on an internal benchmark.
The announcement cites AlphaFold as proof of what clear standards enable. The protein-structure prediction model, which achieved unprecedented accuracy, won the 2024 Nobel Prize in Chemistry, and its predictions are open to researchers worldwide. But in AMR, where AI is already used to identify resistant pathogens, discover novel antibiotics, track spread and support clinical decisions, no accepted framework for model comparison exists. The result: researchers cannot tell which model is better, hospitals do not know what to adopt, and regulators are stuck without criteria.
The programme will build evaluation frameworks and benchmarking approaches for high-priority AMR applications. That means defining how to test a model, what evidence is needed to prove consistent performance across populations, care settings and diverse datasets. At the same time, the team will work with the global AMR community so that data is collected, cleaned and reported in ways that support serious evaluation — not ad-hoc formats that break comparisons.
This does not start from zero. The programme rests on a joint report by the initiative and Google DeepMind, "Harnessing Artificial Intelligence to Tackle Antimicrobial Resistance", and on the Google DeepMind Academic Fellowship in AI and AMR, whose first recipient was recently appointed associate professor of bioelectronic medicine at Imperial College London. A knowledge base and researcher connections already exist; the task now is to turn them into a standard.
"AI's potential against AMR is immense," said Prof. Alison Holmes, director of the initiative, "but without confidence that models represent the populations and environments in which they will be deployed, they cannot be adopted with confidence." Agata Laydon, science lead at the Google DeepMind Impact Accelerator, added that objective, robust evaluation is the key to translating potential into real-world impact, and that the support aims to build the critical infrastructure the global community needs.