Follow

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Subscribe

Biological AI models: new paradigms to leverage the languages of life

image

Biological Artificial Intelligence (AI) models are advancing fastest in data-rich areas like protein structure, while progressing at a slower pace where data is less abundant and structured, such as single-cell biology. The EU is equipped with strong scientific base and computing capacity to develop biology AI models, but more efforts are needed to improve intra-EU collaboration, coordination and data governance, new research shows.

The findings are published in the JRC report Artificial Intelligence for Biology: Capabilities, Readiness, and Policy Implications. It provides an empirically grounded assessment of the current landscape of AI models applied to biological data, such as DNA, RNA, and Proteins. 

Emerging properties of biological AI models

AI is transforming biological research, enabling breakthroughs in genomics, proteomics, drug discovery, and beyond. AI models trained on biological data are advancing rapidly, outpacing existing regulatory and governance frameworks. This rapid progress raises critical questions about scientific maturity, deployment readiness, and governance.

Particularly advanced biological AI models are in place for protein-centric applications such as structure prediction, function annotation, and molecular design, relevant in fields of drug design and discovery. Such progress of protein models (eg. AlphaFold) was made possible by decades-long commitment of a dedicated research community, and thanks to curated data from repositories including the Protein Data Bank and UniProt, supported by European Research Infrastructures like the European Molecular Biology Laboratory (EMBL).

In comparison, areas such as single cell biology, which holds clinical relevance among other in characterising tumours to help predict response to immunotherapy, remain less developed. One of the main reasons behind is limited and less standardised data. 

Maturity and readiness for real-world use

Knowing that an AI model performs well in scientific benchmarks does not tell policymakers, investors or healthcare providers whether that model is actually ready to be used in practice. An added value of the report is that it provides a framework to assess deployment readiness, combining domain-specific maturity with technology readiness levels (TRL). 

The results show that high-profile models such as AlphaFold and ESM3 are domain-mature but remain at low-to-mid TRL. While being the most advanced in development in their research field, they have not been certified for clinical or industrial deployment. Indeed, none of the surveyed models has been subject of an integrated readiness assessment.

This “maturity paradox”, a term coined by the authors, highlights insufficiently addressed validation of the full innovation pipeline, including governance, and benchmarking of integration in real-world applications. The divergence between domain maturity and technology readiness can pose several risks, including biosecurity concerns among others, and requires close attention, particularly where publicly available models could be misused for applications such as pathogen design or toxin engineering.

Building and resourcing biological AI models

Building biological AI models depends on three main factors: training data, computational infrastructure, and collaboration. 

Training data are unevenly distributed across domains and regions. Some datasets are standard within a field, such as the Protein Data Bank or UniProt resources (such as UniRef) for protein models, while RNA, single-cell, and clinical domains rely on a smaller and more varied set of sources.

Protein models increasingly rely on synthetic data from the AlphaFold Database, indicating a growing dependence on predicted data rather than exclusively curated, experimentally derived data. The US and EU host most training datasets, followed by China, the UK, and Switzerland. Fragmented repositories make reproducibility and interoperability difficult. 

Access to computing resources also remains uneven. Industry typically trains models on larger datasets and can rely on more hardware and longer training times than academia, creating a significant resource gap. Europe has sufficient high-performance computing capacity thanks to The European High Performance Computing Joint Undertaking (EuroHPC JU) and its network of supercomputing facilities. The recently established AI Factories, aimed at boosting AI development and uptake across the EU, leverage the EuroHPC JU capacity. 

Collaboration patterns further shape the field: academia participates in the development of 85% of the surveyed models, while industry participates in nearly 40%. However, as industry becomes more active, this may reflect a broader trend towards the non-disclosure of proprietary models and datasets, as only 17% of industry-only developed models release training code. 

Geographically, intra-EU collaboration is limited compared to the EU collaboration with the US, China, and the UK. These three regions are the most dominant in model development. Among the global top 20 model developers, the only EU representative is the Technical University of Munich.

Policy considerations

The report highlights four recommendations for greater EU policy action

  • broaden support for emerging promising research topics, like single cell biology and model architectures (such as multimodal models), while aligning AI for biology with wider EU goals in health, food security, and the circular bioeconomy
  • strengthen biological data infrastructure through better coordination and transparent data quality assessment
  • strengthen Europe’s position by supporting European biological AI foundation models as public goods, further encouraging strategic intra-EU collaboration
  • develop frameworks that assess both scientific maturity and technology readiness, including clinically relevant benchmarks and clearer regulatory pathways. 

Related content

Artificial Intelligence for Biology: Capabilities, Readiness, and Policy Implications

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
EIC CoC Label