Follow

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Subscribe

Bioinformatics: rigorous quality controls, a key to effective public health decision-making

image

In an era where genomics – studying the DNA of living organisms – plays a pivotal role in public health decision-making, the reliability of bioinformatics analyses of large-scale biological data has never been more critical. 

This became especially clear during the COVID-19 pandemic. In the urgency to generate and share the virus’ genomic data, many workflows prioritised speed over consistency. Laboratories worldwide processed millions of samples using rapidly developed sequencing protocols and heterogeneous pipelines. However, without harmonised quality assurance mechanisms, some results, although fast, proved incorrect or misleading. Even subtle inaccuracies can compromise diagnostic precision, epidemiological modelling, and even vaccine development on a worldwide scale.

Quality in bioinformatics operates at multiple levels. First, it concerns the integrity of the raw data and the ability to detect and correct artefacts introduced during sequencing or data processing. Second, and equally important, it concerns the systematic evaluation of the computational tools themselves: without rigorous benchmarking (i.e., the controlled assessment of tool performance under diverse and realistic conditions) there is a genuine risk that researchers and surveillance systems draw inaccurate conclusions from suboptimal or misused methods. Both dimensions are essential for trustworthy, evidence-based genomic surveillance.

A JRC research team has directly addressed the two above-mentioned complementary dimensions – the integrity of raw data and the systemic assessment of computational tools – through three interlocking scientific contributions. The three studies ultimately highlight the importance of adopting harmonised strategies for bioinformatic data analyses in surveillance systems:

  • Ensuring robust trustworthy bioinformatics workflows. 
    PathoSeq-QC was developed as an example of modular bioinformatics workflow that goes beyond standard pipelines. Unlike traditional tools that focus on raw read quality, PathoSeq-QC assesses genomic homogeneity, cross-validates variant calls using multiple methods, and flags uncertainties that could compromise results. Successfully applied to SARS-CoV-2, H5N1, and Oropouche virus datasets, this workflow embodies transparency, reproducibility, and scalability – key principles for trust in science-driven policy.
  • Verification of real-world surveillance pipelines to avoid potential artefacts. 
    A collaboration with Slovenian laboratories revealed how undetected pipeline errors can lead to false alarms. Investigating SARS-CoV-2 sequences that appeared to contain dangerous mutations affecting PCR and antigen test detection, it was discovered that these “mutations” were artefacts introduced by a bioinformatics pipeline. This case underscores the need for continuous monitoring and harmonisation of data workflows to prevent silent failures that could misguide public health responses.
  • Benchmarking bioinformatics tools for accurate target detection.
    Even with high-quality data, the choice of analytical tool can affect results. In a systematic benchmarking study, four bioinformatics tools for detecting SARS-CoV-2 subgenomic RNAs (sgRNAs)– a key marker of active viral replication – were evaluated. Using synthetic datasets that emulate a wide range of sequencing conditions together with real‑world wastewater samples, we generated a comprehensive benchmarking resource. This dataset captures the diversity of mutation profiles, sequencing protocols and sample origins that affect downstream analyses, thereby providing laboratories with a standard reference set on which to evaluate and compare the performance of new detection tools. To support future research, a standardised benchmarking datasets has been made publicly available, providing a reusable resource for developing and validating new sgRNA detection tools.

These three contributions are interlinked: the wastewater data in the benchmarking study was processed with PathoSeq-QC; the artefacts identified in the Slovenian sequences are precisely the type of signal that PathoSeq-QC is designed to flag; and the benchmarking framework mirrors the cross-validation logic embedded in PathoSeq-QC’s architecture. 

Together, they constitute a multi-layered quality assurance paradigm, from raw data to tool selection, and they could become standard practice in genomic surveillance. Taken together, the three JRC contributions form a coherent and complementary response to a shared challenge: ensuring that bioinformatics delivers reliable, reproducible, and actionable results in public health contexts.

As omics technologies transition from research to regulatory applications, the JRC emphasises the need for:

  • A wider adoption of decision-support tools like PathoSeq-QC to enhance data comparability in health surveillance,
  • Cross-validation mechanisms embedded in all bioinformatics workflows, and
  • Standardised benchmarking datasets as community resources to improve tool validation.

Background

The COVID-19 pandemic highlighted both the power and the pitfalls of rapid genomic data generation – where speed sometimes came at the cost of accuracy. It has driven efforts to embed omics technologies in regulatory frameworks and global surveillance strategies. The European Food Safety Authority (EFSA) plans to routinely adopt -omics technologies in risk assessment by 2030, while the World Health Organization (WHO) has adopted a Global genomic surveillance strategy for pathogens with pandemic and epidemic potential 2022 – 2032

The approach taken by EFSA and WHO aims to ensure robust quality controls, essential to prevent misleading conclusions that could impact diagnostics, epidemiological modelling, and vaccine development. Episodes ranging from misidentified genotypes to spurious variant calls and erroneous drug‑resistance predictions, illustrate how artefacts or sub‑optimal bioinformatic pipelines can translate into misguided public‑health actions, delayed interventions, or inappropriate clinical treatments.

Related content

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows

How benchmarking of bioinformatics tools is essential for informed workflow selection: a case study on SARS-CoV-2 subgenomic RNA detection

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
EIC CoC Label