PubMed HealthSearch

SEARCH · PubMed Health

Results for “Computational scientific discovery”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

9 recordsLinked to original sources

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning

Immune dysregulatory disorders: perspective from solving a diagnostic odyssey.

PURPOSE OF REVIEW: Inborn errors of immunity (IEIs), once considered rare disorders characterized primarily by recurrent infections, are now recognized as a rapidly expanding group of diseases encompassing autoimmunity, autoinflammation, allergy, malignancy, and immune dysregulation. Advances in next-generation sequencing, functional immunology, and systems biology have revealed overlap between traditionally distinct disease categories and highlighted the complexity of genotype-phenotype relationships. RECENT FINDINGS: While this evolution has led to the discovery of hundreds of previously unrecognized disorders, it has also challenged conventional diagnostic paradigms and demonstrated how patients may have care spread across multiple specialties, without a clear medical home. These discoveries have also highlighted ongoing challenges translating scientific findings to the clinic including difficulties in accessing genomic testing, interpretation of variants of uncertain significance, impacts of incomplete penetrance and somatic mosaicism, and limited availability of specialized functional assays. Emerging computational approaches, including artificial intelligence, offer opportunities to accelerate diagnosis but cannot replace comprehensive clinical evaluation or longitudinal physician-patient relationships. SUMMARY: This perspective examines how the diagnostic odyssey for immune dysregulatory disorders has evolved, side-by-side with the changing framework for diagnosing rare immune diseases. We propose an integrated approach combining clinical phenotyping, genomics, functional validation, and multidisciplinary expertise to unite ongoing discovery between clinicians and scientists, diagnostics, and patient outcomes.

diagnostic odyssey

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

NAViFluX: a visualization‑centric platform for interactive analysis, refinement and design of genome‑scale metabolic networks.

MOTIVATION: Genome-scale metabolic network (GSMN) models enable flux-based metabolite fate discovery, metabolic engineering, drug target identification, and multi-omics integration. However, programming requirements, architectural complexity, and limited visualization support impede its adoption by the broader scientific community. Existing tools exclusively specialize in GSMN analyses or visualization while lacking important features such as pathway-specific views, database-integrated refinement, and comprehensive enrichment and perturbation analyses. RESULTS: Here, we present NAViFluX (metabolic Network Analysis and Visualization of Flux), a visualization-centric, web browser-based tool that unifies native pathway/subsystem map generation, interactive model refinement via KEGG/BiGG, pathway merging and modules for flux computations, topology, and functional enrichment all within network views. Using three independent case studies on Escherichia coli, the utility of NAViFluX for characterization of nutrient-specific metabolic adaptations, enhancing gene essentiality predictions and interpretability, and rational design of an optimized carbon-fixing metabolic state is demonstrated. AVAILABILITY AND IMPLEMENTATION: All source code and supplementary files associated with the case studies are publicly available via Zenodo at https://zenodo.org/records/19107831. NAViFluX can be easily installed as a standalone software through https://github.com/bnsb-lab-iith/NAViFluX.

Metabolic Networks and Pathways

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

Influence of Major Histocompatibility Complex (MHC) Diversity on Immune Modulation, Pathogenesis, and Control of Lumpy Skin Disease Virus.

INTRODUCTION: Lumpy Skin Disease Virus (LSDV), a member of the genus Capripoxvirus within the family Poxviridae, is an economically important transboundary viral pathogen affecting cattle and water buffalo. The disease causes severe production losses through decreased milk yield, infertility, hide damage, reduced growth performance, and occasional mortality. The rapid geographic spread of LSDV, together with its vectorborne transmission and emerging recombinant strains, has intensified the need for improved understanding of viral pathogenesis, host immune responses, and effective prevention strategies. In particular, the role of the bovine Major Histocompatibility Complex (BoLA/MHC) in regulating antiviral immunity, disease susceptibility, and vaccine responsiveness has gained increasing scientific attention. METHODS: This review summarises the published literature related to the epidemiology, transmission, structure, pathogenesis, diagnosis, prevention, and control of LSDV, with special emphasis on the immunological and molecular role of bovine MHC molecules. Relevant studies concerning BoLA-mediated antigen presentation, immunoinformaticsbased epitope prediction, vaccine development, antiviral drug repurposing, molecular docking, genomic surveillance, and diagnostic approaches, including PCR- and ELISAbased assays, were critically evaluated. Recent advances in computational biology, molecular virology, and host-pathogen interaction studies were also reviewed. RESULTS: The reviewed studies demonstrate that Lumpy Skin Disease Virus (LSDV) possesses a complex double-stranded DNA genome enabling immune modulation and efficient transmission through arthropod vectors such as mosquitoes, ticks, and biting flies. Disease progression involves systemic viral replication, vascular injury, dermal necrosis, and inflammatory skin lesions. Real-time PCR remains the most sensitive diagnostic method for early detection, while ELISA supports surveillance. Evidence highlights the central role of bovine Major Histocompatibility Complex (BoLA) molecules in antigen presentation and T-cell activation. Computational studies identified promising BoLA-binding epitopes and repurposed antiviral candidates, including ivermectin, theaflavin, canagliflozin, and tepotinib, for future therapeutic development. DISCUSSION: Current evidence indicates that effective LSDV control requires integration of molecular diagnostics, vector management, vaccination, and host immunogenetics. BoLAguided immunoinformatics provides promising opportunities for developing multi-epitope vaccines, although experimental validation remains essential. Similarly, repurposed antiviral candidates require comprehensive in vivo and pharmacological evaluation before clinical application. Future research should focus on elucidating viral immune-evasion mechanisms, validating predicted epitopes, and translating computational findings into practical vaccines and therapeutics for sustainable disease control. CONCLUSION: Lumpy Skin Disease continues to pose a major threat to global cattle health and livestock economies. Advances in molecular diagnostics, genomic surveillance, antiviral drug discovery, and BoLA-guided vaccine design provide promising opportunities for improved disease control. Understanding the interaction between LSDV and the bovine MHC system is essential for developing next-generation vaccines, immunotherapeutics, and precision disease-management strategies. Future research should prioritise experimental validation of predicted epitopes, large-scale vaccine trials, and mechanistic studies on host-virus immune interactions to establish effective and sustainable global control programs for LSDV.

BoLA

NLCD: A method to discover nonlinear causal relations among genes.

Distinguishing correlation from causation is a fundamental challenge in many scientific fields, including biology, especially when interventions like randomized controlled trials are infeasible and only observational data are available. Methods based on statistical tests of conditional independence within the Mendelian Randomization framework can detect causality between two observed variables that are each associated with a third instrumental variable. However, these methods for detecting causal relationships between traits (e.g., two gene expression or clinical traits associated with a genetic variant, all observed in the same population) often assume a linear relationship, thereby hindering the discovery of causal gene networks from genomics data. We have developed NLCD, a method for NonLinear Causal Discovery from genomics data based on nonlinear regression modeling and conditional feature importance scoring. NLCD uses these techniques to extend the statistical tests in an existing linear causal discovery method called the Causal Inference Test (CIT). We benchmarked NLCD against current state-of-the-art methods: CIT, Findr, and MRPC. On simulated datasets, NLCD performs comparably to most methods in detecting linear relations (Average AUPRC (Area Under the Precision-Recall Curve) of NLCD = 0.94, CIT = 0.94, Findr = 0.94, and MRPC = 0.99), and outperforms them in detecting nonlinear (sine and sawtooth type) relations between two genes (Average AUPRC of NLCD = 0.76, CIT = 0.60, Findr = 0.56, and MRPC = 0.73). When tested on a nonlinear subset of a yeast genomic dataset to recover known causal relations involving transcription factors, NLCD and CIT performed comparable to each other and slightly better than Findr and MRPC (Average AUPRC of NLCD = 0.82, CIT = 0.81, Findr = 0.71, and MRPC = 0.54). On application to a human genomic dataset, NLCD revealed active causal gene pairs (IRF1 → PSME1 and HLA-C → HLA-T) in the muscle tissue, and clarified the promises and challenges in discovering causal gene networks in tissues under in vivo human settings.

Humans