PubMed HealthSearch

SEARCH · PubMed Health

Results for “Computational Biology”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Biomathematical enzyme kinetics model of prebiotic autocatalytic RNA networks: degenerating parasite-specific hyperparasite catalysts confer parasite resistance and herald the birth of molecular immunity.

Catalysis and specifically autocatalysis are the quintessential building blocks of life. Yet, although autocatalytic networks are necessary, they are not sufficient for the emergence of life-like properties, such as replication and adaptation. The ultimate and potentially fatal threat faced by molecular replicators is parasitism; if the polymerase error rate exceeds a critical threshold, even the fittest molecular species will disappear. Here we have developed an autocatalytic RNA early life mathematical network model based on enzyme kinetics, specifically the steady-state approximation. We confirm previous models showing that these second-order autocatalytic cycles are sustainable, provided there is a sufficient nucleotide pool. However, molecular parasites become untenable unless they sequentially degenerate to hyperparasites (i.e. parasites of parasites). Parasite resistance-a parasite-specific host response decreasing parasite fitness-is acquired gradually, and eventually involves an increased binding affinity of hyperparasites for parasites. Our model is supported at three levels; firstly, ribozyme polymerases display Michaelis-Menten saturation kinetics and comply with the steady-state approximation. Secondly, ribozyme polymerases are capable of sustainable auto-amplification and of surmounting the fatal error threshold. Thirdly, with growing sequence divergence of host and parasite catalysts, the probability of self-binding is expected to increase and the trend towards cross-reactivity to diminish. Our model predicts that primordial host-RNA populations evolved via an arms race towards a host-parasite-hyperparasite catalyst trio that conferred parasite resistance within an RNA replicator niche. While molecular parasites have traditionally been viewed as a nuisance, our model argues for their integration into the host habitat rather than their separation. It adds another mechanism-with biochemical precision-by which parasitism can be tamed and offers an attractive explanation for the universal coexistence of catalyst trios within prokaryotes and the virosphere, heralding the birth of a primitive molecular immunity.

Kinetics

Metabolic network reconstruction as a resource for analyzing Salmonella Typhimurium SL1344 growth in the mouse intestine.

Nontyphoidal Salmonella strains (NTS) are among the most common foodborne enteropathogens and constitute a major cause of global morbidity and mortality, imposing a substantial burden on global health. The increasing antibiotic resistance of NTS bacteria has attracted a lot of research on understanding their modus operandi during infection. Growth in the gut lumen is a critical phase of the NTS infection. This might offer opportunities for intervention. However, the metabolic richness of the gut lumen environment and the inherent complexity and robustness of the metabolism of NTS bacteria call for modeling approaches to guide research efforts. In this study, we reconstructed a thermodynamically constrained and context-specific genome-scale metabolic model (GEM) for S. Typhimurium SL1344, a model strain well-studied in infection research. We combined sequence annotation, optimization methods and in vitro and in vivo experimental data. We used GEM to explore the nutritional requirements, the growth limiting metabolic genes, and the metabolic pathway usage of NTS bacteria in a rich environment simulating the murine gut. This work provides insight and hypotheses on the biochemical capabilities and requirements of SL1344 beyond the knowledge acquired through conventional sequence annotation and can inform future research aimed at better understanding NTS metabolism and identifying potential targets for infection prevention.

Salmonella typhimurium

How does date-rounding affect phylodynamic inference for public health?

Phylodynamic analyses infer epidemiological parameters from pathogen genome sequences for enhanced genomic surveillance in public health. Pathogen genome sequences and their associated sampling dates are the essential data in every analysis. However, sampling dates are usually associated with hospitalisation or testing and can sometimes be used to identify individual patients, posing a threat to patient confidentiality. To lower this risk, sampling dates are often given with reduced date-resolution to the month or year, which can potentially bias inference. Here, we introduce a practical guideline on when date-rounding biases the inference of epidemiologically important parameters across a diverse range of empirical and simulated datasets. We show that the direction of bias varies for different parameters, datasets, and tree priors, while compounding with lower date-resolution and higher substitution rates. We also find that bias decreases for datasets with longer sampling intervals, implying that our guideline is most applicable to emerging datasets. We conclude by discussing future solutions that prioritise patient confidentiality and propose a method for safer sharing of sampling dates that translates them them uniformly by a random number.

Humans

Coarse-grained model of serial dilution dynamics in synthetic human gut microbiome.

Many microbial communities in nature are complex, with hundreds of coexisting strains and the resources they consume. We currently lack the ability to assemble and manipulate such communities in a predictable manner in the lab. Here, we take a first step in this direction by introducing and studying a simplified consumer resource model of such complex communities in serial dilution experiments. The main assumption of our model is that during the growth phase of the cycle, strains share resources and produce metabolic byproducts in proportion to their average abundances and strain-specific consumption/production fluxes. We fit the model to describe serial dilution experiments in hCom2, a defined synthetic human gut microbiome with a steady-state diversity of 63 species growing on a rich media, using consumption and production fluxes inferred from metabolomics experiments. The model predicts serial dilution dynamics reasonably well, with a correlation coefficient between predicted and observed strain abundances as high as 0.8. We applied our model to: (i) calculate steady-state abundances of leave-one-out communities and use these results to infer the interaction network between strains; (ii) explore direct and indirect interactions between strains and resources by increasing concentrations of individual resources and monitoring changes in strain abundances; (iii) construct a resource supplementation protocol to maximally equalize steady-state strain abundances.

Gastrointestinal Microbiome

Understanding disease-associated metabolic changes in human colonic epithelial cells using the iColonEpithelium metabolic reconstruction.

The colonic epithelium plays a key role in the host-microbiome interactions, allowing uptake of various nutrients and driving important metabolic processes. To unravel detailed metabolic activities in the human colonic epithelium, our present study focuses on the generation of the first cell-type-specific genome-scale metabolic model (GEM) of human colonic epithelial cells, named iColonEpithelium. GEMs are powerful tools for exploring reactions and metabolites at the systems level and predicting the flux distributions at steady state. Our cell-type-specific iColonEpithelium metabolic reconstruction captures genes specifically expressed in the human colonic epithelial cells. iColonEpithelium is also capable of performing metabolic tasks specific to the colonic epithelium. A unique transport reaction compartment has been included to allow for the simulation of metabolic interactions with the gut microbiome. We used iColonEpithelium to identify metabolic signatures associated with inflammatory bowel disease. We used single-cell RNA sequencing data from Crohn's Diseases (CD) and ulcerative colitis (UC) samples to build disease-specific iColonEpithelium metabolic networks in order to predict metabolic signatures of colonocytes in both healthy and disease states. We identified reactions in nucleotide interconversion, fatty acid synthesis and tryptophan metabolism were differentially regulated in CD and UC conditions, relative to healthy control, which were in accordance with experimental results. The iColonEpithelium metabolic network can be used to identify mechanisms at the cellular level, and we show an initial proof-of-concept for how our tool can be leveraged to explore the metabolic interactions between host and gut microbiota.

Humans

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning

Informing agent-based models with spatial data using convolutional autoencoders.

MOTIVATION: Spatial computational models such as agent-based models (ABMs) offer powerful in silico tools to study tumor dynamics, yet imaging data are still rarely used to inform these models directly. RESULTS: We present an ABM optimization framework that leverages convolutional encoders to compare spatial patterns between experimental imaging data and ABM-generated outputs within a shared latent space. This quantitative comparison was used to estimate ABM parameters across three datasets, ranging from synthetic data to 3D tumoroid-T cell co-culture microscopy and histopathology images from The Cancer Genome Atlas skin cutaneous melanoma samples. Estimated parameters were evaluated using data-derived features and experimental knowledge, including experimental conditions and gene expressions. Simulations using optimized parameters reproduced key spatial features of the training images, such as tumor boundary complexity and tumor-tumor neighborhood structure. Together, these results demonstrate a flexible framework for ABM parameter optimization using spatial data across modalities, enabling systematic investigation of how spatial architecture influences tumor progression and immune interactions. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/SysBioOncology/ AutoencoderABM under the GPL-3.0 license, with corresponding data sets at https://zenodo.org/records/19022344.

Autoencoder

Transcriptome-wide analysis reveals sequence selection to avoid mRNA aggregation in E. coli.

The stability of RNA base pairing and its limited four-letter code create an intrinsic potential for promiscuous RNA-RNA interactions. In vitro, such interactions drive RNA to self-assemble into aggregates. This raises a fundamental unanswered question: within a confined cellular volume at physiological mRNA abundances, how much aggregation would arise from sequence-encoded chemistry alone? Here, we establish this baseline with large-scale kinetic simulations of the E. coli transcriptome. Our simulations reveal that sequence-encoded base-pairing energetics is sufficient to generate a dynamic network of large aggregates, organized by long, multivalent mRNA hubs. Strikingly, evolutionary analysis shows that native E. coli sequences exhibit clear signatures of selection to counteract this propensity: they fold more stably, minimize unstructured regions, and form weaker intermolecular contacts than dinucleotide-preserving controls. These findings demonstrate that maintaining transcriptome solubility has been a significant, previously unrecognized constraint shaping genome evolution, and provide a new lens to interpret cellular RNA management.

Biological Sciences (Biophysics and Computational

Identification of a Nonribosomal Peptide Analog With Activity Against Multiple Gram-Positive Bacteria via a Synthetic Bioinformatic Natural Product Discovery Approach.

Nonribosomal peptide (NRP) antibiotics exhibit potent biological activities and are broadly used in clinical therapy. Because most microorganisms are difficult to culture and many antibiotic biosynthetic genes are silent, traditional activity tracking approaches face major limitations in the discovery of novel NRPs. Here, based on a synthetic bioinformatic natural product (syn-BNP) discovery approach that integrates bioinformatics and chemical synthesis, a novel nonribosomal peptide synthetase (NRPS) gene cluster from the genome of Rhodococcus erythropolis D-1 was mined. A putative NRP scaffold synthesized by the NRPS encoded by this cluster was predicted. Through chemical synthesis and four rounds of structure-activity relationship (SAR) studies, 37 NRP analogs were ultimately generated. Among these analogs, ZURJC28 shows activity against multiple Gram-positive bacteria, including two drug-resistant strains. Mechanistic studies and metabolomics analyses revealed that ZURJC28 exerts membrane-disruptive activity associated with interaction with phosphatidylglycerol (PG)-enriched Gram-positive membranes, leading to membrane damage and widespread metabolic dysregulation. ZURJC28 also shows low cytotoxicity and low hemolytic activity, suggesting its preliminary in vitro safety profile.

Gram-Positive Bacteria

Protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique.

The advent of single-cell multi-omics sequencing technology makes it possible for researchers to leverage multiple modalities for individual cells. Here, we present a protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique called moETM. We describe steps for data preprocessing, multi-omics integration, inclusion of prior pathway knowledge, and cross-omics imputation. As a demonstration, we used the single-cell multi-omics data collected from bone marrow mononuclear cells (GSE194122) as in our original study. For complete details on the use and execution of this protocol, please refer to Zhou et al.1.

Deep Learning

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny

Agentic AI for Spatial Omics.

This highlight summarises recent advances in agentic artificial intelligence (AI) systems for spatial omics analysis. These systems are compared along two central tensions: autonomy versus accountability, and adaptability versus reproducibility. We argue that progress will depend not on maximising automation, but on defining where autonomy is appropriate.

Artificial Intelligence

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

Drug target ontology to classify and integrate drug discovery data.

BACKGROUND: One of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome. RESULTS: As part of that effort, we have developed a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while supporting manual knowledge acquisition. A modular, hierarchical implementation facilitate ontology development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships. CONCLUSIONS: DTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/ , Github ( http://github.com/DrugTargetOntology/DTO ), and the NCBO Bioportal ( http://bioportal.bioontology.org/ontologies/DTO ). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

Biological Ontologies

A generalized higher-order correlation analysis framework for multi-omics network inference.

Multiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest.

Genomics

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models