PubMed HealthSearch

SEARCH · PubMed Health

Results for “evolutionary scale model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ESMpHLA: Evolutionary Scale Model-Based Deep Learning Prediction of HLA Class I Binding Peptides.

The recognition of endogenous peptides by HLA class I plays a crucial role in CD8+ T cell immune responses and human adaptive cell immune. Thus, the prediction of HLA class I-peptide binding affinities is always the core issue for the research of immune recognition and vaccine development. In this study, an evolutionary scale model (ESM) combined with parallel CNN blocks and a cross attention mechanism was used to construct a novel ESMpHLA model for predicting HLA class I binding peptides. Based on the 91,560 binding peptides of 41 HLA-A alleles, 56,731 of 50 HLA-B alleles and 2444 of 10 HLA-C alleles, the ESMpHLA model was successfully established and achieved satisfying prediction performances with the overall accuracy and AUC values of 0.874 and 0.938 for the test dataset. The results indicate that the ESMpHLA model performs well in dealing with different HLA class I 2-field alleles as well as the peptides with different lengths. Then, the generalisation ability of the ESMpHLA model was validated by an independent test dataset compiled from recent IEDB weekly benchmark datasets. The results showed that the ESMpHLA model achieved the highest ROC-AUC and PR-AUC values when compared with the latest BVMHC, CapsNet-MHC, STMHCpan and BVLSTM models. In addition, two ensemble models were also established by integrating the above 5 deep learning models using soft-voting and hard-voting strategies.

Humans

SimHumanity: Using SLiM 5.0 to run whole-genome simulations of human evolution.

The reconstruction of human evolutionary history has undergone repeated advances, each made possible by methodological innovations. In recent decades, genetic and genomic data played a central role in the reconstruction of major evolutionary events such as the out-of-Africa migration, and genetic simulations of human evolutionary history have come to play a major role in testing more specific hypotheses including proposed patterns of migration and admixture with archaic hominins. Increasing computational power has allowed human evolutionary history to be modeled at ever-larger scales, but simulations that encompass the complete human genome, including sex chromosomes and mitochondrial DNA, have been difficult due to the lack of support for whole-genome models in commonly used evolutionary simulation frameworks. With the recent introduction of SLiM 5 such simulations are now straightforward to construct, allowing the easy simulation of humans at whole-genome scale under different demographic models and evolutionary dynamics. We here present three versions of a reusable, customizable, open-source SLiM 5 model for simulating the molecular evolution of the full human genome. We also show some simple analyses of results from the model, to illustrate its utility. We hope this model, which we have nicknamed "SimHumanity" in jest, will facilitate further progress in the field of human evolutionary simulations.

SLiM

Large-scale genomic analysis places Chinese CC398 as a persistent human-associated MSSA lineage apart from the dominant global LA-MRSA clade.

Staphylococcus aureus clonal complex (CC)398 has emerged as a dominant livestock-associated methicillin-resistant S. aureus (LA-MRSA) lineage worldwide; however, its evolutionary trajectory and regional diversification remain incompletely understood. We developed a core-genome multilocus sequence typing (cgMLST) scheme with hierarchical clustering and applied it to over 30,000 S. aureus genomes, revealing frequent cross-border transmission of CC398. Subsequent time-calibrated phylogenetic analysis placed the most recent common ancestor at 1942 (95% CI: 1939-1945), with the human-to-livestock host jump around 1969 (95% CI: 1968-1972). Chinese CC398 exhibits a distinct trajectory: unlike the LA-MRSA lineages dominating Europe and North America, Chinese isolates are predominantly human-associated methicillin-susceptible S. aureus (HA-MSSA), forming unique East Asia-specific phylogroups (SAP1, SAP2, and AP1-AP3), with distinct resistance and virulence profiles. The LA lineage remains limited in China, with multinational mixed clusters emerging only after 2019. Analysis of global transmission networks revealed a significant correlation between LA-CC398 spread and international trade in fresh swine products, while no such correlation was observed for the human-associated lineage. Beyond the established lineage markers tet(M) and scn, our analysis identified additional differentially distributed genes, including cadC-a chromosomal cadmium resistance regulator-as a novel HA-lineage-enriched gene whose functional role in host adaptation remains to be determined. This study reveals that CC398 followed fundamentally different evolutionary paths in China versus Western countries, challenging a one-size-fits-all model of its dissemination.IMPORTANCEThis study illustrates how large-scale microbial genomics can resolve the evolutionary origins and regional diversification of bacterial pathogens. By applying a novel cgMLST scheme to over 30,000 S. aureus genomes, we show that CC398 followed fundamentally different evolutionary paths in China versus Western countries-challenging the prevailing model of uniform global dissemination-and that livestock-associated MRSA expansion is closely linked to international trade in fresh pork products. These findings highlight the need for integrated surveillance across human, animal, and trade interfaces to anticipate the emergence and spread of zoonotic pathogens.

Staphylococcus aureus

In silico analysis and comparison of the metabolic capabilities of different organisms by reducing metabolic complexity.

BACKGROUND: Understanding how metabolic capabilities diverge across microbial species is essential for deciphering community function, ecological interactions, and the design of synthetic microbiomes. Despite shared core pathways, microbial phenotypes can differ markedly due to evolutionary adaptations and metabolic specialization. Genome-scale metabolic models (GEMs) provide a systems-level framework to explore these differences; however, their complexity hinders direct comparison. RESULTS: We introduce NIS (Neidhardt-Ingraham-Schaechter), a computational workflow that integrates the redGEM, lumpGEM, and redGEMX algorithms to systematically reduce genome-scale models into biologically interpretable modules. This approach enables direct, quantitative comparison of fueling pathways, biomass biosynthetic routes, and environmental exchange processes while retaining essential metabolic information. We first demonstrate the utility of NIS by analyzing Escherichia coli and Saccharomyces cerevisiae, which revealed both conserved and divergent strategies in central metabolism, biosynthetic cost, and substrate utilization. We then applied NIS to the core honeybee gut microbiome, uncovering distinct metabolic traits, functional redundancy, and complementarity that help explain auxotrophy, cross-feeding interactions, and microbial coexistence. CONCLUSIONS: NIS provides an automated, scalable, and reproducible framework for dissecting microbial metabolic networks beyond gene content or taxonomy. By linking metabolism to ecological function, NIS offers new opportunities to interpret microbial community dynamics and to support the rational design of microbiomes in health, agriculture, and environmental applications. Video Abstract.

Metabolic Networks and Pathways

Predictive evolutionary genomics: principles, validation, and practice.

Climate change and habitat loss are driving rapid evolutionary responses in populations world-wide, which creates an urgent need for evolutionary forecasting in conservation and agriculture. Such forecasting can be categorized into three time scales: trait-based models that use multivariate quantitative genetic equations to project correlated phenotypic responses up to c. 20 generations, allele-based analyses that model allele frequency dynamics up to 100 generations, and composite adaptation scores that aggregate many small effects to yield predictions across longer horizons. However, these approaches have remained largely disconnected. Here, we present a Bayesian framework that integrates these three complementary approaches for evolutionary prediction. Our framework combines genomic, phenotypic, and environmental data to yield probabilistic predictions with explicit uncertainty. We show how predictive evolutionary forecasts can be validated with experimental evolution, field experimentation, historical specimens, and reciprocal transplants. These validated forecasts can help advance conservation and agricultural programmes by helping predict which populations are at risk of future extinction, optimizing breeding programmes for future climates, and planning ecosystem management under environmental change. By supporting a shift towards more predictive approaches in evolutionary biology, this framework may help improve our ability to manage biodiversity and food security in a changing world.

Genomics

A pangenome framework uncovers the role of deletions in repeated evolution of cave-derived traits.

Structural variants (SVs) are increasingly recognized as key contributors to adaptive evolution, yet they remain underexplored compared with single-nucleotide variation. To understand how large-scale genomic changes shape repeated evolution, we leveraged multiple levels of sequence data across the powerful evolutionary model system of the Mexican tetra fish (Astyanax mexicanus). We constructed one of the first pangenome graphs from a naturally evolving vertebrate, enabling comprehensive discovery of SVs among 120 fish from 11 populations. We discover substantial amounts of structural variation and explore the roles of genomic biases and selection in shaping the distribution of these variants. More than 2400 high-confidence cave-specific deletions are enriched in biological pathways involved in vision, metabolism, and behavior and cluster nonrandomly in quantitative trait loci linked to cavefish traits. Additionally, 67 genes harbor unique deletions between independent cavefish lineages. These reused genes show evidence of population-specific selection (99% contain selective sweeps compared with 8%-15% in genes lacking SVs), indicating that deletions likely rose in frequency through repeated positive selection rather than drift. Together, these results reveal that recurrent deletion events have repeatedly contributed to the evolution of cave-adapted phenotypes and highlight deletions as underexplored contributors of adaptive evolution in extreme environments.

Animals

Seasonal fluctuations in fitness result in severe reductions in effective population size.

Genetic evidence for fluctuating selection has begun to accumulate for different species over the past few decades, especially for the Drosophila genus where studies have reported hundreds of loci undergoing putatively adaptive oscillations across successive seasons. However, most theoretical and simulation studies of fluctuating selection have relied on abstract or weakly parameterized models, making it difficult to assess their relevance for natural populations. In this study, we simulate multilocus seasonally fluctuating selection acting on standing genetic variation under a recently developed model and examine its effect on the variance effective population size (Ne) at a genome-wide scale. By recapitulating genomic, demographic, and evolutionary parameters from natural Drosophila populations in our simulations, we were able to reproduce allele frequency oscillations reported in recent studies and show that these lead to ∼50% genome-wide reductions in Ne. We also demonstrate that Ne reductions are well predicted by the maximum frequency amplitude among all adaptively fluctuating loci, and that the frequency amplitudes are largely determined by the number of adaptively fluctuating loci and the strength of their epistatic interactions. Our results demonstrate that fluctuating selection can substantially reduce effective population size and underscore the importance of temporally variable selection in shaping genome-wide patterns of variation beyond classical models.

Drosophila melanogaster

Seasonal fluctuations in fitness result in severe reductions in effective population size.

Genetic evidence for fluctuating selection has begun to accumulate for different species over the past few decades, especially for the Drosophila genus where studies have reported hundreds of loci undergoing putatively adaptive oscillations across successive seasons. However, most theoretical and simulation studies of fluctuating selection have relied on abstract or weakly parameterized models, making it difficult to assess their relevance for natural populations. In this study, we simulate multilocus seasonally fluctuating selection under a recently developed model and examine its effect on the variance effective population size (Ne ) at a genome-wide scale. By recapitulating genomic, demographic, and evolutionary parameters from natural Drosophila populations in our simulations, we were able to reproduce allele frequency oscillations reported in recent studies and show that these lead to ~50% genome-wide reductions in Ne . We also demonstrate that Ne reductions are well predicted by the maximum frequency amplitude among all adaptively fluctuating loci, and that the frequency amplitudes are largely determined by the number of adaptively fluctuating loci and the strength of their epistatic interactions. Our results demonstrate that fluctuating selection can substantially reduce effective population size and underscore the importance of temporally variable selection in shaping genome-wide patterns of variation beyond classical models.

Drosophila melanogaster

Concurrent ecological and evolutionary processes contribute to mutualism breakdown between legumes and rhizobia.

Though they jointly shape community responses to environmental perturbations, ecology and evolution are often examined separately, even in microorganisms where both occur over short timescales. Here we examine ecological and evolutionary responses to 33 years of nitrogen fertilization using the legume-rhizobium mutualism. Pairing a manipulative inoculation study with full-length 16S rRNA gene amplicon sequencing and structural equation modeling allows us to synthesize across biological scales: whole bacterial community, genus Rhizobium, Rhizobium ASVs, and symbiosis plasmids. Clover's preferred partner decreases in N-addition soils, limiting host growth, while a diverse and largely uncharacterized Rhizobium community increases. This ecological change is compounded by a concurrent evolutionary degradation of symbiont partner quality via changing frequencies of symbiotic plasmids. Ecological (rarer symbionts) and evolutionary (inferior symbionts) processes each accounted for roughly half of this loss of host benefit, revealing that ecology and evolution jointly shape mutualism breakdown over the short timescales typical of microbial systems.

ecology

Zebrafish relatives as models for functional comparative genetics and genomics.

Closely related species, such as danionin fishes of the Danio, Danionella and Devario genera, often differ in their biology despite their shared evolutionary history, providing a platform for defining the molecular basis for the divergence of phenotypic traits. Such an approach requires the availability of large-scale genomic data, which have been provided by recent reports detailing the genomes of several danionins. Facilitated by the large number of genetic tools that are available for manipulation of the most studied member of this subgroup - the zebrafish, Danio rerio - the danionins have emerged as a useful comparative model system. Here we review their phylogeny and outline the phenotypic traits that are distinct to individual species or genera. We highlight how functional genetic tools such as interspecies hybridization, mutagenesis and transgenesis, as well as the recently reported genome assemblies, have enabled new avenues for hypothesis-driven and technology-driven exploration that collectively establish danionins as important genetic models for understanding a wide range of evolutionary innovations.

Journal Article

Evolutionary Reorganization of Transcriptomic Architecture Across a UVB Tolerance Gradient in Fish.

Environmental stressors such as ultraviolet radiation impose strong selective pressures on organisms, yet how adaptation to such stressors shapes transcriptomic responses at the network level remains poorly understood. Although stratospheric ozone is recovering globally, substantial regional variation in UV exposure persists, particularly in high-altitude environments where extreme UV levels can occur. Here, we compared three fish models representing distinct biological responses to UVB exposure: wild-type zebrafish (Danio rerio), a melanin-deficient zebrafish mutant (nacre) lacking a major protective mechanism against UVB damage, and the high-altitude Andean killifish Orestias ascotanensis, a species naturally exposed to extreme UVB radiation. Together, these models define a gradient spanning physiological protection, impaired protection, and evolutionary adaptation to UVB stress. Using RNA-seq and protein-protein interaction networks, we show that transcriptomic responses differ markedly across this gradient. Wild-type and nacre zebrafish exhibited relatively limited transcriptomic changes (∼2%-2.4% of genes changing), whereas O. ascotanensis displayed a large-scale and highly coordinated response (∼21.6% of genes changing) characterized by functionally specialized networks enriched in DNA repair pathways. These differences involved not only transcriptomic magnitude but also marked reorganization of transcriptomic architecture. Integration with positive selection analyses revealed that positively selected genes were concentrated within highly interconnected regions of transcriptomic networks, consistent with adaptation involving network reorganization. Furthermore, ortholog-based analyses suggest that adaptive responses involve differential reorganization of a conserved functional background. Together, our results support a model in which adaptation to environmental stress is associated with the reorganization of conserved transcriptomic networks across physiological and evolutionary contexts, providing a systems-level perspective on the molecular basis of adaptation.

UVB radiation

Design of highly functional genome editors by modelling CRISPR-Cas sequences.

Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology and human health. CRISPR-based gene editors derived from microorganisms, although powerful, often show notable functional tradeoffs when ported into non-native environments, such as human cells1. Artificial-intelligence-enabled design provides a powerful alternative with the potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models2 trained on biological diversity at scale, we demonstrate successful precision editing of the human genome with a programmable gene editor designed with artificial intelligence. To achieve this goal, we curated a dataset of more than 1 million CRISPR operons through systematic mining of 26 terabases of assembled genomes and metagenomes. We demonstrate the capacity of our models by generating 4.8× the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate that an artificial-intelligence-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 to facilitate broad, ethical use across research and commercial applications.

CRISPR-Cas Systems

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer

ScITree: Scalable Bayesian inference of transmission tree from epidemiological and genomic data.

Phylodynamic models capture joint epidemiological-evolutionary dynamics during an outbreak, providing a powerful tool to enhance understanding and management of disease transmission. Existing phylodynamic approaches, however, mostly rely on various non-mechanistic or semi-mechanistic approximations of the underlying epidemiological-evolutionary process. Previous work by Lau and colleagues has shown that full Bayesian mechanistic models, without relying on these approximations, can enable highly accurate joint inference of the epidemiological-evolutionary dynamics including the unobserved transmission tree. However, the Lau method faces major computational bottlenecks. As the volume of genomic data collected during outbreaks continues to grow, it is crucial to develop scalable yet accurate phylodynamic methods. Here we propose a new Bayesian phylodynamic model, overcoming the major scalability issue in the previous method and enabling a readily deployable, yet accurate, phylodynamic modeling framework. Specifically, we develop a scalable spatio-temporal phylodynamic framework for inferring the transmission tree (ScITree) and other key epidemiological parameters considering the infinite sites assumption in modeling mutation on the sequence level, in contrast to the Lau method in which mutation was modeled explicitly on the nucleotide level. Our approach features full Bayesian implementation utilizing an exact likelihood to mechanistically integrate epidemiological and evolutionary processes. We develop a computationally-efficient data-augmentation Markov Chain Monte Carlo algorithm, inferring key model parameters and unobserved dynamics including the transmission tree. We assess performance of our method using multiple simulated outbreak datasets. Our results indicate that our method can achieve high inference accuracy, comparable to the performance of the Lau method. Additionally, our method scales significantly more efficiently for large outbreaks, with computing time increasing linearly with outbreak size, compared to the exponential scaling of the Lau method. We also demonstrate our method's utility by applying our validated modeling framework to a dataset describing a foot-and-mouth disease outbreak in the UK. Our results show that our method is able to generate estimates of the transmission dynamics consistent with those from the prior method, further demonstrating the robustness of our new approach. In summary, our method provides a computationally-efficient, highly scalable, accurate modeling framework for inferring the joint spatio-temporal dynamics of epidemiological and evolutionary processes, facilitating timely and effective outbreak responses in space and time. Our method is implemented in our R package ScITree.

Bayes Theorem

Species Distribution Models Support Distinct and Non-Random Climatic Constraints on Globally Distributed Generalist Fungi.

Fungi play essential roles in ecosystems as pathogens, mutualists, and ubiquitous decomposers. However, like many important microbes, the spatial distribution of species and natural populations remains poorly understood compared to plants and animals. Many fungi are described as global generalists because they occur across wide geographic areas, but it remains unclear how and if these species are constrained by climate or geographic barriers. In this study, we used Species Distribution Models to infer the global climatic suitability of three common and globally distributed fungi: Aspergillus flavus, Penicillium chrysogenum and Aspergillus fumigatus. Models were constructed using global occurrence data from the Global Biodiversity Information Facility and were trained with Bioclimatic variables from the WorldClim dataset. All species' models showed high prediction fit, with predicted occurrence concentrated in the temperate and subtropical regions and broadly structured patterns. Each species showed distinct predicted distributions, but they displayed considerable spatial overlap on a global scale. Together, these results demonstrate that even apparently globally occurring and generalist fungal species occupy climatically structured niches. This study highlights the utility of SDMs and it provides a framework for future studies integrating ecological, genomics and evolutionary perspectives among the difficult to assess geographically widespread and common fungi.

comparative biogeography

Predicting dynamic expression patterns in budding yeast with a fungal DNA language model.

Predicting gene expression from DNA sequence remains challenging due to complex regulatory codes. We introduce a masked DNA language model pretrained on 165 fungal genomes closely related to budding yeast that captures conserved regulatory grammar. Fine-tuning the LM on yeast RNA-seq data-including high-resolution transcriptional regulator induction time courses generated in this study-yielded Shorkie, a model that substantially improves gene expression prediction compared to baselines trained without self-supervision. Shorkie identified canonical transcription factor (TF) binding motifs and tracked their usage across induction experiments. Furthermore, Shorkie accurately predicted variant effects, outperforming leading sequence-to-expression models in cis-eQTL classification and achieving high concordance with massively parallel reporter assays. Interpretability analyses revealed Shorkie's ability to resolve promoter dynamics, splicing signals, and temporal changes in regulatory motif usage. This framework demonstrates that evolutionary-scale pretraining combined with transfer learning substantially improves our ability to decode gene regulation from sequence, providing insights into noncoding variants and regulatory networks.

Journal Article

Introgression shapes the genomic conflict landscape of Malus, providing evidence for a reticulate backbone in a woody crop lineage.

Phylogenomic discordance is widespread across plants, but its evolutionary significance is often obscured when conflict is treated primarily as analytical noise rather than as evidence of underlying processes. In woody lineages in particular, incomplete lineage sorting, introgression, and genome duplication can interact over long timescales to produce complex genomic histories that are not adequately summarized by a strictly bifurcating tree. Here, we use Malus as a model woody genus to investigate how these processes structure conflict across a genus-scale, accession-based phylogenomic framework. Using broad taxon sampling, hundreds of nuclear loci, plastid genomes, and genome-wide SNP summaries, we reconstruct a robust nuclear backbone for sampled Malus lineages and evaluate where discordance is concentrated and which processes best explain it. Nuclear analyses resolve eight major clades, whereas conflict is non-random and localized to recurrent hotspots rather than evenly distributed across the tree. Cytonuclear discordance is similarly concentrated, especially around Clade H, represented by sampled accessions of M. tschonoskii, where localized plastid-nuclear disagreement is consistent with candidate plastid capture or organellar introgression. Multiple complementary analyses further indicate that the strongest conflict is not explained by ILS alone, but instead reflects lineage-structured introgression, while polyploid complexes represent additional localized sources of evolutionary complexity. Together, these results provide evidence for a reticulate genomic backbone in Malus and show how integrating nuclear, plastid, and genome-wide conflict analyses can help distinguish background discordance from process-specific signals in woody plant radiations. Several lineage-level reticulation hypotheses identified here should now be tested with broader population-level sampling and curated reference accessions.

Malus

The nature of extraversion: a genetical analysis.

A biometrical-genetical analysis of twin data to elucidate the determinants of variation in extraversion and its components, sociability and impulsiveness, revealed that both genetical and environmental factors contributed to variation in extraversion, to the variation and covariation of its component scales, and to the interaction between subjects and scales. A large environmental correlation between the scales suggested that environmental factors may predominate in determining the unitary nature of extraversion. The interaction between subjects and scales depended more on genetical factors, which suggests that the dual nature of extraversion has a strong genetical basis. A model assuming random mating, additive gene action, and specific environmental effects adequately describes the observed variation and covariation of sociability and impulsiveness. Possible evolutionary implications are discussed.

Adult