PubMed HealthSearch

SEARCH · PubMed Health

Results for “evolutionary scale model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning

REvolutionH-tl 2.0: A fast and robust tool for decoding evolutionary gene histories.

REvolutionH-tl is a fast, scalable, and integrated software platform for inferring orthology relationships, gene trees, species trees, and reconciled evolutionary scenarios directly from sequence data. Built upon the formal framework of best match graphs (BMGs), REvolutionH-tl predicts orthogroups and orthologous gene pairs with high accuracy, requiring neither precomputed trees nor multiple external tools. The software reconstructs event-labeled gene and species trees, seamlessly integrating reconciliation to produce fast, accurate, and biologically insightful evolutionary scenarios. Through extensive benchmarking on synthetic datasets with known ground truth, REvolutionH-tl outperforms or matches the accuracy of established tools such as OrthoFinder, Proteinortho, RAxML, GeneRax, and RANGER-DTL, while achieving significantly lower runtimes. A key innovation of REvolutionH-tl is its built-in support for detailed, publication-ready visualizations, which allow users to explore genome evolution dynamics, orthogroup composition, and reconciliation results with clarity and ease. These visual features position REvolutionH-tl as the first platform of its kind to combine analytical precision with intuitive interpretability. The software is open-source, cross-platform, and freely available at https://pypi.org/project/revolutionhtl/, providing a robust solution for large-scale evolutionary analyses in comparative genomics.

Software

Effects of social anxiety and facial expression on habituation of the electrodermal orienting response.

The present research examined electrodermal orienting to happy and angry faces as a function of social anxiety and threat of shock. A preliminary study using 569 undergraduate participants developed an adequate set of normative data of social anxiety for the Willoughby questionnaire (WQ) for use in subject selection. Electrodermal activity was measured in both high and low socially anxious subjects (N = 85) during exposure to 10 presentations of an angry face intermixed with 10 presentations of a happy face. Threat of shock (no-shock, shock work-up only, and shock work-up plus threat) was also manipulated. Skin conductance responses (SCRs) which occurred within 1-4 s of stimulus onset and trials-to-habituation constituted the data of primary interest. Although trials-to-habituation did not differ between angry and happy facial expressions, SCRs were larger to the angry face than to the happy face in both high and low socially anxious subjects. No differences in SCR magnitude were found as a function of threat of shock. The implications of these results for Ohman's functional-evolutionary model of social phobia are discussed, and alternative explanations in terms of prepotency and prior learning are examined.

Adolescent

Sequential gene loss promotes expansion of monophasic Salmonella Typhimurium ST34.

Understanding the genetic factors facilitating emergence of infectious diseases is critical, however, mechanisms underlying expansion of pathogenic bacterial variants remain unclear. Here we performed a large-scale genomic analysis of 44,597 Salmonella Typhimurium genomes and observe that sequential gene loss in monophasic Salmonella Typhimurium (mSTM) ST34 explains its clonal expansion as an increasingly prevalent zoonotic lineage. Functional and in vivo competition experiments show that a frameshift mutation in dinB, a polymerase for translesion DNA synthesis, leads to transcriptional changes affecting flagellin gene expression and subsequent loss of the flagellin-encoding fljB, altering the requirements for gut infection. Temporal evolutionary modelling supports a role for gene loss events in a specific chronological order for mSTM ST34 expansion. Our findings reveal a stepwise pathoadaptation model underpinning clonal global spread, providing mechanistic insights relevant to forecasting future pandemics.

Journal Article

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified ∼218,000 OGs supported by expression evidence, while ∼330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans

In genomes we trust: Assessing genomic reliability within the family Nectriaceae.

Reliable evolutionary inference increasingly depends on public genome resources, and the effects of uneven assembly quality, incomplete metadata, and biased taxonomic sampling remain poorly quantified. Using the species-rich fungal lineage Nectriaceae as a model system, we analysed 1530 genome sequence assemblies to assess metadata completeness, sampling representation, and genome quality. One-third of the assemblies lacked essential metadata, sequencing was heavily skewed toward a few agriculturally important lineages, and sampling of many genera was limited or nonexistent. BUSCO and QUAST metrics revealed substantial heterogeneity in assembly quality, with widespread fragmentation and numerous assemblies falling outside expected quality thresholds. From 763 single-copy orthologs identified in 576 higher-quality genomes, we reconstructed a phylogenomic backbone and quantified gene- and site-level concordance across the tree. Although major clades were broadly recovered, extensive gene-tree discordance and a polyphyletic Fusarium nisikadoi species complex revealed unresolved boundaries and conflict among loci. These results show how data quality, incomplete sampling, and discordant genomic histories can constrain phylogenomic resolution, and provide a general framework for improving comparative genomic resources and large-scale evolutionary inference.

Gene-tree discordance

First chromosome-level genome assembly of the colonial chordate model Botryllus schlosseri (Tunicata).

BACKGROUND: Botryllus schlosseri (Tunicata) is a colonial, laboratory model tunicate recognized for its remarkable developmental diversity, its regenerative abilities, and its peculiar genetically determined allorecognition system governed by a polymorphic locus controlling chimerism and cell parasitism. RESULTS: We report the first chromosome-level genome assembly of B. schlosseri subclade A1. By integrating long and short reads with Hi-C scaffolding, we produced both a phased diploid genome assembly and a conventional collapsed consensus sequence of 533 Mb. Of this total length, 96% belonged to 16 chromosome-scale scaffolds, with a BUSCO completeness score of 91.4%. We then compared our assembly with other high-quality tunicate genomes, revealing some synteny conservation but also extensive genomic rearrangements and a general loss of colinearity. CONCLUSIONS: The chromosome-level resolution of this assembly enhances our understanding of genome organization in colonial modular organisms. Comparative analyses highlight the dynamic nature of tunicate genomes, with conserved macrosynteny yet extensive microsyntenic rearrangements and scrambling, underscoring their rapid evolutionary trajectory. This high-quality genome assembly provides a valuable resource for exploring the unique biological features of colonial chordates, including their exceptional regenerative abilities and complex allorecognition system.

Animals

Transgenerational continuity: Persistence as a dimension of inheritance and evolution.

Transgenerational continuity (TC) describes the persistence of inherited molecular architectures across generations. Progress in identity-by-descent (IBD) detection, recombination dynamics, and epigenetic research highlights the growing need for a more comprehensive model of inheritance. This theoretical framework synthesizes evidence from genomics, population studies, and epigenetics to outline how inherited molecular architectures, which are transmitted through IBD, together with heritable epigenetic modifications, can preserve ancestral information across generations. IBD captures genomic continuity across three nested scales, where recent familial segments link close relatives, population-level haplotypes are shared across cohorts, and archaic fragments from Neanderthal and Denisovan admixture persist as molecular fossils of ancient lineages. Although recombination and selection reshape these regions, their persistence across time scales highlights the evolutionary durability of genomic continuity. Epigenetic memory reflects regulatory persistence, whereby molecular modifications can preserve functional states across cell divisions and sometimes across generations. Together with familial and population-level IBD persistence and the long-term retention of introgressed haplotypes, these findings demonstrate that inherited molecular architectures can persist across multiple timescales. Evolutionary processes shape this persistence. Purifying selection preferentially removes deleterious inherited variants, whereas positive selection can favor the persistence of functionally relevant genomic architectures. From this perspective, evolutionary dynamics arise not only from the generation of variation, but also from the differential persistence of inherited molecular architectures through selection. Transgenerational continuity therefore provides a conceptual framework in which persistence serves as an explanatory dimension of inheritance and evolution that complements variation and explains the persistence of biological identity across generations and evolutionary time.

Biological identity

The uncertainty principle as an evolutionary engine.

Heisenberg's uncertainty principle in quantum mechanics underlies the genesis of evolutionary variability. When the uncertainty principle is coupled with the incontrovertible principle of the conservation of energy and material resources, there appears an uncertainty relationship between local fluctuations in the quantities to be conserved on a global scale and the rate of their local variation. Since the local fluctuations are accompanied by the non-vanishing rate of variation because of the uncertainty relationship, they generate subsequent fluctuations. Generativity latent in the uncertainty relationship is non-random and ubiquitous all through various evolutionary stages from abiotic synthesis of monomers and polymers up to the emergence of behavior-induced variability of organisms.

Amino Acid Sequence

Translating functional molecular knowledge into crop-breeding success.

Historical plant breeding, which optimizes phenotypes through selective crossing guided by phenotypic evaluation and molecular markers, is limited by evolutionary constraints that hinder rapid crop improvement. A new paradigm, precision breeding, circumvents these limitations by targeting genetic variants through functional molecular knowledge. To generate this knowledge at scale, sequence-based deep learning leverages high-quality genome sequence data to predict variant effects at base-pair resolution. When linked to agronomically important traits, these predictions enable breeders to prioritize variants for precision selection or editing. Although it is still in the early stages of development, we foresee three key applications for this approach: introgressing genes from distant breeding pools, purging deleterious mutations and designing new plant ideotypes. Looking ahead, refined computational models will facilitate targeted editing and the systematic redesign of complex physiological processes to address emerging breeding goals under shifting environmental conditions.

Crops, Agricultural

[Syphilis and human treponemes: a long evolutionary history revealed by paleogenomics].

Recent discoveries in paleogenomics have revolutionized our understanding of syphilis and other human treponematoses. Far from being a pathogen that suddenly appeared in Europe in the late Middle Ages, we now know that Treponema pallidum has been circulated among human populations for millennia. Ancient genomes recovered from pre-Columbian contexts in the Americas show that major treponemal lineages had already diversified well before the modern era, often in the absence of recognizable skeletal lesions. Genomic analyses further indicate that treponemal diversity is not the result of extensive genetic acquisition, but rather of small-scale modulation of a highly conserved genome, notably via antigenic variation involving the tpr gene family. Combined with data on endemic treponematoses, congenital syphilis, and historical pathology collections, these findings support a model in which syphilis, yaws, and bejel represent context-dependent expressions of an ancient treponemal continuum, with implications for diagnosis, epidemiology, and vaccine design.

Humans

ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation.

MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).

Semantics

Estimation of demography and mutation rates from one million haploid genomes.

As genetic sequencing costs have plummeted, datasets with sizes previously unthinkable have begun to appear. Such datasets present opportunities to learn about evolutionary history, particularly via rare alleles that record the very recent past. However, beyond the computational challenges inherent in the analysis of many large-scale datasets, large population-genetic datasets present theoretical problems. In particular, the majority of population-genetic tools require the assumption that each mutant allele in the sample is the result of a single mutation (the "infinite-sites" assumption), which is violated in large samples. Here, we present DR EVIL, a method for estimating mutation rates and recent demographic history from very large samples. DR EVIL avoids the infinite-sites assumption by using a diffusion approximation to a branching-process model with recurrent mutation. This approach results in tractable likelihoods that are accurate for rare alleles. We show that DR EVIL performs well in simulations and apply it to rare-variant data from one million haploid samples. We identify mutation-rate heterogeneity even after accounting for trinucleotide context and methylation status. We also predict that at modern sample sizes, the alleles at most polymorphic sites with high mutation rates represent the descendants of multiple mutation events.

Haploidy

Deciphering the mosaic genome of sugarcane cultivars through polyploid admixture inference with AdmixPoly.

BACKGROUND: Characterizing population structure and admixture events between ancestral groups plays a key role in understanding the evolutionary history of species and crops. Most tools for inferring admixture have been developed for diploids and are not suitable for polyploids, in particular those with high and mixed ploidy such as Saccharum. RESULTS: Here we present AdmixPoly, an R-package designed to infer admixture in polyploid species both at the genome-wide scale and locally along chromosomes. We compare AdmixPoly with state-of-the-art methods using simulations, demonstrating its precision and computational efficiency. Notably, local admixture inference in complex scenarios, such as high ploidy levels, large numbers of ancestral groups and alleles per marker is enabled through efficient approximations of emission and transition probabilities within a hidden Markov model framework. We apply this approach to characterize the contributions of wild Saccharum species to the complex polyploid genome of modern sugarcane cultivars. A panel of wild and cultivated Saccharum accessions is genotyped for 80K genomic regions, each revealing approximately 50 read-scale haplotypes. CONCLUSIONS: The results reveal that most of the approximately 12 copies of each basic chromosome in modern cultivars are derived from the domesticated species Saccharum officinarum, with one to four copies typically contributed by distinct subgroups of the wild species Saccharum spontaneum. In addition, contributions from an unknown wild Saccharum group originating from the Pacific were identified in most cultivars. The conserved pattern of these introgressions suggests that they can be traced back to the early stages of sugarcane breeding approximately a century ago.

Saccharum

The sociobiology of bereavement: a reply to Littlefield and Rushton.

This article offers a critique of Littlefield and Rushton's (1986) application of sociobiological principles to bereavement following the death of a child. The following general issues are considered: (a) whether behavior is always adaptive and (b) the distinction between proximate and ultimate explanations. It is argued that grief is a maladaptive by-product of another, adaptive feature and that hypotheses about the severity of grief are best derived from proximate considerations rather than genetic relatedness. The use of a single-item rating scale to measure grief is questioned, and it is noted that interspouse reliabilities reported in the article were low, a problem not solved (as claimed) by aggregation. Criticisms are made of the specific hypotheses, notably in terms of their origins in sociobiological theory. It is argued that functional hypotheses are not alternatives to proximate mechanisms, but enable some proximate mechanisms to be viewed from the perspective of evolutionary biology.

Adaptation, Biological

Dissecting fluctuating selection: A unified population and quantitative genetics framework.

One of the longstanding debates in evolutionary biology is the effect of fluctuating selection on genetic changes in populations. However, the extent to which these periodic forces influence organisms at both genomic and phenotypic levels remains unclear. Despite the compelling evidence of fluctuating selection from recent studies, there is a disconnect between empirical and theoretical findings concerning the underlying mechanisms due to the limited evidence regarding the scale and processes that generate genome-wide oscillations. This study aims to elucidate how both genetic factors (e.g. heritability, number of causative loci) and ecological factors (e.g. season length, the difference in the phenotypic optima between seasons, population size dynamics) drive fluctuating selection and to identify the parameters that produce consistent oscillatory patterns. We developed a modeling framework integrating quantitative and population genetics to simulate a population under various selection regimes. We applied spectral analysis to detect periodicity, indicating cyclical selective environments. Our simulations highlight the conditions sustaining oscillations in allele frequencies over time. Spectral analysis successfully identifies the periodic patterns from allele frequency trajectories, even under highly complex selection regimes. Not only does our study clarify the conditions that yield oscillatory behaviors, but these parameters can also potentially be estimated in natural populations, providing a possibility of empirically testing these models.

Fluctuating selection

Dissecting fluctuating selection: A unified population and quantitative genetics framework.

One of the longstanding debates in evolutionary biology is the effect of fluctuating selection on genetic changes in populations. However, the extent to which these periodic forces influence organisms at both genomic and phenotypic levels remains unclear. Despite the compelling evidence of fluctuating selection from recent studies, there is a disconnect between empirical and theoretical findings concerning the underlying mechanisms due to the limited evidence regarding the scale and processes that generate genome-wide oscillations. This study aims to elucidate how both genetic factors (e.g. heritability, number of causative loci) and ecological factors (e.g. season length, the difference in the phenotypic optima between seasons, population size dynamics) drive fluctuating selection and to identify the parameters that produce consistent oscillatory patterns. We developed a modeling framework integrating quantitative and population genetics to simulate a population under various selection regimes. We applied spectral analysis to detect periodicity, indicating cyclical selective environments. Our simulations highlight the conditions sustaining oscillations in allele frequencies over time. Spectral analysis successfully identifies the periodic patterns from allele frequency trajectories, even under highly complex selection regimes. Not only does our study clarify the conditions that yield oscillatory behaviors, but these parameters can also potentially be estimated in natural populations, providing a possibility of empirically testing these models.

Fluctuating selection

Evolution of ribosomal RNA gene copy number on the sex chromosomes of Drosophila melanogaster.

A diverse array of cellular and evolutionary forces--including unequal crossing-over, magnification, compensation, and natural selection--is at play modulating the number of copies of ribosomal RNA (rRNA) genes on the X and Y chromosomes of Drosophila. Accurate estimates of naturally occurring distributions of copy numbers on both the X and Y chromosomes are needed in order to explore the evolutionary end result of these forces. Estimates of relative copy numbers of the ribosomal DNA repeat, as well as of the type I and type II inserts, were obtained for a series of 96 X chromosomes and 144 Y chromosomes by using densitometric measurements of slot blots of genomic DNA from adult D. melanogaster bearing appropriate deficiencies that reveal chromosome-specific copy numbers. Estimates of copy number were put on an absolute scale with slot blots having serial dilutions both of the repeat and of genomic DNA from nonpolytene larval brain and imaginal discs. The distributions of rRNA copy number are decidedly skewed, with a long tail toward higher copy numbers. These distributions were fitted by a population genetic model that posits three different types of exchange events--sister-chromatid exchange, intrachromatid exchange, and interchromosomal crossing-over. In addition, the model incorporates natural selection, because experimental evidence shows that there is a minimum number of functional elements necessary for survival. Adequate fits of the model were found, indicating that either natural selection also eliminates chromosomes with high copy number or that the rate of intrachromatid exchange exceeds the rate of interchromosomal exchange.

Animals