PubMed HealthSearch

SEARCH · PubMed Health

Results for “evolutionary scale model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

ScITree: Scalable Bayesian inference of transmission tree from epidemiological and genomic data.

Phylodynamic models capture joint epidemiological-evolutionary dynamics during an outbreak, providing a powerful tool to enhance understanding and management of disease transmission. Existing phylodynamic approaches, however, mostly rely on various non-mechanistic or semi-mechanistic approximations of the underlying epidemiological-evolutionary process. Previous work by Lau and colleagues has shown that full Bayesian mechanistic models, without relying on these approximations, can enable highly accurate joint inference of the epidemiological-evolutionary dynamics including the unobserved transmission tree. However, the Lau method faces major computational bottlenecks. As the volume of genomic data collected during outbreaks continues to grow, it is crucial to develop scalable yet accurate phylodynamic methods. Here we propose a new Bayesian phylodynamic model, overcoming the major scalability issue in the previous method and enabling a readily deployable, yet accurate, phylodynamic modeling framework. Specifically, we develop a scalable spatio-temporal phylodynamic framework for inferring the transmission tree (ScITree) and other key epidemiological parameters considering the infinite sites assumption in modeling mutation on the sequence level, in contrast to the Lau method in which mutation was modeled explicitly on the nucleotide level. Our approach features full Bayesian implementation utilizing an exact likelihood to mechanistically integrate epidemiological and evolutionary processes. We develop a computationally-efficient data-augmentation Markov Chain Monte Carlo algorithm, inferring key model parameters and unobserved dynamics including the transmission tree. We assess performance of our method using multiple simulated outbreak datasets. Our results indicate that our method can achieve high inference accuracy, comparable to the performance of the Lau method. Additionally, our method scales significantly more efficiently for large outbreaks, with computing time increasing linearly with outbreak size, compared to the exponential scaling of the Lau method. We also demonstrate our method's utility by applying our validated modeling framework to a dataset describing a foot-and-mouth disease outbreak in the UK. Our results show that our method is able to generate estimates of the transmission dynamics consistent with those from the prior method, further demonstrating the robustness of our new approach. In summary, our method provides a computationally-efficient, highly scalable, accurate modeling framework for inferring the joint spatio-temporal dynamics of epidemiological and evolutionary processes, facilitating timely and effective outbreak responses in space and time. Our method is implemented in our R package ScITree.

Bayes Theorem

Species Distribution Models Support Distinct and Non-Random Climatic Constraints on Globally Distributed Generalist Fungi.

Fungi play essential roles in ecosystems as pathogens, mutualists, and ubiquitous decomposers. However, like many important microbes, the spatial distribution of species and natural populations remains poorly understood compared to plants and animals. Many fungi are described as global generalists because they occur across wide geographic areas, but it remains unclear how and if these species are constrained by climate or geographic barriers. In this study, we used Species Distribution Models to infer the global climatic suitability of three common and globally distributed fungi: Aspergillus flavus, Penicillium chrysogenum and Aspergillus fumigatus. Models were constructed using global occurrence data from the Global Biodiversity Information Facility and were trained with Bioclimatic variables from the WorldClim dataset. All species' models showed high prediction fit, with predicted occurrence concentrated in the temperate and subtropical regions and broadly structured patterns. Each species showed distinct predicted distributions, but they displayed considerable spatial overlap on a global scale. Together, these results demonstrate that even apparently globally occurring and generalist fungal species occupy climatically structured niches. This study highlights the utility of SDMs and it provides a framework for future studies integrating ecological, genomics and evolutionary perspectives among the difficult to assess geographically widespread and common fungi.

comparative biogeography

Predicting dynamic expression patterns in budding yeast with a fungal DNA language model.

Predicting gene expression from DNA sequence remains challenging due to complex regulatory codes. We introduce a masked DNA language model pretrained on 165 fungal genomes closely related to budding yeast that captures conserved regulatory grammar. Fine-tuning the LM on yeast RNA-seq data-including high-resolution transcriptional regulator induction time courses generated in this study-yielded Shorkie, a model that substantially improves gene expression prediction compared to baselines trained without self-supervision. Shorkie identified canonical transcription factor (TF) binding motifs and tracked their usage across induction experiments. Furthermore, Shorkie accurately predicted variant effects, outperforming leading sequence-to-expression models in cis-eQTL classification and achieving high concordance with massively parallel reporter assays. Interpretability analyses revealed Shorkie's ability to resolve promoter dynamics, splicing signals, and temporal changes in regulatory motif usage. This framework demonstrates that evolutionary-scale pretraining combined with transfer learning substantially improves our ability to decode gene regulation from sequence, providing insights into noncoding variants and regulatory networks.

Journal Article

Introgression shapes the genomic conflict landscape of Malus, providing evidence for a reticulate backbone in a woody crop lineage.

Phylogenomic discordance is widespread across plants, but its evolutionary significance is often obscured when conflict is treated primarily as analytical noise rather than as evidence of underlying processes. In woody lineages in particular, incomplete lineage sorting, introgression, and genome duplication can interact over long timescales to produce complex genomic histories that are not adequately summarized by a strictly bifurcating tree. Here, we use Malus as a model woody genus to investigate how these processes structure conflict across a genus-scale, accession-based phylogenomic framework. Using broad taxon sampling, hundreds of nuclear loci, plastid genomes, and genome-wide SNP summaries, we reconstruct a robust nuclear backbone for sampled Malus lineages and evaluate where discordance is concentrated and which processes best explain it. Nuclear analyses resolve eight major clades, whereas conflict is non-random and localized to recurrent hotspots rather than evenly distributed across the tree. Cytonuclear discordance is similarly concentrated, especially around Clade H, represented by sampled accessions of M. tschonoskii, where localized plastid-nuclear disagreement is consistent with candidate plastid capture or organellar introgression. Multiple complementary analyses further indicate that the strongest conflict is not explained by ILS alone, but instead reflects lineage-structured introgression, while polyploid complexes represent additional localized sources of evolutionary complexity. Together, these results provide evidence for a reticulate genomic backbone in Malus and show how integrating nuclear, plastid, and genome-wide conflict analyses can help distinguish background discordance from process-specific signals in woody plant radiations. Several lineage-level reticulation hypotheses identified here should now be tested with broader population-level sampling and curated reference accessions.

Malus

The nature of extraversion: a genetical analysis.

A biometrical-genetical analysis of twin data to elucidate the determinants of variation in extraversion and its components, sociability and impulsiveness, revealed that both genetical and environmental factors contributed to variation in extraversion, to the variation and covariation of its component scales, and to the interaction between subjects and scales. A large environmental correlation between the scales suggested that environmental factors may predominate in determining the unitary nature of extraversion. The interaction between subjects and scales depended more on genetical factors, which suggests that the dual nature of extraversion has a strong genetical basis. A model assuming random mating, additive gene action, and specific environmental effects adequately describes the observed variation and covariation of sociability and impulsiveness. Possible evolutionary implications are discussed.

Adult

Structure of yeast triosephosphate isomerase at 1.9-A resolution.

The structure of yeast triosephosphate isomerase (TIM) has been solved at 3.0-A resolution and refined at 1.9-A resolution to an R factor of 21.0%. The final model consists of all non-hydrogen atoms in the polypeptide chain and 119 water molecules, a number of which are found in the interior of the protein. The structure of the active site clearly indicates that the carboxylate of the catalytic base, Glu 165, is involved in a hydrogen-bonding interaction with the hydroxyl of Ser 96. In addition, the interactions of the other active site residues, Lys 12 and His 95, are also discussed. For the first time in any TIM structure, the "flexible loop" has well-defined density; the conformation of the loop in this structure is stabilized by a crystal contact. Analysis of the subunit interface of this dimeric enzyme hints at the source of the specificity of one subunit for another and allows us to estimate an association constant of 10(14)-10(16) M-1 for the two monomers. The analysis also suggests that the interface may be a particularly good target for drug design. The conserved positions (20%) among sequences from 13 sources ranging on the evolutionary scale from Escherichia coli to humans reveal the intense pressure to maintain the active site structure.

Amino Acid Sequence

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Mutation-selection balance and the evolutionary advantage of sex and recombination.

Mutation-selection balance in a multi-locus system is investigated theoretically, using a modification of Bulmer's infinitesimal model of selection on a normally-distributed quantitative character, taking the number of mutations per individual (n) to represent the character value. The logarithm of the fitness of an individual with n mutations is assumed to be a quadratic, decreasing function of n. The equilibrium properties of infinitely large asexual populations, random-mating populations lacking genetic recombination, and random-mating populations with arbitrary recombination frequencies are investigated. With 'synergistic' epistasis on the scale of log fitness, such that log fitness declines more steeply as n increases, it is shown that equilibrium mean fitness is least for asexual populations. In sexual populations, mean fitness increases with the number of chromosomes and with the map length per chromosome. With 'diminishing returns' epistasis, such that log fitness declines less steeply as n increases, mean fitness behaves in the opposite way. Selection on asexual variants and genes affecting the rate of genetic recombination in random-mating populations was also studied. With synergistic epistasis, zero recombination always appears to be disfavoured, but free recombination is disfavoured when the mutation rate per genome is sufficiently small, leading to evolutionary stability of maps of intermediate length. With synergistic epistasis, an asexual mutant is unlikely to invade a sexual population if the mutation rate per diploid genome greatly exceeds unity. Recombination is selectively disadvantageous when there is diminishing returns epistasis. These results are compared with the results of previous theoretical studies of this problem, and with experimental data.

Animals

Expanding the human proteome with microproteins and peptideins.

A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1-3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of 'peptideins' as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome.

Humans

Outer membrane changes enable evolutionary escape from bacterial predation.

Antimicrobial resistance (AMR) is a threat to modern medicine. To combat AMR pathogens, natural predators like bacteriophages and predatory bacteria have gained interest recently. Predatory bacterium Bdellovibrio bacteriovorus is ubiquitous and has a broad prey range. It is particularly potent at killing many AMR Gram-negative bacterial pathogens featured on the WHO priority list. However, it is currently unclear whether prey bacteria can evolve genetically-determined resistance against predation by B. bacteriovorus. Here, we show that the model bacterium Escherichia coli K-12 consistently evolves resistance against B. bacteriovorus during experimental evolution. Selection for resistance scaled positively with predation pressure and was widespread after two cycles of predator exposure. Similar to antibiotics, predation resistance was costly, manifesting in a trade-off between predation resistance and fitness in the absence of predators. Genetic analysis combined with proteomics identified mutations that lead to the down-regulation of the outer membrane porin OmpF as a common resistance mechanism. In addition, a rarer mutation in cell envelope lipopolysaccharide-modifying enzyme WaaF also conferred predation resistance, likely by a pleiotropic effect, which included OmpF down regulation. While our study uncovers evolutionary and mechanistic aspects of prey escape from predation, it also highlights that the high cost of resistance reflects a handicap for the pathogen and can thus be exploited to increase treatment sustainability. Altogether, our work generates essential knowledge in ecologically important predator-prey interactions and can advance predatory bacteria as "living antibiotics" to combat AMR.

Bdellovibrio bacteriovorus

Neuronal models of cognitive functions.

Understanding the neural bases of cognition has become a scientifically tractable problem, and neurally plausible models are proposed to establish a causal link between biological structure and cognitive function. To this end, levels of organization have to be defined within the functional architecture of neuronal systems. Transitions from any one of these interacting levels to the next are viewed in an evolutionary perspective. They are assumed to involve: (1) the production of multiple transient variations and (2) the selection of some of them by higher levels via the interaction with the outside world. The time-scale of these "evolutions" is expected to differ from one level to the other. In the course of development and in the adult this internal evolution is epigenetic and does not require alteration of the structure of the genome. A selective stabilization (and elimination) of synaptic connections by spontaneous and/or evoked activity in developing neuronal networks is postulated to contribute to the shaping of the adult connectivity within an envelope of genetically encoded forms. At a higher level, models of mental representations, as states of activity of defined populations of neurons, are discussed in terms of statistical physics, and their storage is viewed as a process of selection among variable and transient pre-representations. Theoretical models illustrate that cognitive functions such as short-term memory and handling of temporal sequences may be constrained by "microscopic" physical parameters. Finally, speculations are offered about plausible neuronal models and selectionist implementations of intentions.

Animals

Volition, deception, and the evolution of justice.

Criminal justice is inextricably associated with the attributive concept of volition. Although the voluntary-involuntary distinction is subjectively vivid, causal research shows its poles to be inseparable, i.e., the dichotomy is deceptive. Why a bulwark of civilization should be founded on paradox, may be clarified by examining the role of self-deception in man's evolutionary heritage. Natural selection for an optimal degree of self-deception probably occurred, both to facilitate deception of others and to foster human cooperation. This contributed to the evolution of psychiatric disorders, the voluntary-involuntary continuum, and large scale social systems. Society and its members reach an equilibrium within the truth-deception continuum, manifest in individuals by conscious versus unconscious and voluntary versus involuntary, and in society by tension between what actually occurs (realism) and its organizing ideals (idealism). Three legal models of criminal justice are understood in this context: The (1) utilitarian, most realistic, is essential to social survival but vulnerable to abuse; (2) rehabilitative, at an opposite idealistic pole, better supports the image of social beneficence that helps to bind society's members; (3) retributive, most heavily grounded in volition, puts greater emphasis on individual autonomy, and reciprocally modulates the other models. All are legitimized by evolutionary traditions that antedate homo sapiens, and none is sufficient in itself. Elements of all three models necessarily coexist within any existing society, their relative strength varying with its collective values, prosperity, and perceived safety.

Criminal Law

Size and shape of the cerebral cortex in mammals. II. The cortical volume.

The geometry of the brain and cerebral cortex in mammals has been studied from an evolutionary perspective and is described in mathematical terms. The volume of the cerebral cortex, in contrast to the cortical surface area, scales to brain volume in a similar way, irrespective of the degree of cortical folding. Among mammals, Cetacea form a subgroup, in that their volumetric data fit an isometric model better than an allometric model. An index of corticalization is presented which contains information about both the mass of interconnective nerve fibers and the degree of intracortical processing. It is shown, furthermore, that a semilogarithmic equation appropriately describes the relationship between mean cortical thickness and brain volume. Finally, allometric equations between brain volume and cortical parameters, which can be used for predictive purposes, are presented.

Animals

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

The evolutionary origin of feathers.

Previous theories relating the origin of feathers to flight or to heat conservation are considered to be inadequate. There is need for a model of feather evolution that gives attention to the function and adaptive advantage of intermediate structures. The present model attempts to reveal and to deal with, the spectrum of complex questions that must be considered. In several genera of modern lizards, scales are elongated in warm climates. It is argued that these scales act as small shields to solar radiation. Experiments are reported that tend to confirm this. Using lizards as a conceptual model, it is argued that feathers likewise arose as adaptations to intense solar radiation. Elongated scales are assumed to have subdivided into finely branched structures that produced a heat-shield, flexible as well as long and broad. Associated muscles had the function of allowing the organism fine control over rates of heat gain and loss: the specialized scales or early feathers could be moved to allow basking in cool weather or protection in hot weather. Subdivision of the scales also allowed a close fit between the elements of the insulative integument. There would have been mechanical and thermal advantages to having branches that interlocked into a pennaceous structure early in evolution, so the first feathers may have been pennaceous. A versatile insulation of movable, branched scales would have been a preadaptation for endothermy. As birds took to the air they faced cooling problems despite their insulative covering because of high convective heat loss. Short glides may have initially been advantageous in cooling an animal under heat stress, but at some point the problem may have shifted from one of heat exclusion to one of heat retention. Endothermy probably evolved in conjunction with flight. If so, it is an unnecessary assumption to postulate that the climate cooled and made endothermy advantageous. The development of feathers is complex and a model is proposed that gives attention to the fundamental problems of deriving a branched structure with a cylindrical base from an elongated scale.

Animals

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

Coevolution to the edge of chaos: coupled fitness landscapes, poised states, and coevolutionary avalanches.

We introduce a broadened framework to study aspects of coevolution based on the NK class of statistical models of rugged fitness landscapes. In these models the fitness contribution of each of N genes in a genotype depends epistatically on K other genes. Increasing epistatic interactions increases the rugged multipeaked character of the fitness landscape. Coevolution is thought of, at the lowest level, as a coupling of landscapes such that adaptive moves by one player deform the landscapes of its immediate partners. In these models we are able to tune the ruggedness of landscapes, how richly intercoupled any two landscapes are, and how many other players interact with each player. All these properties profoundly alter the character of the coevolutionary dynamics. In particular, these parameters govern how readily coevolving ecosystems achieve Nash equilibria, how stable to perturbations such equilibria are, and the sustained mean fitness of coevolving partners. In turn, this raises the possibility that an evolutionary metadynamics due to natural selection may sculpt landscapes and their couplings to achieve coevolutionary systems able to coadapt well. The results suggest that sustained fitness is optimized when landscape ruggedness relative to couplings between landscapes is tuned such that Nash equilibria just tenuously form across the ecosystem. In this poised state, coevolutionary avalanches appear to propagate on all length scales in a power law distribution. Such avalanches may be related to the distribution of small and large extinction events in the record.

Biological Evolution

Collective posterior inference from highly variable empirical replicates.

High-throughput experimental platforms now routinely generate data from dozens or hundreds of independent observations. Simulation-based inference (SBI) offers a powerful framework for estimating model parameters from such complex datasets, but standard methods struggle to scale to the noisy multiple-replicates regime without incurring prohibitive computational costs or careful hyperparameter tuning. Here, we introduce a new method for fast and robust collective posterior inference from multiple independent replicates using a robust product-of-experts aggregation scheme that automatically mitigates the influence of outliers. Evaluating it on synthetic and empirical evolutionary datasets, we find it achieves state-of-the-art estimation accuracy and computational efficiency, including inference from noisy observations. Our method is compatible with any SBI framework, providing a scalable, plug-and-play solution for inference from noisy multiple-replicate datasets.

Computational Biology