PubMed HealthSearch

SEARCH · PubMed Health

Results for “demographic inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Two blind spots in the demographic inference of human origins from genomic data.

Ancient DNA and new inference methods have transformed the study of human origins, but consensus has not followed. Evidence increasingly indicates that hominin populations were pervasively structured and admixed, so complexity rather than simplicity is the appropriate prior. Here I highlight two blind spots that impede resolving that complexity. First, every inference passes through summaries of the data, and those summaries bound what can be recovered. Second, the space of candidate models is vast, yet competing model classes are rarely fit to common data, so a reported best model carries little evidence about untested model classes. This second blind spot reflects practice rather than data. It can be narrowed by testing competing models against withheld summaries and by reporting the models that were tried and rejected rather than only the winner.

Journal Article

Inferring the demographic history of Chinese and Indian rhesus macaque (Macaca mulatta) populations from PacBio HiFi long-read sequencing data.

The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time ∼140,000 generations ago from an ancestral population of ∼65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of ∼220,000 individuals in the Chinese population and ∼14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

Cercopithecidae

Faster inference of complex demographic models from large allele frequency spectra.

MOTIVATION: Demographic inference from the joint site frequency spectrum is limited by computation when many populations or many samples are analyzed. RESULTS: We present momi3, a JAX-based method for inferring complex demographic models from large allele frequency spectra. It supports continuous migration, GPU execution, automatic differentiation, standardized demographic model input, and genealogical pruning. These changes yield speedups up to 1000× over existing methods and enable analysis of archaic admixture models using hundreds of human genomes. AVAILABILITY AND IMPLEMENTATION: momi3 is implemented in Python/JAX as part of demestats. Source code is available at https://github.com/jthlab/demestats; documentation is available at https://demestats.readthedocs.org.

Humans

Accessible, realistic genome simulation with selection using stdpopsim.

Selection is a fundamental evolutionary force that shapes patterns of genetic variation across species. However, simulations incorporating realistic selection along heterogeneous genomes in complex demographic histories are challenging, limiting our ability to benchmark statistical methods aimed at detecting selection and to explore theoretical predictions. stdpopsim is a community-maintained simulation library that already provides an extensive catalog of species-specific population genetic models. Here we present a major extension to the stdpopsim framework that enables simulation of various modes of selection, including background selection, selective sweeps, and arbitrary distributions of fitness effects (DFE) acting on annotated subsets of the genome (for instance, exons). This extension maintains stdpopsim's core principles of reproducibility and accessibility while adding support for species-specific genomic annotations and published DFE estimates. We demonstrate the utility of this framework by comparing methods for demographic inference, DFE estimation, and selective sweep detection across several species and scenarios. Our results demonstrate the robustness of demographic inference methods to selection on linked sites, reveal the sensitivity of DFE-inference methods to model assumptions, and show how genomic features, like recombination rate and functional sequence density, influence power to detect selective sweeps. This extension to stdpopsim provides a powerful new resource for the population genetics community to explore the interplay between selection and other evolutionary forces in a reproducible, user-friendly framework.

Journal Article

Pleistocene island connectivity did not enhance dispersal or impact population size change in Galápagos geckos.

Patterns of biodiversity on remote archipelagos are largely shaped by intra-archipelago colonization followed by in situ diversification. Pleistocene sea-level fluctuations purportedly enhanced gene flow among terrestrial organisms by increasing connectivity during periods of lower sea level. Furthermore, changes in sea-level are hypothesized to impact population sizes as a result of fluctuations in island sizes. Here, we used genomic data to test the role of Pleistocene island connectivity on the diversification and demographics of leaf-toed geckos (Phyllodactylus) endemic to the Galápagos. Consistent with previous studies, we found that present diversity of Galápagos Phyllodactylus stems from three independent dispersal events. Contrary to the hypothesis of Pleistocene-driven diversification, we found no correspondence between lineage divergence and island connectivity. Furthermore, we found no evidence of introgression; demographic modelling indicated that all species increased rapidly in effective population size in the period 20-150 ka, and these inferred demographic expansions were largely asynchronous and apparently unassociated with species or island age. Collectively, these results indicate that more complex abiotic and/or biotic factors may better explain the recent demographic history of Phyllodactylus and underscore the need for additional population genomic studies of terrestrial taxa to understand the impact of past climate cycles on Galápagos island communities.

Animals

Homoploid Hybrid Speciation in a Marine Pelagic Fish.

Homoploid hybrid speciation (HHS) is an enigmatic evolutionary process where new species arise through hybridisation of divergent lineages without changes in chromosome number. Although increasingly documented in various taxa and ecosystems, convincing cases of HHS in marine fishes have been lacking. This study presents a possible case of HHS in a pelagic marine fish based on comprehensive genomic, morphological, and ecological analyses. Population genomics, species tree estimation, and tests of introgression and admixture identified three sympatric clusters in Megalaspis cordyla in the western Pacific and the admixed nature of one cluster between the others. Moreover, model-based demographic inference favoured a hybrid speciation scenario over introgression for the origin of the admixed cluster. While contemporary gene flow suggested partial reproductive isolation, examination of occurrence data and ecologically relevant morphological characters suggested ecological differences between the clusters, potentially contributing to the reproductive isolation and niche partitioning in sympatry. The clusters are also morphologically distinguishable and thus can be taxonomically recognised as separate species. The hybrid cluster is restricted to the coasts of Taiwan and Japan, where all three clusters coexist. The parental clusters are additionally found in lower latitudes, where they display non-overlapping distributions. Given the geographical distributions, estimated times of species formation, and patterns of historical demographic changes, we propose that the Pleistocene glacial cycles were the primary driver of HHS in this system. We also develop an ecogeographic model of HHS in marine coastal ecosystems, including a novel hypothesis to explain the initial stages of HHS.

Animals

ADAMIXTURE: adaptive first-order optimization for biobank-scale genetic clustering.

MOTIVATION: Estimating genetic clusters from sequencing data is a fundamental task in population and medical genetics, enabling demographic inference and adjustment for population structure in association studies. ADMIXTURE, a widely used model-based clustering method, employs an accelerated Expectation-Maximization (EM) algorithm to infer population parameters; however, its computational demands scale poorly, limiting its usefulness for modern biobank-sized datasets. While recent EM acceleration strategies employing second-order quasi-Newton schemes preserve accuracy, they remain computationally intensive. Conversely, EM-free approaches that prioritize speed often compromise solution quality. RESULTS: We introduce ADAMIXTURE, a novel optimization framework that integrates the EM algorithm with Adaptive Moment Estimation (Adam). Unlike traditional acceleration methods, ADAMIXTURE utilizes first-order gradients with adaptive learning rates derived from raw and squared moments to approximate curvature information, bypassing the computational overhead of Hessian approximations. This approach surpasses the convergence efficiency of second-order methods while maintaining the low computational complexity of first-order updates. Across simulated and large-scale empirical datasets, ADAMIXTURE demonstrates substantial reductions in wall-clock runtime and enhanced scalability compared to state-of-the-art methods, while maintaining comparable or improved inference accuracy. Its GPU implementation runs in under 2 h on half a million samples and variants, a two order of magnitude speedup over current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/AI-sandbox/ADAMIXTURE.

Clustering Algorithms

Population genomics of Plasmodium malariae from 4 African countries.

BACKGROUNDMalaria caused by Plasmodium malariae is geographically widespread and sometimes associated with prolonged infection, yet little is known about its genomic epidemiology.METHODSWe performed hybrid capture and whole-genome sequencing of 77 isolates collected from Cameroon (n = 7), the Democratic Republic of the Congo (n = 16), Nigeria (n = 4), and Tanzania (n = 50) between 2015 and 2021, analyzing parasite genetic population structure and demography.RESULTSThere is no evidence of geographic population structure. Nucleotide diversity was significantly lower than in colocalized P. falciparum isolates, while linkage disequilibrium was significantly higher. Genome-wide selection scans identified no erythrocyte invasion ligands or antimalarial resistance orthologs as top hits; however, targeted analyses of these loci revealed evidence of selective sweeps around 4 erythrocyte invasion ligands and 6 antimalarial resistance orthologs. Demographic inference modeling suggests that African P. malariae is recovering from a bottleneck.CONCLUSIONP. malariae is genomically atypical among human Plasmodium spp. and lacks strong population structure in Africa. The low diversity has potential impacts on understanding persistent versus new infection through genomic epidemiology.FUNDINGBill & Melinda Gates Foundation (grant 002202), USAID/PMI through Jhpiego and CDC, NIH (T32AI007151, T32AI070114, R01AI107949, R01AI129812, R21 AI148579, R01AI137395, R21AI152260, R01AI132547, and K24AI134990), and the DELTAS Africa initiative (DELGEME grant 107740/Z/15/Z).

Plasmodium malariae

Genetic structure and demographic history of house mice in western Europe inferred using whole-genome sequences.

The western house mouse, Mus musculus domesticus, is a human commensal and an outstanding model organism for studying a wide variety of traits and diseases. However, we have few genomic resources for wild mice and only a rudimentary understanding of the demographic history of house mice in Europe. Here, we sequenced 59 whole genomes of mice collected from England, Scotland, Wales, Guernsey, northern France, Italy, Portugal and Spain. We combined this dataset with 24 previously published sequences from southern France, Germany and Iran and compared patterns of population structure and inferred demographic parameters for house mice in western Europe to patterns seen in humans. Principal component and phylogenetic analyses identified three genetic clusters in western European mice. Admixture and f-branch statistics identified historical gene flow between these genetic clusters. Demographic analyses suggest a shared history of population bottlenecks prior to 20 000 years ago. Estimated divergence times between populations of house mice from western Europe ranged from 1500 to 5500 years ago, in general agreement with the zooarchaeological record. These results correspond well with key aspects of contemporary human population structure and the history of migration in western Europe, highlighting the commensal relationship of this important genetic model.

Animals

Repeated evolution on oceanic islands: comparative genomics reveals species-specific processes in birds.

Understanding the interplay between genetic drift, natural selection, gene flow, and demographic history in driving phenotypic and genomic differentiation of insular populations can help us gain insight into the speciation process. Comparing patterns across different insular taxa subjected to similar selective pressures upon colonizing oceanic islands provides the opportunity to study repeated evolution and identify shared patterns in their genomic landscapes of differentiation. We selected four species of passerine birds (Common Chaffinch Fringilla coelebs/canariensis, Red-billed Chough Pyrrhocorax pyrrhocorax, House Finch  Haemorhous mexicanus and Dark-eyed/island Junco Junco hyemalis/insularis) that have both mainland and insular populations. Changes in body size between island and mainland populations were consistent with the island rule. For each species, we sequenced whole genomes from mainland and insular individuals to infer their demographic history, characterize their genomic differentiation, and identify the factors shaping them. We estimated the relative (Fst) and absolute (dxy) differentiation, nucleotide diversity (π), Tajima's D, gene density and recombination rate. We also searched for selective sweeps and chromosomal inversions along the genome. All species shared a marked reduction in effective population size (Ne) upon island colonization. We found diverse patterns of differentiated genomic regions relative to the genome average in all four species, suggesting the role of selection in island-mainland differentiation, yet the lack of congruence in the location of these regions indicates that each species evolved differently in insular environments. Our results suggest that the genomic mechanisms involved in the divergence upon island colonization-such as chromosomal inversions, and historical factors like recurrent selection-differ in each species, despite the highly conserved structure of avian genomes and the similar selective factors involved. These differences are likely influenced by factors such as genetic drift, the polygenic nature of fitness traits and the action of case-specific selective pressures.

Animals

Accounting for recombination rate variation improves inference of barrier loci and reveals the role of both natural and sexual selection in an incipient bird radiation.

Examining genomic patterns of differentiation across lineage pairs at different stages of the speciation continuum, in combination with recombination maps, can help disentangle the effects of linked and divergent selection and identify lineage-specific targets of selection that may act as barrier loci during speciation. Here, we apply this framework to genomic data from African and Indian Ocean bird species of the genus Zosterops (Zosteropidae) to identify candidate barrier loci between ecologically, phenotypically, and genetically distinct Reunion gray white-eye (Zosterops borbonicus) parapatric geographic forms. Using analyses that account for recombination rate variation, we show that putative targets of divergent selection are primarily located on the Z chromosome, except in comparisons between geographic forms that differ in their ecologies. Functional annotation revealed that candidate barrier loci between forms with similar environmental niches are associated with genes involved in song formation and immune function, whereas those between forms with different environmental niches are associated with adaptation to altitude, morphology, and song behavior. Our results highlight the combined roles of natural and sexual selection in the evolution of reproductive barriers in this incipient species radiation.

Animals

Estimation of demography and mutation rates from one million haploid genomes.

As genetic sequencing costs have plummeted, datasets with sizes previously unthinkable have begun to appear. Such datasets present opportunities to learn about evolutionary history, particularly via rare alleles that record the very recent past. However, beyond the computational challenges inherent in the analysis of many large-scale datasets, large population-genetic datasets present theoretical problems. In particular, the majority of population-genetic tools require the assumption that each mutant allele in the sample is the result of a single mutation (the "infinite-sites" assumption), which is violated in large samples. Here, we present DR EVIL, a method for estimating mutation rates and recent demographic history from very large samples. DR EVIL avoids the infinite-sites assumption by using a diffusion approximation to a branching-process model with recurrent mutation. This approach results in tractable likelihoods that are accurate for rare alleles. We show that DR EVIL performs well in simulations and apply it to rare-variant data from one million haploid samples. We identify mutation-rate heterogeneity even after accounting for trinucleotide context and methylation status. We also predict that at modern sample sizes, the alleles at most polymorphic sites with high mutation rates represent the descendants of multiple mutation events.

Haploidy

The evolutionary origins of the parthenogenetic lizard Aspidoscelis tesselatus.

Most vertebrate species reproduce sexually. The whiptail lizards (Aspidoscelis) are a notable exception; at least 11 of the 45 recognized species are parthenogenetic. Here, we focus on one such species (Aspidoscelis tesselatus) as a case study to understand how parthenogenetic species originate and evolve. Using genome-wide sequence data and ecological niche modelling, we find that A. tesselatus likely arose from a single hybrid speciation event between A. scalaris and A. marmoratus less than 500,000 years ago. The geographic ranges of A. tesselatus and its parental species overlap currently, and niche modelling shows this zone of sympatry was even broader during the period when A. tesselatus likely formed. We additionally show evidence that A. tesselatus has a dynamic genome post-formation, with de novo mutations, introgression, and double-strand break associated events all contributing to variation within the species. These results show that asexual lineages can continue to be shaped by ongoing genomic and ecological dynamics, illuminating the processes that influence transitions in reproductive mode.

asexuality

Genomic exploration of the journey of Plasmodium vivax in Latin America.

Plasmodium vivax is the predominant malaria parasite in Latin America. Its colonization history in the region is rich and complex, and is still highly debated, especially about its origin(s). Our study employed cutting-edge population genomic techniques to analyze whole genome variation from 620 P. vivax isolates, including 107 newly sequenced samples from West Africa, Middle East, and Latin America. This sampling represents nearly all potential source populations worldwide currently available. Analyses of the genetic structure, diversity, ancestry, coalescent-based inferences, including demographic scenario testing using Approximate Bayesian Computation, have revealed a more complex evolutionary history than previously envisioned. Indeed, our analyses suggest that the current American P. vivax populations predominantly stemmed from a now-extinct European lineage, with the potential contribution also from unsampled populations, most likely of West African origin. We also found evidence that P. vivax arrived in Latin America in multiple waves, initially during early European contact and later through post-colonial human migration waves in the late 19th-century. This study provides a fresh perspective on P. vivax's intricate evolutionary journey and brings insights into the possible contribution of West African P. vivax populations to the colonization history of Latin America.

Plasmodium vivax

Habitat Specialisation Impacts Clownfish Demographic Resilience to Pleistocene Sea-Level Fluctuations.

Habitat fragmentation and loss are key threats to biodiversity, yet their impacts on marine species remain poorly understood. Clownfishes, which rely on sea anemones for shelter and reproduction, provide an interesting model to explore how ecological specialisation mediates species responses to habitat perturbations. We used whole-genome data from 382 individuals across 10 species with varying host specialisations to reconstruct demographic histories and infer spatial genetic structure to assess the impact of Pleistocene sea-level fluctuations. Generalist species, associated with multiple hosts, maintained stable effective population sizes () and population connectivity during habitat fragmentation, reflecting resilience to environmental instability. In contrast, specialists experienced severedeclines and genetic structuring, driven by their dependence on specific hosts, without signs of population recovery following habitat reconnection. Spatial genomic analyses identified the Indonesian Through-Flow as a key dispersal corridor and the Coral Triangle as a critical hub of genetic diversity, while continental shelves and extensive open ocean regions appeared as barriers to gene flow. Our findings reveal how host specialisation shapes clownfish population dynamics, emphasising the importance of incorporating ecological dependencies into conservation assessments and deepening our understanding of species responses to ecological constraints and environmental changes over evolutionary timescales.

Animals

Targeted population genomics uncovers demographic history and genetic divergence in north American wild cranberry.

Wild populations of North American cranberry (Vaccinium macrocarpon Aiton) are reservoirs of genetic variation that may contribute to the improvement of breeding-relevant traits. However, the extent to which wild genetic variation is geographically structured and represented in elite germplasm remains unclear. We analysed 179 wild cranberry accessions from the upper Midwest and Eastern North America to estimate nucleotide diversity (π), population structure, and loci associated with genetic differentiation and environmental variables using a genome-informed targeted genotyping panel. Additionally, 14 demographic scenarios were evaluated using site-frequency-spectrum-based inference to identify historical events that could explain current genetic diversity. We observed extremely low nucleotide diversity within the targeted panel (π = 5 × 10-6). Rare allele distributions strongly influenced π and Tajima's D values, suggesting constrained diversity in the genomic regions assayed that is not captured by heterozygosity-based estimates alone. However, we interpreted these results as conservative lower bounds on genome-wide neutral diversity because the targeted panel is enriched for genic and conserved regions. A clear separation between the Midwest and East populations was observed, with inbreeding coefficients ranging from -0.13 to 0.15. Furthermore, site frequency spectrum inference from the targeted panel supported a demographic scenario consistent with a significant population reduction ≈15-14 thousand years ago (kya), followed by a divergence between the two regions ≈12 kya, and an asymmetric gene flow ≈1.3 kya. We detected 254 candidate loci showing regional allele-frequency differentiation. Several of these loci colocalized with candidate genes linked to stress response, development, and metabolic processes. To evaluate the representation of geographically differentiated wild alleles in a breeding context, we analysed Rutgers breeding materials (n = 484) and found that this panel is enriched for common alleles in Eastern wild populations. These findings indicate regionally structured allele-frequency variation in wild cranberry, with potential relevance to environmental response and breeding. This study extends prior wild cranberry population-genetic research by providing targeted-panel estimates of diversity, comparisons of demographic models, and breeding insights on geographically differentiated alleles, while highlighting the importance of conserving wild cranberry germplasm for use in modern breeding programs.

Journal Article

A General Framework for Branch Length Estimation in Ancestral Recombination Graphs.

Inference of Ancestral Recombination Graphs (ARGs) is of central interest in the analysis of genomic variation. ARGs can be specified in terms of topologies and coalescence times. The coalescence times are usually estimated using an informative prior derived from coalescent theory, but this may generate biased estimates and can also complicate downstream inferences based on ARGs. Here we introduce, POLEGON, a novel approach for estimating branch lengths for ARGs which uses an uninformative prior. Using extensive simulations, we show that this method provides improved estimates of coalescence times and lead to more accurate inferences of effective population sizes under a wide range of demographic assumptions (population expansion, bottleneck, split, etc). It also improves other downstream inferences including estimates of mutation rates. We apply the method to data from the 1000 Genomes Project to investigate population size histories and differential mutation signatures across populations. We also estimate coalescence times in the HLA region, and show that they exceed 30 million years in multiple segments.

Ancestral Recombination Graph

Advancing translational exposomics: bridging genome, exposome and personalized medicine.

Understanding the interplay between genetic predisposition and environmental and lifestyle exposures is essential for advancing precision medicine and public health. The exposome, defined as the sum of all environmental exposures an individual encounters throughout their lifetime, complements genomic data by elucidating how external and internal exposure factors influence health outcomes. This treatise highlights the emerging discipline of translational exposomics that integrates exposomics and genomics, offering a comprehensive approach to decipher the complex relationships between environmental and lifestyle exposures, genetic variability, and disease phenotypes. We highlight cutting-edge methodologies, including multi-omics technologies, exposome-wide association studies (EWAS), physiology-based biokinetic modeling, and advanced bioinformatics approaches. These tools enable precise characterization of both the external and the internal exposome, facilitating the identification of biomarkers, exposure-response relationships, and disease prediction and mechanisms. We also consider the importance of addressing socio-economic, demographic, and gender disparities in environmental health research. We emphasize how exposome data can contextualize genomic variation and enhance causal inference, especially in studies of vulnerable populations and complex diseases. By showcasing concrete examples and proposing integrative platforms for translational exposomics, this work underscores the critical need to bridge genomics and exposomics to enable precision prevention, risk stratification, and public health decision-making. This integrative approach offers a new paradigm for understanding health and disease beyond genetics alone.

Humans