PubMed HealthSearch

SEARCH · PubMed Health

Results for “population structure”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Association mapping in structured populations.

The use, in association studies, of the forthcoming dense genomewide collection of single-nucleotide polymorphisms (SNPs) has been heralded as a potential breakthrough in the study of the genetic basis of common complex disorders. A serious problem with association mapping is that population structure can lead to spurious associations between a candidate marker and a phenotype. One common solution has been to abandon case-control studies in favor of family-based tests of association, such as the transmission/disequilibrium test (TDT), but this comes at a considerable cost in the need to collect DNA from close relatives of affected individuals. In this article we describe a novel, statistically valid, method for case-control association studies in structured populations. Our method uses a set of unlinked genetic markers to infer details of population structure, and to estimate the ancestry of sampled individuals, before using this information to test for associations within subpopulations. It provides power comparable with the TDT in many settings and may substantially outperform it if there are conflicting associations in different subpopulations.

Alleles

Genome-wide SNP-based genomic diversity and population structure analysis in alpaca populations from Europe and Peru.

This study aimed to analyze the genetic diversity and population structure of alpacas in Germany, Switzerland, and Austria (German-speaking regions, GSR) and to compare with that of the country of origin of the species (Peru). A total of 179 animals from GSR and 151 from Peru were genotyped with a species-specific 76k SNP array. The observed and expected heterozygosity was 0.305 and 0.311 for GSR and 0.310 and 0.312 for Peru. The mean FROH values were 0.029 for GSR and 0.023 for Peru. In general, results show that breeders in both analyzed regions efficiently maintain genetic diversity. Principal component analysis identified the GSR and Peru populations as separate from each other, but the relative proximity of both clusters indicates the shared genetic heritage. FST and XPEHH methods identified genomic regions under selection for traits such as coat color and adaptation. Genome-wide association studies comparing black and brown with white or gray alpacas identified associated genome regions containing the ASIP and KIT genes, respectively. The association of a recently identified keratin locus on chromosome 16 with differences in fleece type in alpacas was confirmed, while the putative causality of a TRPV3 variant was rejected.

Animals

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

Mitochondrial genome-derived microsatellites reveal genetic diversity and population structure in Callery pear populations.

Callery pear (Pyrus calleryana Decne.; PC) possesses many desirable characteristics valued in managed landscapes. This has driven the release of numerous cultivars, including both hybrids and selections derived from native populations. The extensive planting of PC cultivars in managed areas has contributed to the widespread occurrence of invasive individuals across a broad range of habitats in the eastern United States (US). Self-incompatibility, tolerance to various environmental conditions, pathogen and pest resistance, intraspecific hybridization among the cultivars, possible interspecific hybridization with other Pyrus species, and seed dispersal by various vertebrates have contributed to the spread and persistence of PC across diverse environments. Because effective and environmentally appropriate management options remain limited, improved understanding of PC genetics may help inform management strategies. Previous studies have characterized PC diversity using nuclear genomic short sequence repeats (gSSRs), however, neither a mitochondrial genome resource nor mitochondrial short sequence repeats (mtSSRs) have been developed for this purpose. Here, we assembled a mitochondrial genome of 485,892 bp and used five mtSSRs to characterize mitochondrial diversity and population structure among accessions from the species' native range in Asia (n = 72), southeastern US escapees (SNesc; n = 90), Tennessee escapees (TNesc; n = 90), and US-released commercial cultivars (UScult; n = 69 representing 14 unique cultivars). We found a high genetic diversity (He = 0.728) and evidence of genetic structure in PC. In distance-based and multivariate analyses, UScult occupied an intermediate position between the Asian populations and the US escapees. The observed mitochondrial diversity among samples assigned to PC cultivars is consistent with a complex genetic landscape and may reflect distinct maternal lineages, cultivar-labeling or record-keeping discrepancies, and/or technical variation. This study underscores the need for broader genomic investigations using authenticated cultivar reference material and high-resolution nuclear markers to resolve cultivar ancestry, validate true-to-name identity, and inform species management.

Genetic Variation

Population structure in Kanoya population, Japan.

The mean inbreeding coefficients found for Minami-cho (366 couples) and Shinsei-cho (511 couples) were 0.00307 and 0.00191, respectively. The mean inbreeding coefficient decreased and the mean marital distance increased as the year of marriage becomes more recent. The mean distances and their standard deviations between birthplaces of mates, father-offspring, mother-offspring, and sibs are 69.05 +/- 229.64, 73.09 +/- 246.66, 49.81 +/- 158.43, and 39.53 +/- 159.51 km, respectively, at Minami-cho. These values are 188.45 +/- 387.05, 187.79 +/- 562.59, 148.26 +/- 326.35, and 73.93 +/- 225.92 km, respectively, at Shinsei-cho. The dimensionality of migration is closest to one dimension.

Consanguinity

treestructure: an R package to detect population structure in phylogenetic trees.

MOTIVATION: How population structure can shape genetic diversity is a longstanding problem in population genetics. While the use of geographic locations, when available, can help answer some of these questions, it is still difficult to determine population structure when such metadata are not available or when the potential population structure is not easily observed. Here, we present an updated version of treestructure, an R package that implements a statistical test based on coalescent theory to detect unobserved population structure in a time-scaled phylogenetic tree. AVAILABILITY: treestructure is available at CRAN at https://cloud.r-project.org/web/packages/treestructure/ and at https://emvolz-phylodynamics.github.io/treestructure/.

Phylogeny

Recovering the precolonial population structure of Khoe-San descendant populations.

San populations from Botswana and Namibia retain exceptional linguistic, cultural, and genetic diversity, but few Khoisan-speaking groups remain south of the Kalahari Desert. However, historically, far southern Africa was home to many San and Khoekhoe groups. Popular opinion often implies that such populations do not contribute to the ancestry of contemporary South Africans. Here, we characterize the genetic ancestry of self-identified South African Coloured groups and reconstruct precolonial and colonial population structures from 620 newly sampled individuals. These groups retain the majority of Khoe-San genetic ancestry (>48%), suggesting the persistence of Khoe-San ancestry to the present day. By isolating the Khoe-San ancestry component, we show that it is intermediate between the ≠Khomani San and Nama and distinct from Kalahari Khoe-San populations. We also find that signatures of the Indian Ocean slave trade can be traced to Indonesian islands such as Sulawesi, Java, and Flores, while the South Asian ancestry is regionally nonspecific.

Humans

Genealogy of neutral genes and spreading of selected mutations in a geographically structured population.

In a geographically structured population, the interplay among gene migration, genetic drift and natural selection raises intriguing evolutionary problems, but the rigorous mathematical treatment is often very difficult. Therefore several approximate formulas were developed concerning the coalescence process of neutral genes and the fixation process of selected mutations in an island model, and their accuracy was examined by computer simulation. When migration is limited, the coalescence (or divergence) time for sampled neutral genes can be described by the convolution of exponential functions, as in a panmictic population, but it is determined mainly by migration rate and the number of demes from which the sample is taken. This time can be much longer than that in a panmictic population with the same number of breeding individuals. For a selected mutation, the spreading over the entire population was formulated as a birth and death process, in which the fixation probability within a deme plays a key role. With limited amounts of migration, even advantageous mutations take a large number of generations to spread. Furthermore, it is likely that these mutations which are temporarily fixed in some demes may be swamped out again by non-mutant immigrants from other demes unless selection is strong enough. These results are potentially useful for testing quantitatively various hypotheses that have been proposed for the origin of modern human populations.

Animals

Genetic Diversity and Population Structure of Urban and Rural Goshawks.

Urbanization poses a growing threat to biodiversity with potential impacts on species' genetic diversity and population structure. The Eurasian goshawk (Astur gentilis) is traditionally a forest-dwelling raptor that has recently established breeding populations in urban environments such as Helsinki, Finland. Here, we investigated genetic diversity and population structure across urban, suburban, and rural goshawk populations in Finland using 10 microsatellite markers and 72 individuals sampled between 1990 and 2020. Genetic diversity, measured by heterozygosity and allelic richness, was similar among populations. Genetic differentiation was low to moderate (F ST = 0.022-0.074) and statistically non-significant. Despite urbanization, contemporary urban goshawks showed genetic similarity to adjacent contemporary non-urban goshawks, while greater differentiation was observed between temporally separated populations. Consistent with this pattern, clustering supported K = 2 as the primary level of genetic structure, separating the contemporary urban and surrounding populations from the earlier surrounding and rural populations. Given the limited marker set and sample sizes, these findings are interpreted as broad-scale patterns rather than definitive evidence of fine-scale population structure. Further studies using larger sample sizes and genome-wide markers are needed to resolve population connectivity and the longer-term genetic effects of urbanization.

Astur gentilis

Phylogenetically diverse introgression drives subtle population structure in Pacific rockfishes.

Genomic methods have shown that admixture and introgression is common across animal taxa. Pacific rockfishes, genus Sebastes, are group of commercially important species that primarily inhabit inshore, shelf, and slope habitats along the North American west coast. Among these, Copper and Quillback Rockfishes (abbreviated to Copper and Quillback) are closely related species known to hybridize, particularly within the Salish Sea in North America's Pacific Northwest. Here, we investigate genetic population structure and introgression patterns in Copper and Quillback from Alaska to California. Using whole-genome resequencing (WGS) across a broad geographic range, we seek to (1) compare population structure between these species, and (2) assess how introgression affects population structure patterns. Our analyses reveal that Copper exhibit much higher levels of population differentiation compared to Quillback, especially separating Salish Sea samples from all other populations. In contrast, Quillback populations appear to be nearly panmictic, with lower overall differentiation. Surprisingly, we detected signatures of introgression from 13 other rockfish species in Copper and 16 species in Quillback. This introgression was highly regional suggesting hybridization depended on geographic context and congener ranges. Yelloweye Rockfish introgression drives the strongest signal of regional population structure in Quillback. These findings provide novel insights into the range-wide genetic structure of these species and highlight that hybridization in Sebastes is phylogenetically broader than previously appreciated.

Journal Article

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs = 0.14, Ht = 0.217, I = 0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst = 0.319) and moderate gene flow (Nm = 1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated ΔK using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K = 2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

Population structure and connectivity among coastal and freshwater Kelp Gull (Larus dominicanus) populations from Patagonia.

The genetic identification of evolutionary significant units and information on their connectivity can be used to design effective management and conservation plans for species of concern. Despite having high dispersal capacity, several seabird species show population structure due to both abiotic and biotic barriers to gene flow. The Kelp Gull is the most abundant species of gull in the southern hemisphere. In Argentina it reproduces in both marine and freshwater environments, with more than 100,000 breeding pairs following a metapopulation dynamic across 140 colonies in the Atlantic coast of Patagonia. However, little is known about the demography and connectivity of inland populations. We aim to provide information on the connectivity of the largest freshwater colonies (those from Nahuel Huapi Lake) with the closest Pacific and Atlantic populations to evaluate if these freshwater colonies are receiving immigrants from the larger coastal populations. We sampled three geographic regions (Nahuel Huapi Lake and the Atlantic and Pacific coasts) and employed a reduced-representation genomic approach to genotype individuals for single-nucleotide polymorphisms (SNPs). Using clustering and phylogenetic analyses we found three genetic groups, each corresponding to one of our sampled regions. Individuals from marine environments are more closely related to each other than to those from Nahuel Huapi Lake, indicating that the latter population constitutes the first freshwater Kelp Gull colony to be identified as an evolutionary significant unit in Patagonia.

Humans

Comparative genomic analysis reveals distinct population structure in Legionella anisa.

Legionella anisa has been frequently isolated from engineered water systems; however, its population structure remains understudied compared to Legionella pneumophila. Here, we generated complete genome sequences for four L. anisa isolates recovered from a healthcare facility in Rimouski, Canada. Further the population structure of this species was investigated by performing comparative genomic analyses of the genomes generated in this study together with publicly available L. anisa genomes. Genome-wide phylogenetic analysis revealed the presence of three distinct clades separated by substantial genetic divergence (∼500 SNP), with the Rimouski isolates forming a tightly clustered group, suggesting a clonal lineage. Comparative pangenome analysis indicated moderate core genome conservation accompanied by a highly variable accessory genome (∼50%). The isolates characterized in this study harbored multiple plasmids encoding genes associated with conjugation, heavy metal resistance, and other stress-related functions, suggesting potential roles in environmental persistence. Previous studies have shown that L. anisa can proliferate within protozoan host cells, although outcomes vary depending on the host species. Our isolates showed efficient proliferation within Acanthamoeba castellanii, but not within Vermamoeba vermiformis, under the conditions tested. Together, these findings underscore the genomic diversity of this understudied Legionella species and provide a framework for future investigations regarding environmental persistence and potential pathogenicity.

Legionella anisa, Whole genome sequencing

Population structure of Barra (Outer Hebrides).

Historical demography, surname concordance (isonymy), migration, and genealogy give a consistent description of population structure. The census size has averaged about 1400 over the last five centuries. Conjoined with an effective migration rate of 3-05 per generation as estimated by three different methods, this gives an evolutionary size of 638, random kinship of 0-008 and inbreeding of 0-007 relative to the rest of Britain. The population structure of Barra is similar to other British isolates in the recent past, but an order of magnitude less inbred than slash-and-burn agriculturalists and Pacific Islanders. Some consequences for rare genes and polymorphisms are discussed.

Emigration and Immigration

Population structure and mitochondrial DNA gene flow in Old World populations of Drosophila subobscura.

An extensive survey of mitochondrial DNA (mtDNA) restriction polymorphism in 156 isofemale lines from 29 different geographic populations of Drosophila subobscura distributed throughout the Old World was carried out. Ten restriction enzymes were used, five of which revealed restriction site polymorphism. Of the 31 restriction sites detected, 13 were found to be polymorphic. Comparisons with the mtDNA map of Drosophila yakuba indicate that the variable sites are mainly concentrated in protein genes, especially those corresponding to the NADH complex. A total of 13 different haplotypes were observed, two of which (haplotypes I and II) are quite frequent and widely distributed throughout the populations, whereas the other 11 with the exception of VIII, which deserves special attention, are each restricted to one population only and occur at low frequencies. The observed distribution of haplotypes, corroborated by a parsimonious unrooted tree, suggests an ancient origin of haplotypes I and II in the continent. In order to compare genetic structure according to mtDNA and allozymes, the 10 populations with higher population sizes were studied for 10 polymorphic allozymes also. One striking result is the high degree of population structure of the mtDNA when compared to that obtained for allozymes. If an island model is assumed, estimates of gene flow give values of 0.013 and 1.89 migrants per generation for mtDNA and allozymes, respectively. What is apparent from these estimates is that Drosophila subobscura populations are effectively subdivided for mtDNA genes at migration rates at which nuclear genes (allozymes) are almost panmictic.

Alleles

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals

Population structure, gene flow and natural selection in populations of Euphydryas phaeton.

An examination of seven proteins, presumably encoded by seven structural gene loci, in three local populations of the supposedly sedentary and colonial butterfly, Euphydryas phaeton revealed that three (43 per cent) were polymorphic with three to five alleles each. In addition to this high level of heterozygosity, no statistically significant differences in allele frequencies were found at two of the three polymorphic loci. Since the effective breeding size in each population was estimated to range from as few as 20 to 200 individuals, it appears that some level of gene flow between populations must be invoked to explain the high levels of genetic variability maintained in local populations of this butterfly, despite its apparently colonial nature.

Alleles

Genetic investigation of population structure in Atlantic chub mackerel, Scomber colias Gmelin, 1789 along the West African coast.

Sustainable management of transboundary fish stocks hinges on accurate delineation of population structure. Genetic analysis offers a powerful tool to identify potential subpopulations within a seemingly homogenous stock, facilitating the development of effective, coordinated management strategies across international borders. Along the West African coast, the Atlantic chub mackerel (Scomber colias) is a commercially important and ecologically significant species, yet little is known about its genetic population structure and connectivity. Currently, the stock is managed as a single unit in West African waters despite new research suggesting morphological and adaptive differences. Here, eight microsatellite loci were genotyped on 1,169 individuals distributed across 33 sampling sites from Morocco (27.39°N) to Namibia (22.21°S). Bayesian clustering analysis depicts one homogeneous population across the studied area with null overall differentiation (F ST = 0.0001ns), which suggests panmixia and aligns with the migratory potential of this species. This finding has significant implications for the effective conservation and management of S. colias within a wide scope of its distribution across West African waters from the South of Morocco to the North-Centre of Namibia and underscores the need for increased regional cooperation in fisheries management and conservation.

Animals