PubMed HealthSearch

SEARCH · PubMed Health

Results for “Admixed ancestry”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A genealogy-based approach for revealing ancestry-specific structures in admixed populations.

Elucidating ancestry-specific structures in admixed populations is crucial for comprehending population history and mitigating confounding effects in genome-wide association studies. Existing methods to reveal the ancestry-specific structures generally rely on frequency-based estimates of genetic relationship matrix (GRM) among admixed individuals after masking segments from ancestry components not being targeted for investigation. However, these approaches disregard linkage information between markers, potentially limiting their resolution in revealing structure within an ancestry component. We introduce ancestry-specific expected GRM (as-eGRM), a novel framework for estimating the relatedness within ancestry components between admixed individuals. The key design of as-eGRM consists of defining ancestry-specific pairwise relatedness between individuals based on genealogical trees encoded in the ancestral recombination graph (ARG) and local ancestry calls and then computing the expectation of the ancestry-specific relatedness across the genome. Comprehensive evaluations using both simulated stepping-stone models of population structure and empirical datasets based on three-way admixed Latino cohorts showed that analysis based on as-eGRM robustly outperforms existing methods in revealing the structure in admixed populations with diverse demographic histories, which in turn improves the robustness against confounding due to population structure in association testing.

Humans

Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10 250 participants from the HostSeq cohort.

Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10 250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score ± SD: 0.012 ± 0.052), Black (0.013 ± 0.042), White (0.055 ± 0.022), and South Asian (0.057 ± 0.098) participants, whereas higher scores among Hispanic (0.083 ± 0.060) and Middle Eastern or Central Asian (0.122 ± 0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.

Humans

Genome-wide association studies reveal genetic variants associated with antineoplastic monoterpenoid indole alkaloid accumulation in Catharanthus roseus.

Catharanthus roseus produces pharmacologically important monoterpenoid indole alkaloids (MIAs), yet their natural accumulation is low, limiting therapeutic exploitation. To dissect the genetic basis of natural variation in MIA accumulation, we integrated phenotypic, chemotypic, and genomic analyses of 93 C. roseus accessions sampled from six locations across India, including New Delhi, Lucknow, Jodhpur, Bangalore, and two locations in Gujarat: Navsari and Bardoli. Morphological characterization showed limited differentiation among locations, whereas accessions from Gujarat tended to be taller compared to other locations and more frequently white-flowered. Quantitative HPLC profiling revealed substantial accession- and location-dependent variation in total indole alkaloid levels, with Gujarat accessions showing the highest accumulation, largely driven by vindoline and catharanthine. Genotyping-by-sequencing generated 10,801 high-quality variants comprising 10,087 SNPs and 714 InDels corresponding to an average density of 19.34 variants per Mbp of the genome, revealing three genetic subgroups with overall admixed ancestry and weak geographic stratification. Genome-wide association study (GWAS) using five benchmark models identified 47 variants potentially associated with catharanthine, vindoline, and vinblastine content. These putative candidate loci were located near genes implicated in hormone signaling, mitochondrial function, nitrogen metabolism, and RNA processing, suggesting complex regulatory control of MIA biosynthesis. Notably, two missense variants in a carboxylesterase-like gene were associated with vindoline accumulation, and highly significant intergenic SNP clusters suggested putative regulatory hotspots for vinblastine biosynthesis. These results provide GWAS-based insights into the genetic architecture of MIA metabolism in C. roseus and nominate candidate variants for precision breeding and metabolic engineering to enhance pharmaceutical alkaloid production.

Catharanthus roseus

Genome-Wide and Rare Variant Association Studies of Amblyopia in Admixed American and African Ancestry Groups.

OBJECTIVE: To identify genetic variants associated with amblyopia in African (AFR) and Admixed American (AMR) ancestry groups, expanding on previous studies conducted in European ancestry. DESIGN: Retrospective ancestry-stratified genome-wide association study (GWAS) and gene-level rare variant association study (RVAS). PARTICIPANTS: Participants in the All of Us Research Program from AFR and AMR ancestry groups who had whole-genome sequencing available. Cases and controls were distinguished based on the presence of International Classification of Diseases 9/10/SNOMED diagnosis codes for amblyopia in electronic health records. This yielded ancestry-stratified subsets of 269 cases and 71 585 controls of AMR ancestry and 366 cases and 79 460 controls of AFR ancestry. METHODS: Stratified logistic regression models were adjusted for age, biological sex, and the top 10 principal components of genomic ancestry. GWAS was limited to common variants (minor allele frequency &#x2265;1%), and RVAS was limited to rare variants with coding sequence-altering effects (minor allele frequency >1%, exonic only, excluding synonymous variants) aggregated at the gene level using the SKAT algorithm. Downstream analyses of the significant variants were performed using KEGG and GO pathway analysis and STRING database queries for protein-protein interactions and gene-gene interactions. MAIN OUTCOME MEASURES: Single-nucleotide polymorphisms were determined to have genome-wide significance if P < 5e-8 in the GWAS, and genes were determined to have significant association with amblyopia in the RVAS if P < 8.0 &#xd7; 10-4. RESULTS: In the AMR GWAS, 245 unique single-nucleotide polymorphisms mapping to 97 distinct loci were identified, notably within neurodevelopmental and axonal guidance genes, including ROBO1, SEMA4B, PTPRD, NRXN1, and CAMK2D. The AFR GWAS identified 11 significant variants corresponding to 6 loci mapping primarily to long noncoding RNAs and pseudogenes. The AMR RVAS identified 15 genes, including axonal transport genes (KIF1B and KIF7) and growth factor signaling genes (EGF, ERBIN, and AKAP17A). The AFR RVAS identified a single gene, DLG2, which encodes the postsynaptic protein PSD-93, which promotes the closure of the sensitive period of neuroplasticity for vision in early childhood. CONCLUSIONS: Genetic risk architectures for amblyopia differ across ancestries but fundamentally converge on neurodevelopmental signaling, cortical synapse assembly, and sensitive period plasticity rather than ocular structural dynamics. FINANCIAL DISCLOSURE(S): The authors have no proprietary or commercial interest in any materials discussed in this article.

Amblyopia

DiscoDivas: Leveraging genetic ancestry continuum information to interpolate PRS for admixed populations.

The relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits polygenic risk score (PRS) development and equitable clinical translation for admixed populations. Under the assumption that the most informative PRS model for a genetically homogeneous sample varies linearly in an ancestry continuum space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, genetic ancestries, leveraging multiple PRS models fine-tuned within existing samples, which are mostly of single ancestry, and genetic distance. DiscoDivas treats genetic ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar genetic ancestry samples. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed individuals.

PRS harmonization

Quantitative trait loci mapping of gene expression and chromatin accessibility in primary fibroblasts reveals shared allelic effects between Latin American and European ancestries.

BACKGROUND: Quantitative Trait Locus (QTL) analysis of molecular data has identified genetic variants associated with traits such as gene expression, and colocalization of these functional QTL with GWAS risk loci has offered insights into the genetic basis of human disease. We employed gene expression (RNA-seq) and chromatin accessibility (ATAC-seq) obtained from human primary fibroblasts to investigate quantitative trait loci (QTLs) in cohorts ascertained for bipolar disorder of European (n&#x2009;=&#x2009;150) and Latin American (n&#x2009;=&#x2009;96) ancestries. RESULTS: Leveraging data from three countries of origin (The Netherlands, Colombia, Costa Rica) within our cohort, we characterized differences among individuals at the SNP, gene, and accessible-chromatin levels to compute ancestry-specific expression (e)QTLs and chromatin-accessibility (ca)QTLs. Across ancestries, we observed R2&#x2009;&#x2265;&#x2009;0.93 for eQTL effect sizes and R2&#x2009;&#x2265;&#x2009;0.95 for caQTLs, indicating a high degree of concordance. Integrating chromatin data with expression and genotype information enabled precise fine-mapping of eQTLs, yielding 203 genes with high-confidence (posterior probability&#x2009;>&#x2009;90%) candidate regulatory pathways. In downstream analyses, transcriptome-wide (TWAS) and chromatin-wide (CWAS) association studies with brain- and skin-related GWAS identified 36 TWAS-significant genes and 77 CWAS-significant open chromatin regions. CONCLUSIONS: These findings underscore the shared genetic regulatory mechanisms across European and Latin American ancestries, while demonstrating that ancestry-specific reference panels enhance the accuracy of TWAS and CWAS in diverse populations. More broadly, this study highlights the value of paired multi-omic datasets from diverse cohorts for interpreting disease-associated genetic variation.

Humans

Leveraging local ancestry and cross-ancestry genetic architecture to improve genetic prediction of complex traits in admixed populations.

The broader application of polygenic risk score (PRS) is hindered by the limited transferability of PRS developed in Europeans to non-European populations. While many statistical methods have been developed to improve the performance of PRS in non-European populations, most of them focused on discrete genetic ancestry clusters and did not consider admixed individuals. Admixed individuals pose a unique challenge for PRS calculation due to the complexity of local ancestry and cross-ancestry effect sizes. Here, we present a statistical method called SDPR_admix for calculating PRS in admixed individuals. SDPR_admix characterizes the joint distribution of the effect sizes of a genetic variant with two ancestries to be both zero, ancestry enriched, or shared with correlation. SDPR_admix outperformed other methods in simulations and improved the prediction of real traits in European-African admixed individuals in UK Biobank when trained on the Population Architecture using Genomics and Epidemiology (PAGE) dataset (N = 13,000). Deployment of SDPR_admix on All of Us (N = 52,000) further increased the prediction accuracy by approximately 5-fold on average compared with training on PAGE. This enhancement was achieved with manageable computational time and cost, demonstrating the feasibility of training PRS models on large-scale All of Us data. We provided several examples demonstrating that both ancestral-enriched and shared effects, as included in the SDPR_admix prediction model, are helpful for improving polygenic prediction in admixed populations. We also applied SDPR_admix to construct PRS for admixed Americans with mixture of European and Amerindigenous ancestries and showed that SDPR_admix overall outperformed other methods.

Humans

The paradoxical extinction: Exploring signatures of assortative mating as a possible mechanism that maintains canonical Red Wolf genetic ancestry in the American Gulf Coast canids.

Admixed genomes, particularly those with an evolutionary history of genetic exchange with an endangered or extinct species, are valued for innovative and unconventional conservation actions. Here, we show the substantial conservation value that the admixed canids of the Gulf Coast have as they retain high amounts of contemporary Red Wolf ancestry and unique genetic variation of past Red Wolf lineages (e.g. ghost ancestry). We analyzed 54,439 loci genotyped across the genome of 413 North American canids and investigated the role that assortative mating with respect to ancestry proportions played in the retention of endangered genetic variation. We report high correlations of inter-chromosomal ancestry proportions that varied with geographic location along Texas and Louisiana Gulf Coast populations, with the stronger signatures reported in the latter. We found that models of assortative mating promoted greater ancestry variance compared with random mating leading to increased efficiency of selection for Red Wolf and ghost alleles. Despite the Red Wolf being extinct in the wild, original, and ghost genomic variation persists in Gulf Coast admixed canids. We suggest two conservation strategies that value and preserve this unique and endangered genomic variation through designed breeding programs. Ultimately the incorporation of this ghost genetic variation would be valuable to boost the genetic viability of the ex situ Red Wolf breeding program, create in situ redundancy, and avoid extinction for this endemic American wolf species.

Animals

GWAS for Periodontitis Phenotypes Using Multi-Ancestry All of Us Research Platform.

Periodontitis is a multifactorial inflammatory disease whose pathogenesis is associated with intricate interactions between genetic and environmental factors. Leveraging electronic health records data from the All of Us Research Program, we stratified periodontitis by clinically relevant dimensions: stage, grade, and extent. Based on these phenotypes, we performed a multi-ancestry genome-wide association study, focusing on predominant ancestry populations of African, European, and Admixed American. Our study cohort comprised 3,881 periodontitis patients and a control group of 10,760 patients with dental caries and without periodontitis. Ancestry-specific GWAS revealed significant genetic associations (P<5&#xd7;10-8) in periodontitis grade phenotypes at the LINC00294 and CLMN loci in the African ancestry population and also confirmed via the multi-ancestry meta-analysis. In addition, the XYLT1 locus emerged as a significant signal associated with periodontitis grade phenotype in the admixed American GWAS. Our GWAS comparing periodontitis to dental caries in the admixed American population identified several significant loci, including RABGAP1L, previously linked to immune regulation, DCHS2, a cadherin-related gene involved in bone mineralization and tissue morphogenesis, and OSTM1, known to be crucial for bone remodeling. The findings of our study highlight the potential of integrating EHR and genomic data from large-scale biobanks to achieve informative dental phenotyping, uncover novel molecular insights into periodontal disease, and personalize treatment approaches.

Journal Article

Genetic Ancestry and Colorectal Cancer in the All of Us Dataset.

IMPORTANCE: Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited. OBJECTIVE: To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model. DESIGN, SETTING, AND PARTICIPANTS: This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome. EXPOSURES: Genetically inferred ancestry categories and principal components. MAIN OUTCOMES AND MEASURES: Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost). RESULTS: Among 316&#x202f;624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172&#x202f;327 [54.4%] of European ancestry; 191&#x202f;705 female [61.2%]; 121&#x202f;585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated. CONCLUSIONS AND RELEVANCE: In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.

Aged

Nested Admixture During and After the Trans-Atlantic Slave Trade on the Island of S&#xe3;o Tom&#xe9;.

Human genetic admixture, involving the contact between two or more previously isolated populations, can be a complex process influenced by social dynamics. In this study, we aim to reconstruct complex admixture histories in S&#xe3;o Tom&#xe9;, an island in the Gulf of Guinea where the Portuguese established one of the first plantation-based slave societies. Since the 15th century, migration waves from Africa and Europe, slavery, marooning, and indentured labour led to profound demographic shifts and social stratification on the island. Examining 2.5 million SNPs newly genotyped in 96 S&#xe3;o Tom&#xe9;ans, we observed patterns of genetic differentiation that were more complex than those of other populations descended from enslaved Africans on either side of the Atlantic. Using local ancestry inference and Identical-by-Descent methods, we identified five genetic clusters in S&#xe3;o Tom&#xe9; and reconstructed shared ancestries between each cluster and 70 African and European population samples, including an extensive sample from the Cabo Verde archipelago. Our findings align with historical records, retracing the major slave trade routes and labour-driven migrations after the abolition of slavery. We also identified gene flow between recently admixed groups that were previously isolated on the island. We call this process, creating multiple layers of genetic ancestry in admixed genomes, nested admixture. We suggest that changing social structures in S&#xe3;o Tom&#xe9; transformed the genetic structure of its population and influenced the admixture process. This study demonstrates how successive admixture and isolation events during and after the Trans-Atlantic Slave Trade shaped extant genetic diversity patterns at local scale in Africa.

Humans

Tractor Workflow Pipeline: A Scalable Nextflow Framework for Local Ancestry-Aware Genome-Wide Association Studies.

The routine exclusion of admixed individuals from traditional Genome-Wide Association Studies (GWAS) due to concerns about spurious associations has hindered genetic analyses involving multiple ancestries. Tractor GWAS addresses this issue by incorporating local ancestry into its analysis, empowering identification of ancestry-enriched hits and generating ancestry-specific summary statistics. However, Tractor requires accurate genomic phasing and local ancestry inference as prerequisite steps, which requires additional bioinformatics expertise and decision points regarding reference panel setup. To streamline, harmonize, and automate this process, we present a scalable Nextflow workflow that integrates all necessary steps, minimizing the need for manual intervention while remaining modular and customizable. The workflow supports multiple commonly used tools and offers flexibility in how Tractor is implemented. To demonstrate its utility, we applied this pipeline to analyze 32 blood biomarkers in 6,245 two-way AFR-EUR admixed individuals from the UK Biobank. This pipeline ran efficiently at scale, replicated known associations, and identified novel ancestry-specific loci. These novel associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously missed genetic signals. By enabling the efficient analysis of admixed individuals, our workflow facilitates Tractor use, paving the way for more broader genetic discovery.

Journal Article

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans

A multi-ancestry polygenic risk score for body mass index predicts longitudinal weight change.

BACKGROUND: Identifying individuals at risk for future weight gain is challenging, partly because associations with traditional clinical risk factors may be biased by confounding and reverse causation. Polygenic risk scores (PRS) provide a stable, lifelong measure of genetic predisposition to obesity. However, existing PRS have not been evaluated for their association with longitudinal weight change in adulthood and often lack generalizability across diverse genetic ancestry groups. METHODS: We conducted ancestry-specific genome-wide association study meta-analyses of body mass index (BMI) in populations of European, African or African American, Admixed American, East Asian, and South Asian ancestries and developed ancestry-specific PRS. A multi-ancestry polygenic risk score (MAPRS) was trained using ancestry-specific PRS in a model selection dataset (N&#x2009;=&#x2009;39,685) from the All of Us Research Program (AoU). We evaluated the MAPRS in an independent AoU model evaluation dataset (N&#x2009;=&#x2009;158,743) for BMI prediction and in a separate AoU test dataset (N&#x2009;=&#x2009;78,219) with repeated measurements over 1.5-2.5 years for weight change prediction. The outcomes included change in BMI and&#x2009;&#x2265;&#x2009;10% or&#x2009;&#x2265;&#x2009;5% total body weight (TBW) gain. We further examined the relationship between MAPRS and 12 clinical risk factors commonly comorbid with obesity in relation to weight change. RESULTS: The MAPRS captured 7.05% of the variance in measured BMI in the AoU model evaluation dataset and demonstrated improved generalizability across all non-European genetic ancestry groups. In the AoU test dataset, conditioned on baseline BMI at the second-to-last measurement, a one SD increase in MAPRS was associated with a 0.16 kg/m2 increase in future BMI (standard error&#x2009;=&#x2009;0.012 kg/m2; p-value&#x2009;=&#x2009;2.2&#x2009;&#xd7;&#x2009;10-39), 1.27-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;10% TBW gain (95% CI: 1.24-1.31; p-value&#x2009;=&#x2009;1.4&#x2009;&#xd7;&#x2009;10-55), and 1.15-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;5% TBW gain (95% CI: 1.13-1.18; p-value&#x2009;=&#x2009;2.8&#x2009;&#xd7;&#x2009;10-39). These associations were observed across all genetic ancestry groups and remained highly consistent after adjustment for any clinical risk factor. In contrast, most clinical risk factors demonstrated inconsistent or weaker associations with weight change outcomes. CONCLUSIONS: We developed an MAPRS for BMI that represents a robust and generalizable risk factor for longitudinal weight gain in adulthood, providing a foundation for genetically informed risk stratification and earlier, more targeted obesity prevention strategies.

Humans

Characterizing features of the genetic architecture underlying autism from a multi-ancestry perspective.

Autism spectrum disorder (ASD; MIM 209850) is reported to vary globally from 0.01% in East Asian populations to 4.36% in certain Australian cohorts. Despite high heritability estimates (61-94%), the genetic architecture underlying ASD susceptibility remains poorly characterized across diverse populations, as most genomic studies have initially focused on individuals of European ancestry. To investigate ancestry-specific genetic contributions to ASD, we analyzed whole-genome sequencing data from three independent ASD cohorts. We identified admixed ASD probands (n&#x2009;=&#x2009;1 033) and ancestry-matched controls (n&#x2009;=&#x2009;1 033) and performed admixture mapping (AM). AM using five continental reference populations (European, African, East Asian, South Asian, and Native American) identified five ancestry-specific ASD-susceptibility loci, including one African-related locus at 1p21.2 near S1PR1 and four Native American-associated loci at chromosome 11q13.4. Three of these latter loci were contiguous and encompassed genes previously implicated in ASD, notably SHANK2 and DHCR7, with fine-mapping identifying a significantly associated variant between the two genes (rs77695321; P&#x2009;=&#x2009;1.52 &#xd7; 10&#x207b;&#x2077;). The fourth Native American-associated signal at 11q13.4 overlapped the folate receptor genes FOLR1 and FOLR3, with fine-mapping identifying a genome-wide significant variant (rs7950807; P&#x2009;=&#x2009;5.21 &#xd7; 10&#x207b;&#x2078;). A secondary admixture mapping analysis restricted to Latin American individuals, incorporating 6 487 Brazilian controls, identified 16 additional ancestry-specific loci across seven genomic regions.

Journal Article

Linkage between HLA-B8 and HLA-DQ2.5 Contributes to Ancestry-Dependent Genetic Risk for Celiac Disease.

BACKGROUND: Most genetic studies on celiac disease (CeD) have focused on individuals of European descent. Limited data are available for the Hispanic and black populations. METHODS: We analyzed whole-genome sequencing data, electronic health records (EHR), and laboratory results from the All of Us Research Program. We identified 3,481 individuals with CeD through EHR, self-reporting, or both. Of these, 2,899 carried one of the four well-established risk haplotypes, including 262 of admixed American (89% Hispanic) and 108 of African (70% black) ancestry. Five sex-, age-, and ancestry-matched controls per case were selected for the assessment of genetic and clinical risk factors. RESULTS: An enrichment in the DQB1*02:01 allele was observed in CeD patients across all ancestries, with the strongest association in Europeans (32.3% vs. 11.6%), followed by Americans (18.5% vs. 8.1%) and Africans (15.7% vs. 8.1%). Among individuals carrying the DQ2.5 (DQA1*05:01-DQB1*02:01 haplotype), HLA-B8 was present in 72.3% of Europeans, 42.3% of Admixed Americans, and lower in Africans. This linkage disequilibrium was higher in CeD patients than in controls across all three ancestries. A polygenic risk score distinguished seropositive CeD from controls with 86% accuracy. Incorporating clinical risk factors, including family history, hypothyroidism, diarrhea, vitamin D deficiency, and anemia, increased predictive accuracy to 92%. The model identified 93% of CeD patients with tTG-IgA levels greater than 10 IU/mL. CONCLUSION: Linkage between HLA-B8 and DQ2.5 differs significantly among individuals of European, admixed American, and African ancestry, contributing to ancestry-dependent genetic risk for CeD.

Celiac disease

Genetic Ancestry and Carrier Variant Frequency Enrichment in a Colombian Andean Population: Insights From the Eje Cafetero.

Colombia is one of the most genetically diverse populations in Latin America, and its demographic process has promoted the persistence and local enrichment of deleterious alleles, increasing the frequency of autosomal recessive disorders, particularly in semi-isolated Andean populations such as the Eje Cafetero. However, exome-based reference data from this region remain scarce, limiting ancestry-aware variant interpretation and carrier screening strategies. We aimed to characterize the ancestry proportions of this population using exome data, and to estimate the carrier frequency and distribution of pathogenic and likely pathogenic (P/LP) variants in clinically relevant recessive genes. We conducted a cross-sectional study with whole-exome sequencing (WES) in 316 unrelated individuals from the Colombian Eje Cafetero. P/LP variants were evaluated in 454 genes associated with autosomal recessive disorders. The global ancestry proportions were estimated using a validated panel of 250 exome-compatible ancestry-informative markers. Carrier frequencies were compared against Non-Finnish Europeans (NFE) and Admixed Americans (AMX) from gnomAD v4. The cohort showed predominant European ancestry (mean 51%), followed by Native American (36%) and African (13%) components. We identified 151 carriers of 89 distinct pathogenic variants across autosomal recessive genes. The most frequent variants were SERPINA1 c.863A>T (5.5%), CFTR c.1210-11T>G (3.5%), and PYGM c.1094C>T (1.5%). Also, recurrent variants were significantly enriched compared with both NFE and AMX populations, supporting regional founder effects. This study represents one of the most comprehensive exome-based genetic characterizations of the Colombian Eje Cafetero, revealing ancestry-specific enrichment of clinically relevant autosomal recessive variants driven by founder effects.

Female

Genomic consequences of admixture in an experimentally founded sand lizard population.

Conservation interventions are increasingly required for species threatened by population declines and isolation due to anthropogenic pressures. Small, isolated populations are particularly vulnerable to the loss of genetic diversity, increased inbreeding, and the accumulation of deleterious mutations. Translocations or supplementation of allopatric individuals for genetic rescue may be the only way to increase genetic diversity and increase population persistence via increased adaptive potential. Here, we use an experimentally admixed population of sand lizards on a small island in Sweden as a valuable model of genetic rescue. This population was established approximately 20 years ago (5-6 generations), resulting in increased fecundity and hatchling viability. This population was founded from crossings between individuals from an inbred population from the nearby mainland and individuals sourced from populations in southern Sweden. Low-coverage whole-genome sequencing revealed elevated genetic diversity and reduced realized genetic load in this admixed population relative to the source populations. Ancestry analyses indicated a greater contribution of southern Swedish genetic variation, potentially reflecting the contribution of beneficial adaptive variation from this region that may underlie the positive population effects. This system provides valuable empirical insights into the long-term genomic consequences of genetic rescue in this model vertebrate population.

Journal Article