PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genotype imputation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Selphi, a tool for improving genotype imputation accuracy.

Genotype imputation is a powerful tool for inferring missing genotype data in large-scale genetic studies. Over the last two decades, multiple imputation algorithms have been developed, steadily improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge, largely because existing methods rely on local haplotype matching within genomic windows and do not fully exploit the extended patterns of haplotype sharing that span entire chromosomes. Here we present Selphi, a new genotype imputation algorithm that combines the Positional Burrows-Wheeler Transform (PBWT) with a multi-stage haplotype selection heuristic operating across entire chromosomes. When compared to state-of-the-art methods Beagle 5.4, IMPUTE5, and Minimac4, Selphi showed higher accuracy on the 1000 Genomes Project and TOPMed datasets, across all super-populations and allele frequencies. Similarly, Selphi achieved higher accuracy than Beagle 5.4 on the UK Biobank dataset, which translated into improved concordance with hc-WGS GWAS summary statistics at known trait-associated loci and more accurate polygenic risk scores (PRS). Selphi outputs standard VCF files with genotype dosages (DS), haplotype-specific allele probabilities (AP1, AP2), and a per-variant dosage R-squared quality score (DR2), enabling direct integration with downstream analytical pipelines including standard post-imputation quality filtering.

Genome-Wide Association Study

Effect of founder breeds on genotype imputation accuracy in Canchim cattle.

UNLABELLED: Genotype imputation is a technique used to infer unobserved genotypes based on reference panels, allowing increased marker density and cost-effective optimization for genomic selection. This study aimed to evaluate whether the inclusion of genotypes from the founder breeds Nelore (NE) and Charolais (CH) improves the imputation accuracy in the composite beef cattle breed Canchim (CA). The populations studied consisted of 804 NE, 897 CH, and 392 CA animals, all genotyped using high-density panels (777,962 SNP – single nucleotide polymorphisms). CA animals had their genotypes masked to simulate a medium-density panel (54,609 SNP). Fourteen imputation scenarios were evaluated, varying according to breed, sex, year of birth, and lineage. Imputation accuracy was determined based on the percentage of correctly imputed genotypes (PERC) and the squared Pearson’s correlation between observed and imputed genotypes (R2). PERC values ranged from 66.52% to 97.39% and R² from 0.6352 to 0.9780. The scenarios that included NE, CH, and CA (males or animals born before 2004) as the reference population for imputing CA females or CA animals born after 2004 showed the highest imputation accuracies. Therefore, the use of founder breeds in the reference population improves the accuracy of genotype imputation in CA cattle. The results indicate that a multibreed reference population, incorporating founder breeds, could provide a more robust and informative genetic basis for imputing composite cattle. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13353-026-01060-z.

Animal breeding

The Soifua Manuia reference panel with 2,570 Samoan haplotypes improves genotype imputation quality among Samoans.

Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we used whole-genome sequencing from 1,285 Samoan individuals combined with 1000 Genomes Project (1KGP) individuals to construct an imputation reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. Here we show that this panel yielded up to two times more well-imputed (r2 ≥ 0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum; accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in accuracy over TOPMed-R3 was largest for small haplotypes, reflecting the Samoan panel's ability to capture variation not well tagged by other panels.

Haplotypes

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

Pedigree-assisted genotype imputation enables cost-effective genomic prediction in Penaeus vannamei.

Genomic selection in Penaeus vannamei has long been constrained by the high cost of dense genotyping. To address this limitation, we evaluated genotype imputation from a low-density 1 K panel to a medium-density 55 K panel of the "Yellow Sea Array No. 1" and examined its impact on genomic prediction for harvest body weight in P. vannamei. A four-generation pedigree including 30 great-grandparents, 39 grandparents, 100 parents, and 608 offspring was genotyped using the 55 K panel. A two-step experimental design was implemented to (i) assess the performance of different imputation algorithms under reference population scenarios with varying proportions of siblings, and (ii) compare six alternative reference population structures incorporating parents, ancestors, and siblings. Genotype imputation using the pedigree-based method FImpute v3.0 consistently achieved higher accuracy than the population-based method Beagle v5.5. Using this pedigree-assisted approach, imputation accuracy increased from 0.73 when only parental genotypes were used to 0.84 with the inclusion of 10% siblings, and subsequently plateaued at 0.87-0.90 when sibling representation reached 20%. Across the six reference population structures, imputation accuracy was primarily driven by the availability of parental genotypes, ranging from 0.50 to 0.56 in the absence of parents to 0.88-0.89 when both parents and ancestral generations were included. Accuracy remained high when both parents were available (0.84-0.87 with siblings; 0.73 without siblings) but declined substantially when only one parent was genotyped (0.65-0.68). Imputation accuracy was positively associated with both minor allele frequency (MAF) and linkage disequilibrium (max r2LD), with LD exerting the stronger influence. Heritability estimates derived from imputed 55 K genotypes were highly consistent with those obtained from the original 55 K data (0.39 ± 0.14 vs. 0.41 ± 0.14), indicating that genotype imputation did not compromise variance component estimation. In predictive ability analyses, pedigree-based BLUP (PBLUP) achieved higher predictive ability than genomic BLUP (GBLUP) based on the 1 K panel, with predictive abilities of 0.42-0.44 for PBLUP compared with 0.34-0.35 for GBLUP. Using imputed genotypes for genomic prediction further improved predictive ability relative to the true 1 K panel, yielding values ranging from 0.35 to 0.47. Notably, when parental genotypes were included in the reference population, GBLUP based on imputed genotypes surpassed the predictive ability of PBLUP and approached that achieved with the original 55 K genotypes (0.45-0.47). Collectively, these results provide the first empirical evidence that low- to medium-density genotype imputation, combined with pedigree information, can effectively support genomic prediction in P. vannamei. This study establishes a cost-efficient and scalable framework for implementing genomic selection in P. vannamei and provides a practical reference for the application of genomic selection in other aquaculture species with constrained breeding budgets.

Animals

GCRP: Integrated Global Chicken Reference Panel from 11,951 Chicken Genomes.

Chickens are a crucial source of protein for humans and a popular model animal for bird research. Despite the emergence of imputation as a reliable genotyping strategy for large populations, the lack of a high-quality chicken reference panel has hindered progress in chicken genome research. To address this, here we introduce the first phase of the 100K Global Chicken Reference Panel (100K GCRP). Currently, two panels are available: a comprehensive mix panel (CMP) for domestication diversity research and a commercial breed panel (CBP) for breeding broilers specifically. Evaluation of genotype imputation quality showed that CMP had the highest imputation accuracy compared to imputation using existing chicken panels in Animal-SNPAtlas and Animal Genotype Imputation Database (AGIDB), whereas CBP performed stably in the imputation of commercial populations. Additionally, we found that genome-wide association studies using GCRP-imputed data, whether on simulated or real phenotypes, exhibited greater statistical power. In conclusion, our study indicates that the GCRP effectively fills the gap in high-quality reference panels for chickens, providing an effective imputation platform for future genetic and breeding research. The project includes 11,951 samples and provides services for various applications on its website at http://farmrefpanel.com/GCRP/#/.

Animals

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Archaic ancestry inference in imputed ancient human genomes.

When modern humans expanded from Africa into Eurasia, they interbred with archaic hominins such as Neanderthals and Denisovans. This introgression shaped human evolution, yet most insights have been gained from present-day genomes, leaving little known about how archaic variants evolved after interbreeding. Ancient genomes offer a direct view of this process, but low coverage and poor quality have limited their use. Recent advances in genotype imputation offer a way to overcome these challenges by reconstructing missing information from reference panels and recovering evolutionary signals from low-coverage data. Here, we show that imputation enables accurate detection and quantification of archaic introgression in ancient genomes, improves local archaic ancestry inference, and that regions of archaic ancestry are imputed with especially high accuracy. We further demonstrate that imputed genomes can reconstruct the trajectories of introgressed haplotypes, distinguish populations across time and geography, and identify both known and additional candidates for adaptive introgression.

Humans

Variant harmonization critically determines polygenic score transferability for lipid traits in Samoan populations.

Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa. Polygenic scores (PGSs) for lipid traits offer promise for improved CVD risk prediction; however, their performance in Pacific Islander populations-comprising only 0.002% of genome-wide association study (GWAS) participants as of 2024-remains unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL-C), HDL cholesterol (HDL-C), triglycerides (TGs), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990-2010. PGSs from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed via incremental R2 from linear mixed models with bootstrapped confidence intervals. HDL-C showed the highest performance (incremental R2 5.0%-15.0%), followed by TC (5.0%-10.7%), LDL-C (5.7%-8.6%), and TG (3.5%-7.0%). Critically, meaningful LDL-C performance was achieved only with the genome-wide PRS-CS score (99.6%-99.7% variant matching), while a curated pruning-and-thresholding score achieved ∼9% matching and near-zero performance. These findings establish systematic lipid PGS benchmarks in Samoans, demonstrating meaningful transferability when genome-wide variant coverage is ensured, and highlight variant harmonization as a critical precondition for PGS deployment in underrepresented populations.

Pacific Islanders

New Genetic Loci Implicated in Cardiac Morphology and Function Using Three-Dimensional Population Phenotyping.

BACKGROUND: Cardiac remodeling occurs in the mature heart and is a cascade of adaptations in response to stress, which are primed in early life. A key question remains as to the processes that regulate the geometry and motion of the heart and how it adapts to stress. METHODS: We performed spatially resolved phenotyping using machine learning-based analysis of cardiac magnetic resonance imaging in 47 549 UK Biobank participants. We analyzed 16 left ventricular spatial phenotypes, including regional myocardial wall thickness and systolic strain in both circumferential and radial directions. In up to 40 058 participants, genetic associations across the allele frequency spectrum were assessed using genome-wide association studies with imputed genotype participants, and exome-wide association studies and gene-based burden tests using whole-exome sequencing data. We integrated transcriptomic data from the GTEx project and used pathway enrichment analyses to further interpret the biological relevance of identified loci. To investigate causal relationships, we conducted Mendelian randomization analyses to evaluate the effects of blood pressure on regional cardiac traits and the effects of these traits on cardiomyopathy risk. RESULTS: We found 42 loci associated with cardiac structure and contractility, many of which reveal patterns of spatial organization in the heart. Whole-exome sequencing revealed 3 additional variants not captured by the genome-wide association study, including a missense variant in CSRP3 (minor allele frequency 0.5%). The majority of newly discovered loci are found in cardiomyopathy-associated genes, suggesting that they regulate spatially distinct patterns of remodeling in the left ventricle in an adult population. Our causal analysis also found regional modulation of blood pressure on cardiac wall thickness and strain. CONCLUSIONS: These findings provide a comprehensive description of the pathways that orchestrate heart development and cardiac remodeling. These data highlight the role that cardiomyopathy-associated genes have on the regulation of spatial adaptations in those without known disease.

Humans

Alterations in ether lipid metabolism in obesity revealed by systems genomics of multi-omics datasets.

Ratios between two metabolites are sensitive indicators of metabolic changes. Lipidomic profiling studies have revealed that plasma ether lipids, a class of glycero- and glycerophospho-lipids with reported health benefits, are negatively associated with obesity. Here, we utilized lipid ratios as surrogate markers of lipid metabolism to explore the processes underlying the inverse relationship between ether lipid metabolism and obesity. Plasma lipidomics data from two independent human cohorts (n = 10,339 and n = 4,492) were integrated to assess the associations between 82 lipid ratios and obesity-related markers in males and females. Results were externally validated using mouse transcriptomics data from the Hybrid Mouse Diversity Panel (n = 152-227 across 74 strains). Genome-wide association studies using imputed genotypes from a population cohort (n = 4,492) were performed to examine the genetic architecture of the ratios. Findings showed that waist circumference (WC), body mass index, and waist-hip ratio were inversely associated with total plasmalogens relative to total phospholipids in both sexes. Ratios comprising product-substrate pairs positioned either side of enzymes involved in plasmalogen synthesis and degradation showed positive and negative associations with WC, respectively. Branched-chain fatty acids negatively correlated with WC, while omega-6 polyunsaturated fatty acids exhibited differing associations depending on their position within the pathway. Mouse transcriptomics corroborated these results. Genomics data showed strong associations between ratios containing choline-plasmalogens and single-nucleotide polymorphisms in the transmembrane protein 229B (TMEM229B) gene region. This work demonstrates the utility of lipid ratios in understanding lipid metabolism. By applying the ratios to multi-omic datasets, we identified alterations in enzymatic activity and genetic variants likely affecting ether lipid synthesis in obesity that could not have been obtained from lipidomics data alone. Additionally, we characterized a potential role for TMEM229B, offering new perspectives on ether lipid metabolism and regulation.

Humans

Genetic Determinants of Pulmonary Artery Size in over 50,000 Subjects with and without COPD.

RATIONALE: Pulmonary artery (PA) enlargement is a non-invasive imaging biomarker associated with pulmonary hypertension and mortality in COPD; however, its genetic determinants remain incompletely understood. OBJECTIVES: To characterize the genetic architecture of PA size across COPD-enriched and population-based cohorts. METHODS: We performed genome-wide association analyses of PA diameter using whole-genome sequencing in COPDGene (n=9,418) and ECLIPSE (n=1,859), and imputed-genotype data from the UK Biobank (n=37,073). We replicated lead variants in the Framingham Heart Study (FHS; n=3,289), incorporated all four studies into a joint meta-analysis, and identified independent signals through conditional analyses. Candidate effector genes were prioritized using coding variant annotation, colocalization, and integrative regulatory evidence. MEASUREMENTS AND MAIN RESULTS: We identified 44 independent genome-wide significant PA diameter signals within 39 loci, including 8 variants replicated in FHS, novel associations near FRMD4B, SLC20A2, BORCS7-ASMT, and KCNRG, and 5 signals in conditional analysis including multiple signals at ANO1. Genetic effects were concordant across imaging modalities and cohorts of differing COPD burden. Effector-gene prioritization nominated ABCC8, PDGFD, HMCN1, CCNE1, and TBX20, implicating pathways in vascular remodeling, developmental regulation, smooth muscle and endothelial function, ion-channel signaling, and extracellular matrix organization. Colocalization with pulse pressure GWAS demonstrated substantial shared causal variation between pulmonary and systemic vascular biology. CONCLUSIONS: In this largest genetic study of pulmonary vascular imaging to date, PA diameter exhibits a polygenic architecture consistent across imaging modalities and cohorts of differing COPD burden. The prioritized effector genes bridge rare-variant pulmonary hypertension biology with common-variant systemic vascular biology.

Pulmonary artery diameter

MetaGLIMPSE: Meta-imputation of low-coverage sequencing data for modern and ancient genomes.

The advent of efficient and accurate imputation for low-coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between target and reference samples, leveraging imputation from multiple reference panels improves imputation accuracy; however, individual reference panel genotypes are often privacy protected. Meta-imputation bypasses individual-level data by combining single-panel imputed genotypes through estimating panel- and marker-specific weights. We present a meta-imputation method, MetaGLIMPSE, that combines estimates from multiple reference panels for low-coverage sequencing imputation. Across all our scenarios, for both modern and ancient DNA samples, MetaGLIMPSE consistently outperforms the best single-panel imputation for coverages of 0.1×-8× and across all minor-allele frequencies, equaling the combined panel imputation for some parameters. Finally, MetaGLIMPSE is computationally efficient, meta-imputing 500 whole genomes in 16% of the time of GLIMPSE2.

Humans

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals

A model of assortative mating with partial dominance.

A model of assortative mating incorporating partial dominance is proposed for a single locus with two alleles. It is derived by starting from an arbitrary genotypic distribution and finding symmetric and non-selective mating frequencies which duplicate this distribution. Numerical values are imputed to genotypes, the homozygotes having numerically equal values, opposite in sign, and the heterozygote having a value determined by the gene and heterozygote frequencies. The model is specified in a canonical form which reveals the correlation between mates based on genotypic values, and relates the correlation to the fixation index. It permits negative as well as positive values of the fixation index. It is shown that this general model includes several particular cases, in equilibrium phase, occurring in the literature.

Alleles

Genetic Architecture of Idiopathic Inflammatory Myopathies From Meta-Analyses.

OBJECTIVE: Idiopathic inflammatory myopathies (IIMs, myositis) are rare systemic autoimmune disorders that lead to muscle inflammation, weakness, and extramuscular manifestations, with a strong genetic component influencing disease development and progression. Previous genome-wide association studies identified loci associated with IIMs. In this study, we imputed data from two prior genome-wide myositis studies and analyzed the largest myositis data set to date to identify novel risk loci and susceptibility genes associated with IIMs and its clinical subtypes. METHODS: We performed association analyses on 14,903 individuals (3,206 patients and 11,697 controls) with genotypes and imputed data from the Trans-Omics for Precision Medicine reference panel. Fine-mapping and expression quantitative trait locus colocalization analyses in myositis-relevant tissues indicated potential causal variants. Functional annotation and network analyses using the random walk with restart (RWR) algorithm explored underlying genetic networks and drug repurposing opportunities. RESULTS: Our analyses identified novel risk loci and susceptibility genes, such as FCRLA, NFKB1, IRF4, DCAKD, and ATXN2 in overall IIMs; NEMP2 in polymyositis; ACBC11 in dermatomyositis; and PSD3 in myositis with anti-histidyl-transfer RNA synthetase autoantibodies (anti-Jo-1). We also characterized effects of HLA region variants and the role of C4. Colocalization analyses suggested putative causal variants in DCAKD in skin and muscle, HCP5 in lung, and IRF4 in Epstein-Barr virus (EBV)-transformed lymphocytes, lung, and whole blood. RWR further prioritized additional candidate genes, including APP, CD74, CIITA, NR1H4, and TXNIP, for future investigation. CONCLUSION: Our study uncovers novel genetic regions contributing to IIMs, advancing our understanding of myositis pathogenesis and offering new insights for future research.

Humans

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., &#x2265;10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals

Common genetic variants associated with urinary phthalate levels in children: A genome-wide study.

INTRODUCTION: Phthalates, or dieters of phthalic acid, are a ubiquitous type of plasticizer used in a variety of common consumer and industrial products. They act as endocrine disruptors and are associated with increased risk for several diseases. Once in the body, phthalates are metabolized through partially known mechanisms, involving phase I and phase II enzymes. OBJECTIVE: In this study we aimed to identify common single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) associated with the metabolism of phthalate compounds in children through genome-wide association studies (GWAS). METHODS: The study used data from 1,044 children with European ancestry from the Human Early Life Exposome (HELIX) cohort. Ten phthalate metabolites were assessed in a two-void pooled urine collected at the mean age of 8&#xa0;years. Six ratios between secondary and primary phthalate metabolites were calculated. Genome-wide genotyping was done with the Infinium Global Screening Array (GSA) and imputation with the Haplotype Reference Consortium (HRC) panel. PennCNV was used to estimate copy number variants (CNVs) and CNVRanger to identify consensus regions. GWAS of SNPs and CNVs were conducted using PLINK and SNPassoc, respectively. Subsequently, functional annotation of suggestive SNPs (p-value&#xa0;<&#xa0;1E-05) was done with the FUMA web-tool. RESULTS: We identified four genome-wide significant (p-value&#xa0;<&#xa0;5E-08) loci at chromosome (chr) 3 (FECHP1 for oxo-MiNP_oh-MiNP ratio), chr6 (SLC17A1 for MECPP_MEHHP ratio), chr9 (RAPGEF1 for MBzP), and chr10 (CYP2C9 for MECPP_MEHHP ratio). Moreover, 115 additional loci were found at suggestive significance (p-value&#xa0;<&#xa0;1E-05). Two CNVs located at chr11 (MRGPRX1 for oh-MiNP and SLC35F2 for MEP) were also identified. Functional annotation pointed to genes involved in phase I and phase II detoxification, molecular transfer across membranes, and renal excretion. CONCLUSION: Through genome-wide screenings we identified known and novel loci implicated in phthalate metabolism in children. Genes annotated to these loci participate in detoxification, transmembrane transfer, and renal excretion.

Humans