PubMed HealthSearch

SEARCH · PubMed Health

Results for “genomic inflation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

16 recordsLinked to original sources

Multi-Ancestry Survival GWAS of Substance Use Initiation in the ABCD Study.

BACKGROUND: Substance use initiation in adolescence is influenced by both genetic and environmental factors; however, large-scale genetic studies often treat initiation as a binary outcome and underuse longitudinal timing information. METHODS: We conducted time-to-event (survival) genome-wide association analyses (GWAS) of initiation for four outcomes-alcohol, nicotine, cannabis, and any substance use-using longitudinal follow-up data from the Adolescent Brain Cognitive Development (ABCD) Study. We performed ancestry-stratified GWAS within European (EUR), African (AFR), and Hispanic (HISP) groups, applying consistent quality control and covariate adjustment. Summary statistics were harmonized across ancestries and meta-analyzed using inverse-variance weighted fixed-effects and DerSimonian-Laird random-effects models. We evaluated genomic inflation and heterogeneity (Cochran's Q and I 2), identified independent lead variants at genome-wide and suggestive significance thresholds, and assessed cross-trait overlap of associated loci. RESULTS: In the multi-ancestry meta-analysis, we observed suggestive association signals across traits (minimum p-values: alcohol ~ 1 &#xd7; 10-7, any ~ 1 &#xd7; 10-7, cannabis ~ 5 &#xd7; 10-8, nicotine ~ 1 &#xd7; 10-8). Nicotine initiation showed one genome-wide significant variant in both fixed- and random-effects meta-analyses (p < 5 &#xd7; 10-8). Across traits, suggestive loci demonstrated limited overlap, with the strongest concordance between alcohol and any substance use, consistent with shared liability. Heterogeneity statistics indicated that some loci exhibited cross-ancestry variation in effect estimates. CONCLUSIONS: Survival GWAS leveraging initiation timing can identify genetic signals that may be missed by binary designs and enables principled multi-ancestry synthesis. Our results highlight both shared and trait-specific genetic contributions to early substance initiation and provide a foundation for downstream functional annotation and integrative modeling with environmental risk factors. These findings demonstrate the value of incorporating developmental timing into genetic discovery and provide a framework for integrating longitudinal risk modeling with genomic analyses.

ABCD

Matching Heterogeneous Cohorts by Projected Principal Components Reveals Two Novel Alzheimer's Disease-Associated Genes in the Hispanic Population.

Alzheimer's disease (AD) is the most common form of dementia in elderly, affecting 6.9 million individuals in the United States. Some studies have suggested the prevalence of AD is greater in individuals who self-identify as Hispanic. Focused results are relevant for personalized and equitable clinical interventions. Ethnicity as a stratifying tool in genetic studies is often accompanied by genomic inflation due to heterogeneity. In this study, we report GWAS and meta-analyses conducted among NIAGADS subjects who self-identified as Hispanic and All of Us (AoU) sub-cohorts matched to that cohort, using projected genetically-derived principal components, with and without age and sex. In Hispanic NIAGADS subjects, we identified a common variant in PIEZO2 that was protective for AD with a p-value just beyond genome-wide significance (p = 5.4*10-8). Meta-analyses with genetically-matched AoU participants yielded three (two novel) genome-wide significant AD-associated loci based on rare lead variants: rs374043832 (RGS6/PSEN1), rs192423465 (ASPSCR1), and rs935208076 (GDAP2), which were also nominally significant in AoU sub-cohorts. We thus demonstrate an efficient way to select subjects from large heterogeneous biobank cohorts who are genetically similar to a smaller disease-specific cohort, yielding novel disease-relevant findings.

Journal Article

DNA methylation signatures in skeletal muscle associated with physical function in healthy older adults.

Despite the substantial variability in physical function among older adults, the molecular mechanisms remain poorly characterized, particularly within skeletal muscle. This study aimed to determine the patterns of DNA methylation in skeletal muscle associated with physical function in healthy older adults. We analyzed DNA methylation (EPIC v2 array; 875,554 CpG sites) in skeletal muscle from 92 healthy older adults (median age 74; 62% female). Associations were examined across five phenotypes: Short Physical Performance Battery (SPPB), 6-min walk test (6MWT), handgrip strength, perceived disability (PAT-D), and lifestyle health (modified Life's Essential 8). Linear regression models adjusted for age, sex, race, BMI, and muscle fiber composition. Genomic inflation corrected via the BACON method (FDR&#x2009;<&#x2009;0.05). Gene set enrichment analysis was performed on suggestive hits (FDR&#x2009;<&#x2009;0.1). We identified significant differentially methylated probes (DMPs) and regions (DMRs) across all phenotypes: SPPB (70 DMPs, 22 DMRs), 6MWT (16 DMPs, 566 DMRs), handgrip strength (2 DMRs), PAT-D (19 DMPs, 1 DMR), and lifestyle health (2 DMPs). DMRs largely overlapped promoters. Identified genes overlapped known musculoskeletal and neurological GWAS hits, including RUNX2 and FOXL1 (bone mineral density), IGFBP3 (muscle mass), and NEK1 and SHANK1 (neurological function). Enrichment analysis revealed that 6MWT-associated genes relate to nervous and skeletal system development, while handgrip-associated genes involve cytoskeletal dynamics and protein assembly. Epigenetic variation in aging skeletal muscle is associated with physical function. The enrichment of pathways related to nervous and musculoskeletal development suggests specific epigenetic mechanisms underlying functional decline, offering potential targets for intervention in older adults.

DNA methylation

Two-sample Mendelian randomization study of gut microbiota and inflammatory proteins: Predictive, preventive, and personalized treatment for migraine.

The human gut microbiota is increasingly recognized as a significant factor in the pathogenesis of migraine, potentially via inflammatory pathways. Identifying specific human gut microbiota components associated with migraines, along with the investigation of particular inflammatory proteins, is essential for advancing primary prediction, targeted prevention, and personalized treatment strategies for migraines. We conducted a two-sample Mendelian randomization study using publicly available summary statistics from genome-wide association studies. Data for 473 human gut microbiota taxa were obtained from the Finnish national health survey conducted by the National Institute for Health and Welfare study (FINRISK, n = 5959 European participants). Genome-wide association study data (https://www.ebi.ac.uk/gwas/) for 91 circulating inflammatory proteins were obtained from 14,824 participants across 11 cohorts using the Olink Target 96 Inflammation panel. Migraine outcome data were obtained from the FinnGen R12 release, with cases defined using ICD-10 code G43. All genome-wide association study analyses were adjusted for sex, age, genotyping batch, and 10 genetic principal components to control population stratification (genomic inflation factors: 1.00&#x2013;1.05). Inverse variance-weighted Mendelian randomization was the primary analysis method, with Mendelian randomization-Egger, weighted median, and mode-based methods as sensitivity analyses. Two-step Mendelian randomization mediation analysis quantified the proportion of the effects of human gut microbiota on migraine that are mediated through inflammatory proteins. Thirty-seven bacterial genera were found to be associated with migraine using the inverse variance-weighted method. Of these, 18 genera exhibited a negative association, while 19 genera demonstrated a positive association with migraine risk. Additionally, eight inflammatory proteins were found to increase the risk of migraine. Among human gut microbiota, four were observed to reduce inflammatory protein levels, whereas another four were associated with increased inflammatory protein levels. Additionally, five gut microbiota were identified to influence migraine through inflammatory proteins in both Mendelian randomization analyses. Specifically, Actinobacteria, Brachyspiraceae, CAG-269 sp001915995, and Paraglaciecola were found to affect migraine outcomes via inflammatory proteins, with mediation proportions of 12%, 19%, 15.5%, and 6.7%, respectively. Lawsonibacter sp002161175 was identified to influence migraine risk through Oncostatin-M and SLAM, with mediation proportions of 15.6% and 11.3%, respectively. Our study elucidated the role of specific human gut microbiota alterations in the pathogenesis of migraine and highlighted the mediating effects of inflammatory proteins. Targeting these particular human gut microbiota alterations offers a promising strategy for predictive, preventive, and personalized medicine in migraine management, resulting in substantial clinical advancements.

causality

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N &#x2248; 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1&#x2009;=&#x2009;0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n&#x2009;=&#x2009;305) demonstrated generalizability (AUROC = 0.8105; F1&#x2009;=&#x2009;0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC)&#x2009;=&#x2009;0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article

How negative sampling shapes the performance of transcription factor binding site prediction models.

MOTIVATION: Transcription factors (TFs) are key players in gene regulation and development, where they activate and repress gene expression through DNA binding. Predicting transcription factor binding sites (TFBSs) has long been an active area of research, with many deep learning methods developed to tackle this problem. These models are often trained on TF ChIP-seq data, which is generally seen as only providing positive samples. The choice of datasets and negative sampling techniques is a critical yet often overlooked aspect of this work. RESULTS: In this study, we investigate the impact of different negative sampling techniques on TFBS prediction performance. We create high-quality test datasets based on ChIP-seq and ATAC-seq data, where true negatives can be identified as positions that are accessible but not bound by the TF in question. We then train models using various negative sampling techniques, including genomic sampling, shuffling, dinucleotide shuffling, neighborhood sampling, and cell line specific sampling, simulating cases where matching ATAC-seq data is not available. Our results show that, generally, metrics calculated on training datasets give inflated performance scores. Of the tested techniques, genomic sampling of negatives based on similarity to the positives performed by far the best, although still not reaching the performance of baseline models trained on high-quality datasets. Models trained on dinucleotide shuffled negatives performed poorly, despite being a common practice in the field. Our findings highlight the importance of carefully selecting negative sampling techniques for TFBS prediction, as they can significantly impact model performance and the interpretation of results. AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/NatanTourne/TFBS-negatives (DOI: 10.5281/zenodo.18007567).

Binding Sites

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome&#x2011;wide coverage of only 0.1-5&#xd7;, sWGS data display a pronounced zero&#x2011;inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several&#x2011;fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy&#x2011;number gains (false positives), and true deletions often become indistinguishable from pervasive zero&#x2011;coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics

Intratumoral fungus Neurospora crassa is associated with worsened prognosis in ovarian cancer via modulation of extracellular matrix.

Landmark studies on intratumoral fungi (ITF) have raised concerns due to irreproducible results and data-analysis errors. We aimed to determine whether ITF exist in ovarian cancer (OvCa) and, if so, whether they play a role in disease biology. Formalin-fixed, paraffin-embedded OvCa samples and multiple controls underwent operational decontamination, qPCR, internal transcribed spacer sequencing, and post-hoc data decontamination. We also leveraged updated fungal reads from The Cancer Genome Atlas generated by the TCMbio group, which addressed human-read contamination and artificial inflation, to validate findings and assess prognostic associations. A murine syngeneic model established using mouse ovarian cancer cell line (OVHM) with intratumoral Neurospora crassa injection was established. Transcriptomic and metabolomic analyses were performed to explore mechanisms. Tumor-containing blocks harbored significantly higher fungal loads than environmental controls but had loads comparable to paraffin controls. Applying a two-pass decontamination filter reduced raw sequence features from 9289 amplicon sequence variants (ASVs) to 659 ASVs. We focused on high-abundance features present in human tissues but absent from xenografts and paraffin controls and identified one candidate, N. crassa, associated with unfavorable prognosis in OvCa. Integrating human and murine data, we found Neurospora correlated with eosinophils, whereas N. crassa itself was not immune-related. Neurospora crassa promoted OvCa progression with downregulation of integrin-linked kinase and decreased extracellular matrix-receptor interaction. Most ITF signals are likely contaminants. We identified N. crassa as associated with unfavorable prognosis in OvCa, potentially via modulation of the extracellular matrix.

Neurospora crassa

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals

Detecting Introgression in Shallow Phylogenies: How Minor Molecular Clock Deviations Lead to Major Inference Errors.

Recent theoretical and algorithmic advances in introgression detection, coupled with the growing availability of genome-scale data, have highlighted the widespread occurrence of interspecific gene flow across the tree of life. However, current methods largely depend on the molecular clock assumption-a questionable premise given empirical evidence of substitution rate variation across lineages. While such rate heterogeneity is known to compromise gene flow detection among divergent lineages, its impact on closely related taxa at shallow evolutionary timescales remains poorly understood, likely because these taxa are often assumed to adhere to a molecular clock. To address this gap, we combine theoretical analyses and simulations to evaluate the robustness of widely used site pattern methods (D-statistic and HyDe) to rate variation across phylogenetic timescales. Our results demonstrate that both methods exhibit high sensitivity to even minor deviations from the molecular clock at shallow timescales, complementing previous findings at deeper scales. Specifically, in young phylogenies (with an age of 3 &#xd7; 105 generations) with small population sizes, weak (17% difference) and moderate (33% difference) rate variation can inflate false-positive rates up to 35% and 100%, respectively, using site pattern counts from a 500&#x2005;Mb genome. Employing a more distant outgroup intensifies these spurious signals. Our study demonstrates that summary tests for introgression are pervasively vulnerable to minor rate variations and underscores the critical need for advanced methodologies to disentangle genuine introgression from false signals generated by rate heterogeneity.

Phylogeny

acmgscaler: an R package and Colab for standardized gene-level variant effect score calibration within the ACMG/AMP framework.

MOTIVATION: A genome-wide variant effect calibration method was recently developed under the guidelines of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP), following ClinGen recommendations for variant classification. While genome-wide approaches offer clinical utility, emerging evidence highlights the need for gene- and context-specific calibration to improve accuracy. Building on previous work, we have developed an algorithm tailored to converting functional scores from both multiplexed assays of variant effects (MAVEs) and computational variant effect predictors (VEPs) into ACMG/AMP evidence strengths. RESULTS: Our method is designed to deliver consistent performance across different genes and score distributions, with all variables adaptively determined from the input data, preventing selective adjustments or overfitting that could inflate evidence strengths beyond empirical support. To facilitate adoption, we introduce acmgscaler, a lightweight R package and a plug-and-play Google Colab notebook for the calibration of custom datasets. This algorithmic framework bridges the gap between MAVEs/VEPs and clinically actionable variant classification. AVAILABILITY AND IMPLEMENTATION: The R package and Colab notebook are available at https://github.com/badonyi/acmgscaler.

Software

HiCPotts: An R/Bioconductor package to identify significant interactions in chromosome conformation capture data and model sources of bias.

MOTIVATION: Chromosome Conformation Capture methods, including Hi-C, micro-C or Capture-C, are used to map chromatin interactions genome-wide. Most of the existing computational methods do not account for sources of bias (such as DNA accessibility, GC content or TE content) in the data. RESULTS: We previously developed ZipHiC, a Bayesian method based on the hidden Markov random field (HMRF) model and the Approximate Bayesian Computation (ABC), that uses zero-inflated Poisson distribution to model the noise, signal and false signal of the data and showed that this approach was able to detect bias from DNA accessibility, GC content and TE content in both Hi-C and micro-C data. Here, we present HiCPotts, another Bayesian method based on the HMRF model and the ABC that uses a zero-inflated Negative Binomial distribution instead to model the noise and signal of the data. We systematically show that HiCPotts reduces false positives and increases recovery of true interactions compared to ZipHiC, but also compared to other methods such as FastHiC, Juicer and HiCExplorer. Most importantly, we provide an R/Bioconductor package that allows modelling the noise, signal and false signal using various distributions such as the zero-inflated Negative Binomial (ZINB) and the zero-inflated Poisson distribution (ZIP). AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/HiCPotts/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Approximate Bayesian Computation

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans