PubMed HealthSearch

SEARCH · PubMed Health

Results for “summary statistics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Guide for Exploring Pleiotropic Associations in Genome-Wide Association Studies Using Summary Statistics.

Genome-wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross-phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic-based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.

Genome-Wide Association Study

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans

STABIX: summary-statistic-based GWAS indexing and compression.

MOTIVATION: Genome-wide association studies (GWAS) are widely used to investigate the role of genetics in disease traits, but the resulting file sizes from these studies are large, posing barriers to efficient storage, sharing, and querying. This issue is especially important for biobanks like the UK Biobank that publish GWAS for thousands of traits, increasing the volume of data that must be effectively managed. Current compression and query methods reduce file sizes and allow for quick genomic position-based queries but do not provide utility for quickly finding loci based on their summary statistics. For example, finding all SNVs in a particular p-value range would require decompressing and scanning the whole file. We propose a new tool, STABIX, which introduces summary-statistic-based queries and improves upon the standard bgzip compression and Tabix query tool in both compression ratio and decompression speed. RESULTS: When applied to 10 GWAS files from PanUKBB, STABIX created smaller compressed data and indices than Tabix for all files, where bgzip and tbi files were an average of 1.2 times the size of STABIX compressed files and indexes. In the same 10 files, STABIX per gene decompression was, on average 7× faster than Tabix per gene decompression, and achieved faster per gene decompression times for over 99% of nearly 20,000 genes. AVAILABILITY AND IMPLEMENTATION: Software freely available for download at GitHub: https://github.com/kristen-schneider/stabix/.

Genome-Wide Association Study

Shared genetic architecture of smoking dependence and Crohn's disease: A cross-trait analysis of GWAS summary statistics.

INTRODUCTION: Smoking dependence (SD) and Crohn's disease (CD) are epidemiologically associated, but whether this relationship reflects shared genetic susceptibility remains unclear. METHODS: We conducted a cross-trait genetic analysis of SD and CD using publicly available genome-wide association study (GWAS) summary statistics from European-ancestry populations. Genome-wide genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Pleiotropic variants were identified using PLACO and mapped to genomic loci using FUMA. Regional signal sharing was assessed by Bayesian colocalization. Functional analyses included stratified LDSC, Multi-marker Analysis of GenoMic Annotation (MAGMA), GTEx tissue analysis, and Metascape. Expression-linked candidate genes were prioritized using expression quantitative trait locus (eQTL)-based summary-data-based Mendelian randomization (SMR) with heterogeneity in dependent instruments (HEIDI) testing. Genetically informed spatial mapping of cells for complex traits (gsMap) was used for spatial mapping. RESULTS: SD and CD showed positive genetic correlation by LDSC (rg=0.2090, p=0.0008) and HDL (rg=0.3817, p=0.00106). PLACO identified 81 genome-wide significant pleiotropic SNPs, which were mapped by FUMA to three loci at 1p31.3, 5p13.1, and 12q12, represented by rs11209031, rs1395152, and rs17467116, respectively. MAGMA identified 22 FDR-significant genes, four of which remained Bonferroni significant: LRRK2, TNFRSF6B, ZGPAT, and RP4-583P15.15. Cross-trait tissue analysis showed significant enrichment of the shared genetic signal in whole blood and small intestine, while gene-set analysis highlighted inflammatory response (pbon=1.86×10-5) and T-helper 17 cell differentiation (pbon=7.37×10-4). SMR/HEIDI analysis further prioritized RPS6KB1 as a shared expression-linked candidate. Spatial mapping revealed a prominent signal in the embryonic gastrointestinal tract and gene-specific regional patterns involving LRRK2 and SLC2A13 in the adult mouse brain. CONCLUSIONS: SD and CD showed measurable shared genetic susceptibility, with convergent evidence from pleiotropic loci, immune-inflammatory pathway enrichment, tissue-level associations, and spatial transcriptomic mapping.

Crohn's disease

Beacon Reconstruction Attack: Reconstruction of genomes in genomic data-sharing beacons using summary statistics.

MOTIVATION: Genomic data-sharing beacon protocol, developed by the Global Alliance for Genomics and Health, offers a privacy-preserving mechanism for querying genomic datasets while restricting direct data access. Despite their design, beacons remain vulnerable to privacy attacks. This study introduces a novel privacy vulnerability of the protocol: one can reconstruct large portions of the genomes of all beacon participants by only using the summary statistics reported by the protocol. RESULTS: We introduce a novel optimization-based algorithm that leverages beacon responses and SNP correlations for reconstruction. By optimizing for the SNP correlations and allele frequencies, the proposed approach achieves genome reconstruction with a substantially higher F1-score (70%) compared to baseline methods (45%) on beacons generated using individuals from the HapMap and OpenSNP datasets. We show that reconstructed genomes can be used by downstream applications such as in membership inference attacks against other beacons. Our findings reveal that beacons releasing allele frequencies substantially increase the reconstruction risk, underscoring the need for enhanced privacy-preserving mechanisms to protect genomic data. AVAILABILITY AND IMPLEMENTATION: Our implementation is available at https://github.com/ASAP-Bilkent/Beacon-Reconstruction-Attack.

Genomics

Causal effect of chloride intracellular channel protein 5 on chronic periodontitis: A Mendelian randomization study.

This study aimed to evaluate the potential causal effect of chloride intracellular channel protein 5 (CLIC5) on the risk of chronic periodontitis (CP) using a Mendelian randomization (MR) approach. MR analysis was conducted utilizing publicly available summary statistics from genome-wide association studies summary statistics for CLIC5 and CP. Multiple MR methods, including inverse variance weighted, MR Egger, weighted median and weighted mode, were employed to estimate the causal effects. Sensitivity analyses, comprising leave-one-out and heterogeneity assessments were performed to evaluate the robustness of our findings. This MR analysis consistently revealed a negative association between CLIC5 and CP, with statistical significance achieved using the inverse variance weighted and weighted median methods. The concordant effect estimates obtained from all methodological approaches collectively indicated a potential protective effect of CLIC5 against CP. The sensitivity analyses further confirmed the robustness of these findings. This study provides genetic evidence suggesting a potential causal association between increased CLIC5 levels and decreased risk of CP. These findings augment the existing literature implicating chloride channel proteins in modulating the inflammatory processes pertinent to periodontal health. Further investigation is warranted to decipher the underlying biological mechanisms and to explore the potential of CLIC5 as a therapeutic target for CP.

Chloride Channels

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

Agnostic polygenic prediction of weight loss after bariatric surgery.

A large interindividual variability in weight loss outcomes following bariatric surgery is reported. To ensure optimal management of patients, it is crucial to accurately identify candidates most likely to benefit the most from the intervention. Since genetic variants largely contribute to surgery response, polygenic scores (PGS) derived from genome-wide association studies (GWAS) could constitute valuable tools for clinical decision making. We developed and evaluated PGS to predict the weight loss response in 540 patients with a body mass index (BMI) of 35 kg/m2 or higher who underwent biliopancreatic diversion with duodenal switch. Summary statistics derived from BMI-derived GWAS, together with summary statistics from previously published GWAS of BMI and adiposity features, were used to construct, evaluate, and benchmark weight loss PGS. The full-adjusted BMI PGS model built in the entire cohort explained 39.6% of the mean-over-time excessive body weight loss (%EBWL), while the BMI-PGS built in the training dataset explained 38.9%. All benchmarked PGS based on BMI showed a significant relationship with mean-over-time %EBWL. These findings highlight the potential of BMI PGS in predicting weight loss after bariatric surgery and support their use as promising tools to improve the effectiveness of future antiobesity treatments.

Humans

Freely available genomic datasets for atrial fibrillation research: current resources and analytical pipeline.

Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, characterized by clinical and genetic heterogeneity. Increasing use of genomics and other omics approaches has driven reliance on publicly available AF datasets to advance biological discovery. Thus, this systematic review aimed to identify freely available genomic AF datasets through Mendeley Data and its interconnected repositories, and to characterize the most common analyses performed on these data. The search was conducted in adherence to the PRISMA 2020 guideline. Nineteen freely available genomic AF datasets were identified: Summary statistics for 'Biobank-driven genomic discovery yields new insight into atrial fibrillation biology', hum0014.v8.58qt.v1, AF GWAS in UK Biobank, UK Biobank (Publication 9659), GWAS summary statistics from a 2025 multi-ancestry AF meta-analysis, GSE115574, GSE128188, GSE14975, GSE2240, GSE238242, GSE254133, GSE261170, GSE271748, GSE271839, GSE293813, GSE294456, GSE31821, GSE41177, and GSE79768. The GEO datasets were further examined using differential gene expression, functional enrichment, protein-protein interaction networks, hub gene analysis, microRNA target prediction, and gene clustering, as well as, for the more recently deposited datasets, eQTL colocalization, single-cell/single-nucleus clustering, cell-cell communication analysis, and gene-dosage-dependent transcriptional and electrophysiological profiling. These analyses show some consistency but also considerable heterogeneity in initial conditions, data normalization, and analytical methodological settings. In conclusion, only a limited number of datasets are freely available, so additional, well-characterized and standardized datasets are needed to provide a complete picture of the AF pathology.

Mendeley Data

From aerial drone to quantitative trait locus: leveraging next-generation phenotyping to reveal the genetics of color and height in field-grown Lactuca sativa.

In recent years, accurate and low-cost variant calling has enabled the genotyping of large diversity panels for genome-wide association studies. As a result, phenotyping rather than genotyping is now the rate-limiting step, especially in field experiments. This has created a strong need for high-throughput, accurate, and low-cost in-field phenotyping. Here, we present a genome-wide association study (GWAS) study on 194 field-grown accessions of lettuce (Lactuca sativa). These accessions were non-destructively phenotyped at two time points 15 days apart using a drone equipped with an RGB and multispectral (MSP) camera. Our high-throughput phenotyping approach integrates an RGB- and MSP camera to measure the color and height of lettuce in this large-scale field experiment. We used the mean and other summary statistics, such as median, quantiles, skewness, kurtosis, minimum, and maximum to quantify different aspects of color and height variation in lettuce from the drone images. Using these summary statistics as traits for GWAS, we confirm several previously described genetic associations, now under field conditions, and identify additional novel associations for color and height traits in lettuce.

Lactuca

Genetic Relationship Between Endometriosis and Melanoma.

Epidemiological studies have observed that risk of endometriosis is associated with history of cutaneous melanoma and vice versa. Evidence for shared biological mechanisms between the two traits is limited. The aim of this study was to investigate the genetic correlation and causal relationship between endometriosis and melanoma. Summary statistics from genome-wide association meta-analyses (GWAS) for endometriosis and melanoma were used to estimate the genetic correlation between the traits and Mendelian randomization was used to test for a causal association. When using summary statistics from separate female and male melanoma cohorts we identified a significant positive genetic correlation between melanoma in females and endometriosis (r g = 0.144, se = 0.065, p = 0.025). However, we find no evidence of a correlation between endometriosis and melanoma in males or a combined melanoma dataset. Endometriosis was not genetically correlated with skin color, red hair, childhood sunburn occasions, ease of skin tanning, or nevus count suggesting that the correlation between endometriosis and melanoma in females is unlikely to be influenced by pigmentary traits. Mendelian Randomization analyses also provided evidence for a relationship between the genetic risk of melanoma in females and endometriosis. Colocalization analysis identified 27 genomic loci jointly associated with the two diseases regions that contain different causal variants influencing each trait independently. This study provides evidence of a small genetic correlation and relationship between the genetic risk of melanoma in females and endometriosis. Genetic risk does not equate to disease occurrence and differences in the pathogenesis and age of onset of both diseases means it is unlikely that occurrence of melanoma causes endometriosis. This study instead provides evidence that having an increased genetic risk for melanoma in females is related to increased risk of endometriosis. Larger GWAS studies with increased power will be required to further investigate these associations.

endometriosis

Myeloid Dendritic Cell Counts and Coronary Heart Disease: a Bidirectional Mendelian Randomization Study.

BACKGROUND: Coronary heart disease (CHD) remains a leading cause of morbidity and mortality worldwide, with immune and inflammatory mechanisms playing important roles in its pathogenesis. Dendritic cells (DCs) are key regulators of immune responses; however, the relationship between specific DC subsets and CHD risk remains incompletely understood. METHODS: This study conducted a bidirectional two-sample Mendelian randomization (MR) analysis using publicly available genome-wide association study (GWAS) summary statistics to investigate the potential associations between circulating dendritic cell traits and CHD. Genetic instruments for myeloid dendritic cells (Myeloid DCs) and plasmacytoid dendritic cells (Plasmacytoid DCs), including both absolute counts and relative proportions, were obtained from immune cell GWAS datasets. Summary statistics for CHD were derived from a large European-ancestry population. Multiple MR methods were applied, and sensitivity analyses were performed to assess the robustness of the findings and potential pleiotropic effects. RESULTS: Nominal associations between genetically predicted Myeloid DC counts and CHD risk were observed in the MR-Egger and weighted median analyses, whereas the inverse variance weighted analysis demonstrated no significant association. These nominal associations did not remain statistically significant after correction for multiple testing. No significant associations were observed for Plasmacytoid DC counts or for the relative proportions of either DC subset. Reverse MR analyses were inconclusive due to wide confidence intervals, precluding meaningful inference regarding a causal effect of CHD on DC-related traits. Sensitivity analyses revealed no substantial heterogeneity or horizontal pleiotropy. CONCLUSIONS: This bidirectional MR study explored the potential relationships between circulating dendritic cell traits and CHD risk. Although nominal associations involving Myeloid DC counts were observed in secondary MR analyses, no robust evidence supporting an association remained after correction for multiple testing. Further studies using larger datasets and functional approaches are warranted to clarify the role of dendritic cells in CHD.

Humans

Genetic Evidence Links Sex Hormone-binding Globulin to Total Body Bone Mineral Density at Age 45-60 Years: A Two-sample Mendelian Randomization Study.

The menopausal transition and early postmenopause represent important periods for women's skeletal health, but the genetic relevance of metabolic, behavioral, and hormone-related factors to bone mineral density during midlife remains incompletely understood. This study used publicly available genome-wide association study summary statistics to examine associations between body mass index, 25-hydroxyvitamin D, sex hormone-binding globulin, high-density lipoprotein cholesterol, smoking initiation, and alcohol intake frequency and total body bone mineral density at ages 45-60 years. Exposure genome-wide association study summary statistics were derived from large European-ancestry populations and were not restricted to midlife women, whereas the outcome genome-wide association study captured an age-stratified total body bone mineral density phenotype at age 45-60 years. This age range overlaps with the menopausal transition and early postmenopause in women. Univariable, reverse, and multivariable Mendelian randomization analyses were performed, with inverse-variance weighting as the primary method and complementary sensitivity analyses used to assess heterogeneity, pleiotropy, and result stability. Genetically predicted higher sex hormone-binding globulin was associated with lower total body bone mineral density (β = -0.111, 95% CI: -0.170 to -0.051; P = 0.0003). Reverse Mendelian randomization did not support reverse causation from bone mineral density to sex hormone-binding globulin. Multivariable analyses suggested that this association persisted after adjustment for selected metabolic biomarkers. The other examined exposures did not show consistent evidence of association. These findings provide genetic evidence linking sex hormone-binding globulin to total-body bone mineral density at ages 45-60 years. Further prospective and predictive studies are needed to evaluate its clinical relevance beyond established bone health assessment tools.

Humans

FINEMAP-miss: fine-mapping genome-wide association studies with missing genotype information.

MOTIVATION: The most informative genome-wide association studies (GWAS) are meta-analyses that have combined multiple studies to increase the GWAS sample size. Statistical fine-mapping is a key downstream analysis of GWAS to jointly evaluate the probability of causality of all variants in a genomic region of interest. Current fine-mapping methods are miscalibrated in the meta-analysis setting due to variation in sample size across the variants. RESULTS: We introduce FINEMAP-miss, a new fine-mapping method that extends the FINEMAP model to account for variant-specific missingness. We show that FINEMAP-miss is well-calibrated in meta-analysis simulations where the standard fine-mapping fails. Compared to the summary statistics imputation approach, FINEMAP-miss provides clear improvement when the causal variants have low imputation information or when the sample size or complexity of the meta-analysis setting increase. We successfully apply FINEMAP-miss on a breast cancer GWAS meta-analysis where neither the standard fine-mapping nor the summary statistics imputation are applicable. AVAILABILITY: An open source implementation of FINEMAP-miss as an R package ("finemapmiss") is available at https://github.com/JoonasKartau/finemapmiss. The archived version of FINEMAP-miss used for this publication can be found on Zenodo at https://doi.org/10.5281/zenodo.17492622. SUPPLEMENTARY INFORMATION: is available at the journal's web site.

Genome-Wide Association Study