PubMed HealthSearch

Biomedical subjects

Anurag Verma

Publications and source records attributed to Anurag Verma.

6 recordsLinked to original sources

Multi-population GWAS meta-analysis identifies bladder cancer susceptibility loci and highlights genetic regulation of smoking-related risk.

Bladder cancer is the ninth most common cancer worldwide, caused by genetic and environmental risk factors. Here, we report the findings of a multi-population meta-analysis of genome-wide association studies, including 32,470 individuals with and 1,753,462 without bladder cancer. We identify 70 independent risk loci, of which 43 are novel. Using a 70-marker polygenic risk score (HR = 1.63 per standard deviation), we increase the area under the curve from 0.71 (baseline risk model) to 0.75. Integrative analyses reveal the enrichment of the associated variants within accessible chromatin regions, and of the prioritized genes within pathways for xenobiotic metabolism and smoking behavior. Specifically, we show that the 15q25.1 variant rs71581744-ACCCC/A co-localizes with tissue-specific CHRNA3 expression, modulates mRNA stability, and associates with risk of muscle-invasive bladder cancer among current smokers. Together, these findings substantially expand the known genetic architecture of bladder cancer risk and highlight the germline regulation of smoking behavior as a mechanism driving bladder cancer susceptibility.

Humans

PMBB Geno-Pheno Toolkit: A suite of scalable, reproducible pipelines for cross-biobank association analyses.

Electronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkit's SAIGE family of pipelines - supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing - together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted.

Journal Article

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs.

MOTIVATION: Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. RESULTS: We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Genome-Wide Association Study

Germline Variants Influence Chronic Liver Disease Progression through Distinct Pathways.

Cirrhosis and hepatocellular carcinoma (HCC) are long-term complications of chronic liver disease (CLD). In this large multi-ancestry genome-wide association study of all-cause cirrhosis (35,481 cases, 2.36M controls) and HCC (6,680 cases, 1.76M controls), we identified 27 loci associated with cirrhosis (10 novel) and 11 with HCC (three novel). Three novel cirrhosis loci were replicated in independent cohorts (e.g. FGF21, RPTOR, and IFNL3/4). Fifteen cirrhosis loci exhibited differential effects on cirrhosis risk via underlying etiologies, and six HCC loci influenced HCC risk indirectly via cirrhosis. In a gene-burden analysis of rare variants from whole-genome sequencing data in the VA Million Veteran Program (n=102,677), we identified GSTA5 as a novel cirrhosis-associated gene, while APOB and ATP9B were associated with and replicated for HCC. A high genetic risk score for cirrhosis was associated with a nearly doubled risk of CLD progressing to cirrhosis (HR=1.94, P=2×10-68) and of cirrhosis progressing to HCC (HR=1.65, P=7×10-08). Finally, among individuals with chronic hepatitis C who underwent antiviral therapy, cirrhosis risk was modified by variants in PNPLA3, IFNL3/4, and CD81 following pegylated interferon-α therapy, and by APOE lead variant following direct-acting antiviral therapy. These findings provide new insights into the complex genetic architecture of CLD progression with potential clinical and therapeutic implications.

Journal Article

Bidirectional Risk Modulator and Modifier Variant of Dilated and Hypertrophic Cardiomyopathy in BAG3.

IMPORTANCE: The genetic factors that modulate the reduced penetrance and variable expressivity of heritable dilated cardiomyopathy (DCM) are largely unknown. BAG3 genetic variants have been implicated in both DCM and hypertrophic cardiomyopathy (HCM), nominating BAG3 as a gene that harbors potential modifier variants in DCM. OBJECTIVE: To interrogate the clinical traits and diseases associated with BAG3 coding variation. DESIGN, SETTING, AND PARTICIPANTS: This was a cross-sectional study in the Penn Medicine BioBank (PMBB) enrolling patients of the University of Pennsylvania Health System's clinical practice sites from 2014 to 2023. Whole-exome sequencing (WES) was linked to electronic health record (EHR) data to associate BAG3 coding variants with EHR phenotypes. This was a health care population-based study including individuals of European and African genetic ancestry in the PMBB with WES linked to EHR phenotypes, with replication studies in BioVU, UK Biobank, MyCode, and DCM Precision Medicine Study. EXPOSURES: Carrier status for BAG3 coding variants. MAIN OUTCOMES AND MEASURES: Association of BAG3 coding variation with clinical diagnoses, echocardiographic traits, and longitudinal outcomes. RESULTS: In PMBB (n = 43 731; median [IQR] age, 65 [50-76] years; 21 907 female [50.1%]), among 30 324 European and 11 198 African individuals, the common C151R variant was associated with decreased risk for DCM (odds ratio [OR], 0.85; 95% CI, 0.78-0.92) and simultaneous increased risk for HCM (OR, 1.59; 95% CI, 1.25-2.02), which was confirmed in the replication cohorts. C151R carriers exhibited improved longitudinal outcomes compared with noncarriers as assessed by age at death (hazard ratio [HR], 0.85; 95% CI, 0.74-0.96; median [IQR] age, 71.8 [63.1-80.7] in carriers and 70.3 [61.6-79.2] in noncarriers) and heart transplant (HR, 0.81; 95% CI, 0.66-0.99; median [IQR] age, 56.7 [46.1-63.1] in carriers and 55.6 [45.2-62.9] in noncarriers). C151R was associated with reduced risk of DCM (OR, 0.42; 95% CI, 0.24-0.74) and heart failure (OR, 0.27; 95% CI, 0.14-0.50) among individuals harboring truncating TTN variants in exons with high cardiac expression (n = 358). CONCLUSIONS AND RELEVANCE: BAG3 C151R was identified as a bidirectional modulator of risk along the DCM-HCM spectrum, as well as an important genetic modifier variant in TTN-mediated DCM. This work expands on the understanding of the etiology and penetrance of DCM, suggesting that BAG3 C151R is an important genetic modifier variant contributing to the variable expressivity of DCM, warranting further exploration of its mechanisms and of genetic modifiers in DCM more broadly.

Humans