PubMed HealthSearch

Biomedical subjects

Lifang Hou

Publications and source records attributed to Lifang Hou.

10 recordsLinked to original sources

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

PathwayVote: an R package for robust pathway enrichment analysis for DNA methylation data using a consensus-based voting framework.

MOTIVATION: Pathway enrichment analysis is commonly used to interpret epigenomewide association studies, yet conventional methods often rely on arbitrary thresholds and simplified CpG-gene mappings, making them sensitive to analytical choices and unable to fully leverage CpG-gene relationships Recent advances in expression quantitative trait methylation (eQTM) studies offer a rich resource to refine these mappings, but are rarely utilized in DNA methylation enrichment pipelines. RESULTS: We developed PathwayVote, an R package that implements a voting-based consensus approach and leverages eQTM data to identify robustly enriched pathways. PathwayVote reduces dependence on arbitrary cutoffs and improves sensitivity and reproducibility of enrichment results. AVAILABILITY AND IMPLEMENTATION: PathwayVote is freely available on GitHub (https://github.com/YinanZheng/PathwayVote) under the GPL-3 license and CRAN: https://CRAN.R-project.org/package=PathwayVote. The version of the code corresponding to this manuscript has been archived on Zenodo (https://doi.org/10.5281/zenodo.17209507).

Humans

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246&#xa0;K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246&#xa0;K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86&#xa0;K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

Polygenic scores for obstructive sleep apnoea reveal pathways contributing to cardiovascular disease.

BACKGROUND: Obstructive sleep apnoea (OSA) is a common chronic condition, with obesity its strongest risk factor. Polygenic scores (PGSs) summarise the genetic liability to phenotype and can provide insights into relationships between phenotypes. Recently, large datasets that include genetic data and OSA status became available, providing an opportunity to utilise PGS approaches to study the genetic relationship between OSA and other phenotypes, while differentiating OSA-specific from obesity-specific genetic factors. METHODS: Using race/ethnic diverse samples from over 1.2 million individuals from the Million Veteran Program, FinnGen, TOPMed, All of Us (AoU), Geisinger's MyCode, MGB Biobank, and the Human Phenotype Project, we developed and assessed PGSs for OSA, both without (BMIunadjOSA-PGS) and with adjustment for the genetic contributions of BMI (BMIadjOSA-PGS). FINDINGS: Adjusted odds ratios (ORs) for OSA per 1 standard deviation of the PGSs ranged from 1.38 to 2.75. The associations of BMIadjOSA- and BMIunadjOSA-PGSs with CVD outcomes in AoU shared both common and distinct patterns. Only BMIunadjOSA-PGS was associated with type 2 diabetes, heart failure, and coronary artery disease, while both BMIadjOSA- and BMIunadjOSA-PGSs were associated with hypertension and stroke. Sex stratified analyses revealed that BMIadjOSA-PGS association with hypertension was driven by females (OR = 1.1, p-value = 0.002, OR = 1.01 p-value = 0.2 in males). OSA PGSs were also associated with body fat measures with some sex-specific associations. INTERPRETATION: Distinct components of OSA genetic risk are related and independent of obesity. Sex-specific associations with body fat distribution measures may explain differing OSA risks and associations with cardiometabolic morbidities between sexes. FUNDING: R01AG080598.

Humans

Alterations in DNA Methylation, Proteomic, and Metabolomic Profiles in African Ancestry Populations with APOL1 Risk Alleles.

KEY POINTS: We aimed to elucidate potential methylation, proteomic, and metabolomic mechanisms by which APOL1 variants may be linked to kidney disease. We report distinct methylation profiling between APOL1 risk allele carriers and noncarriers, many near APOL gene family. We report higher APOL1 protein and lower C18:1 cholesteryl ester in two risk allele carriers. BACKGROUND: The APOL1 high-risk haplotype has been associated with CKD and the deterioration of kidney function, particularly in populations with West African ancestry. However, the mechanisms by which APOL1 risk variants increase the risk for kidney disease and its progression have not been fully elucidated. METHODS: We compared methylation (N=3191; 715 [22%] carriers), proteomic (N=1240; 169 [14%] carriers), and metabolomic (N=6309; 674 [11%] carriers) profiles in African and Hispanic/Latino carriers of two APOL1 high-risk alleles (G1/G1, G2/G2, G1/G2) and noncarriers (G0/G0), excluding heterozygotes (G0/G1, G0/G2), from the Population Architecture using Genomics and Epidemiology Consortium and UK Biobank. In each study, the associations between the APOL1 high-risk haplotype and up to 722,719 cytosine-phosphate-guanine (CpG) sites, 2923 proteins, or 836 metabolites were estimated using covariate-adjusted linear regression models, followed by fixed-effects sample size&#x2013;weighted meta-analyses. RESULTS: Significant associations were observed between APOL1 high-risk haplotype and methylation at 52 CpG sites, with 48 located on chromosome 22 and 18 in the vicinity of APOL1&#x2013;4 and MYH9. All significant CpG sites near APOL2 were hypomethylated, whereas those near APOL3 and APOL4 were hypermethylated. APOL1-associated CpG sites were also identified in genes involved in ion transport and mitochondrial stress pathways. Sensitivity analyses indicated consistent yet attenuated effects among heterozygotes, supporting an additive effect of APOL1 risk alleles. Further analyses of the 52 CpG sites identified two near APOL4 exhibiting G1-specific effects, eight associated with CKD but none with eGFR, and three showing heterogeneity by CKD status. In addition, carrying two APOL1 risk alleles was associated with higher plasma APOL1 protein (&#x3b2;=1.12, PFDR = 2.26e-70) and lower C18:1 cholesteryl ester metabolite (Z=&#x2212;4.50, PFDR = 4.83e-3). CONCLUSIONS: Our results demonstrate differential methylation, proteomic, and metabolomic profiles associated with APOL1 high-risk haplotypes.

APOL1

Whole genome sequence-based association analysis of African American individuals with bipolar disorder and schizophrenia.

In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole genome sequencing of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls. To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of BD association with single-variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD GWAS loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.

Journal Article

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58&#x2009;706 SVs in a study sample of 11&#x2009;556 CAD cases and 42&#x2009;907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Cardiovascular Risk Factors and Genetic Risk in Transthyretin V142I Carriers.

BACKGROUND: Nearly 3% to 4% of Black individuals in the United States carry the transthyretin V142I variant, which increases their risk of heart failure. However, the role of cardiovascular (CV) risk factors (RFs) in influencing the risk of clinical outcomes among V142I variant carriers is unknown. OBJECTIVES: This study aimed to assess the impact of CV RFs on the risk of heart failure in V142I carriers. METHODS: This study included self-identified Black individuals without prevalent heart failure from 6 TOPMed (Trans-Omics for Precision Medicine) cohorts, the REGARDS (Reasons for Geographic And Racial Differences in Stroke) study, and the All of Us Research Program. The cohort was stratified based on the V142I genotype and the number of CV RFs (hypertension, diabetes, obesity, and hypercholesterolemia). Adjusted Cox models were used to assess the association of heart failure with the V142I genotype and CV RF profile, taking noncarriers with a favorable CV RF profile as reference. RESULTS: The cross-sectional analysis, including 1,625 V142I carriers among 48,365 Black individuals, found that the prevalence of CV RFs did not vary by V142I carrier status. In the longitudinal analysis, there were 587 (3.2%) V142I carriers among 18,407 Black individuals (median age: 60 years [Q1-Q3: 52-68 years], 63.0% female). Among carriers, the heart failure risk was attenuated with a favorable (0 or 1 RF) CV RF profile (adjusted HR: 2.26; 95%&#xa0;CI: 1.58-3.23) compared with an unfavorable (3 or 4 RFs) CV RF profile (adjusted HR: 4.14; 95%&#xa0;CI: 2.79-6.14). CONCLUSIONS: A favorable CV RF profile lowers but does not abrogate V142I variant-associated heart failure risk. This study highlights the importance of having a favorable CV RF profile among V142I carriers for risk reduction of heart failure.

Aged

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency &#x2264;&#x2009;0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans