PubMed HealthSearch

Biomedical subjects

Robert Kaplan

Publications and source records attributed to Robert Kaplan.

7 recordsLinked to original sources

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

Steroid hormone biosynthesis and dietary related metabolites associated with excessive daytime sleepiness.

BACKGROUND: Excessive daytime sleepiness (EDS) is a complex sleep problem that affects approximately 33% of the United States population. Although EDS usually occurs in conjunction with insufficient sleep and other sleep and circadian disorders, recent studies have shown unique genetic markers and metabolic pathways underlying EDS. Here, we aimed to further elucidate the biological profile of EDS using large-scale single- and pathway-level metabolomics analyses. METHODS: Metabolomics data were available for 877 metabolites in 6071 individuals from the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). EDS was assessed using the Epworth Sleepiness Scale (ESS) questionnaire. We performed linear regression for each metabolite on the continuous ESS score, adjusting for demographic, lifestyle, and physiological confounders, and in sex specific groups. Subsequently, gaussian graphical modelling was performed coupled with pathway and enrichment analyses to generate a holistic interactive network of the metabolomic profile of EDS associations. FINDINGS: We identified seven metabolites belonging to steroids, sphingomyelin, and long-chain fatty acids sub-pathways in the primary model associated with EDS, and an additional three metabolites in the male-specific analysis. INTERPRETATION: Our findings indicate that an EDS metabolomic profile is characterised by endogenous and dietary metabolites within the steroid hormone biosynthesis pathway, with some pathways that differ by sex. These pathways may be useful for understanding the causes or consequences of EDS and related sleep disorders. FUNDING: Details regarding funding supporting this work and all studies involved are provided in the acknowledgements section.

Humans

Polygenic scores for obstructive sleep apnoea reveal pathways contributing to cardiovascular disease.

BACKGROUND: Obstructive sleep apnoea (OSA) is a common chronic condition, with obesity its strongest risk factor. Polygenic scores (PGSs) summarise the genetic liability to phenotype and can provide insights into relationships between phenotypes. Recently, large datasets that include genetic data and OSA status became available, providing an opportunity to utilise PGS approaches to study the genetic relationship between OSA and other phenotypes, while differentiating OSA-specific from obesity-specific genetic factors. METHODS: Using race/ethnic diverse samples from over 1.2 million individuals from the Million Veteran Program, FinnGen, TOPMed, All of Us (AoU), Geisinger's MyCode, MGB Biobank, and the Human Phenotype Project, we developed and assessed PGSs for OSA, both without (BMIunadjOSA-PGS) and with adjustment for the genetic contributions of BMI (BMIadjOSA-PGS). FINDINGS: Adjusted odds ratios (ORs) for OSA per 1 standard deviation of the PGSs ranged from 1.38 to 2.75. The associations of BMIadjOSA- and BMIunadjOSA-PGSs with CVD outcomes in AoU shared both common and distinct patterns. Only BMIunadjOSA-PGS was associated with type 2 diabetes, heart failure, and coronary artery disease, while both BMIadjOSA- and BMIunadjOSA-PGSs were associated with hypertension and stroke. Sex stratified analyses revealed that BMIadjOSA-PGS association with hypertension was driven by females (OR = 1.1, p-value = 0.002, OR = 1.01 p-value = 0.2 in males). OSA PGSs were also associated with body fat measures with some sex-specific associations. INTERPRETATION: Distinct components of OSA genetic risk are related and independent of obesity. Sex-specific associations with body fat distribution measures may explain differing OSA risks and associations with cardiometabolic morbidities between sexes. FUNDING: R01AG080598.

Humans

Heterogeneity of Apolipoprotein B Levels Among Hispanic or Latino Individuals Residing in the US.

IMPORTANCE: Apolipoprotein B (apoB) distribution and its implications as an atherosclerotic cardiovascular disease (ASCVD) risk-enhancing factor among individuals of diverse Hispanic or Latino backgrounds have not been described. OBJECTIVE: To describe the distribution of apoB in the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) cohort and to characterize associations of baseline sociodemographic and clinical variables with apoB and self-identified Hispanic or Latino background. DESIGN, SETTING, AND PARTICIPANTS: The HCHS/SOL was a prospective, population-based cohort study of diverse Hispanic or Latino adults living in the US who were recruited and screened between March 2008 and June 2011. Sampling weights were used to generate a population-based sample of Hispanic or Latino participants aged 18 to 74 years who resided in 4 US metropolitan areas (Bronx, New York; Chicago, Illinois; Miami, Florida; and San Diego, California). ApoB concentration was measured in participants from the HCHS/SOL, and apoB tertiles were compared across demographic groups, including self-identified Hispanic or Latino background. Median percentage continental genetic ancestry (West African, Amerindian, and European) was compared across apoB tertiles. EXPOSURE: ApoB measured in mg/dL from serum or plasma using an immunoturbidimetric assay. MAIN OUTCOMES AND MEASURES: ApoB tertiles were determined, and traditional lipids were evaluated across apoB tertiles. ApoB and traditional lipid measurements were assessed across ASCVD risk categories. Additionally, scatterplots were created to observe correlations between apoB and low-density lipoprotein cholesterol or non-high-density lipoprotein cholesterol. RESULTS: Overall mean (SD) apoB concentration was 99.8 (0.4) mg/dL, with male participants displaying significantly higher mean levels than female participants (102.4 vs 97.4 mg/dL, respectively). Mean (SD) participant age was 41.1 (0.8) years, and 8376 participants (51.9%) were female. ApoB levels were higher among older age groups. There was significant heterogeneity in mean apoB concentrations across self-identified Hispanic or Latino background groups, ranging from 95.1 mg/dL in Dominican individuals to 104.8 mg/dL in Cuban individuals. The prevalence of elevated apoB (&#x2265;130 mg/dL) was greater across higher predicted ASCVD risk categories. Among participants with a 10-year predicted ASCVD risk of 7.5% or higher, 26.5% had an elevated apoB. Median West African ancestry was lower across higher tertiles of apoB. CONCLUSIONS AND RELEVANCE: In this cohort study among participants from the HCHS/SOL, elevated apoB was present in one-quarter of a diverse cohort study of Hispanic or Latino individuals who were at intermediate or high predicted ASCVD risk. Differences in apoB distribution among Hispanic or Latino individuals may have important implications for apoB's use in ASCVD risk assessment.

Adolescent

The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup.

Polygenic risk scores (PRSs) depend on genetic ancestry due to differences in allele frequencies between ancestral populations. This leads to implementation challenges in diverse populations. We propose a framework to calibrate PRS based on ancestral makeup. We define a metric called "expected PRS" (ePRS), the expected value of a PRS based on one's global or local admixture patterns. We further define the "residual PRS" (rPRS), measuring the deviation of the PRS from the ePRS. Simulation studies confirm that it suffices to adjust for ePRS to obtain nearly unbiased estimates of the PRS-outcome association without further adjusting for PCs. Using the TOPMed dataset, the estimated effect size of the rPRS adjusting for the ePRS is similar to the estimated effect of the PRS adjusting for genetic PCs. Similarly, we applied the ePRS framework to six cardiovascular-related traits in the All of Us dataset, and the results are consistent with those from the TOPMed analysis. The ePRS framework can protect from population stratification in association analysis and provide an equitable strategy to quantify genetic risk across diverse populations.

Journal Article

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans