PubMed HealthSearch

Biomedical subjects

Karmel W Choi

Publications and source records attributed to Karmel W Choi.

4 recordsLinked to original sources

Systematic common and rare variant association testing in 392,030 whole genomes in All of Us.

Large-scale genome-wide association studies (GWAS) and rare variant association studies (RVAS) from population biobanks provide valuable resources for gene discovery in complex human traits. We present an analysis of the All of Us Research Program v8 release, which includes whole genome sequencing data and harmonized phenotypic information of 392,030 participants after quality control, enabling a unified investigation of rare and common variants across a spectrum of human traits and diseases. We build an extensive phenome- and genome-wide ("All by All") computational framework to perform GWAS and RVAS on 3,602 phenotypes and identify 49,863 approximately independent, high-quality single-variant and gene-level associations. Meta-analyses of All of Us and UK Biobank, with sample sizes as large as 786,871 participants, further enhance statistical power and find 193 pLoF gene-phenotype associations that are not significant in either cohort alone, including 22 associations not highlighted by previous studies. We also present a public interactive browser that integrates association results for common and rare variants to facilitate interpretation and rapid querying of summary statistics, along with supporting documentation, and a Featured Workspace in the All of Us Researcher Workbench. Our framework will apply to iterative data releases as All of Us grows, empowering researchers worldwide to uncover insights into the functional effects of genetic components on complex traits and diseases.

Journal Article

Can Psychiatric Genetics Advance Without Incorporating a Life Course Perspective?

Psychiatric disorders unfold over the life course; however, genomic studies of these conditions overwhelmingly rely on phenotypes collected at a single time point, often in adulthood. Therefore, genome-wide association studies (GWASs) of psychiatric conditions may miss genetic variants with time-varying relevance to etiology, prevention, and treatment, such as those that influence trajectories of symptoms and behaviors, age at onset, course of treatment response, and the co-evolution of comorbidities. With recent advances in longitudinal biobanks and analytic tools, we posit that incorporating a life course perspective in psychiatric genetics will enable critically relevant insights into each of these areas of investigation. We propose that the current inconsistent portability of polygenic scores across age groups can be reconciled through the design of carefully considered longitudinal GWASs in age-diverse samples. Pioneering longitudinal GWASs in psychiatry have revealed novel genomic signals associated with time-dependent phenotypes that are distinct from those influencing lifetime diagnosis, suggesting that the study of longitudinal phenotypes will complement cross-sectional approaches and empower biological and therapeutic discoveries. Advances in post-GWAS functional annotation resources and analytic approaches now enable us to contextualize the genetic contributions to psychiatric disorders as dynamic age- and exposure-dependent processes. Although longitudinal GWASs pose unique challenges with regard to data availability, selection bias, and missing data, integrating temporality into psychiatric genetics at scale is now attainable and promises to reveal novel biology and therapeutic opportunities for psychiatric conditions.

Cohort study

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results across studies. Here, we performed multiple modeling experiments integrating clinical and demographic data from electronic health records (EHR) and genetic data to understand which decision points may affect performance. Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from two large independent health systems and polygenic risk scores (PRS) were generated across all patients with genetic data in the corresponding biobanks. Crohn's disease was used as the model phenotype based on its substantial genetic component, established EHR-based definition, and sufficient prevalence for model training and testing. We investigated the impact of PRS integration method, as well as choices regarding training sample, model complexity, and performance metrics. Overall, our results show that including PRS resulted in higher performance by some metrics but the gain in performance was only robust when combined with demographic data alone. Improvements were inconsistent or negligible after including additional clinical information. The impact of genetic information on performance also varied by PRS integration method, with a small improvement in some cases from combining PRS with the output of a clinical model (late-fusion) compared to its inclusion an additional feature (early-fusion). The effects of other modeling decisions varied between institutions though performance increased with more compute-intensive models such as random forest. This work highlights the importance of considering methodological decision points in interpreting the impact on prediction performance when including PRS information in clinical models.

Preprint