PubMed HealthSearch

Biomedical subjects

Jordan W Smoller

Publications and source records attributed to Jordan W Smoller.

11 recordsLinked to original sources

Distinguishing different psychiatric disorders using DDx-PRS.

Despite great progress on case-control polygenic prediction, an unmet need remains for a method that genetically distinguishes clinically related disorders (e.g., schizophrenia (SCZ) versus bipolar disorder (BIP) versus major depressive disorder (MDD) versus controls). We introduce differential diagnosis-polygenic risk score (DDx-PRS), which jointly estimates the posterior probabilities of each diagnostic category (e.g., SCZ = 50%, BIP = 25%, MDD = 15%, control = 10%) by modeling variance-covariance structure across disorders, leveraging case-control polygenic risk scores and prior clinical probabilities for each diagnostic category. We applied DDx-PRS to Psychiatric Genomics Consortium SCZ, BIP, MDD and control data, including summary-level training data from three case-control genome-wide association studies (n = 41,917-173,140 cases; total n = 1,048,683) and held-out test data from different cohorts with equal numbers for each diagnostic category (total n = 11,460). DDx-PRS was well calibrated and well powered (consistent with simulations) and produced comparable results to methods that require tuning data. True diagnosis probabilities in the top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, implying appreciable potential for clinical utility in certain settings.

Humans

Unsupervised characterization of 100,272 EHR patients identifies high-risk groups and comorbidities linked to premature aging.

Electronic health records (EHRs) contain extensive multidimensional patient data, presenting challenges for the discovery of novel and meaningful clinical patterns. Unsupervised clustering of high-dimensional clinical data holds great potential for identifying novel clinical patterns. Here, we performed unsupervised clustering and characterized 100,272 patients in the Electronic Medical Records and GEnomics (eMERGE) Network. We identified 70 clusters defined by distinct comorbidity patterns. Meanwhile, age and sex are also strongly associated with patient stratification, influencing phenotype prevalence and onset time. Notably, phenotype onset time accurately predicted chronological age and was significantly associated with overall mortality risk. Besides age and sex, we assessed the contribution of genetic variation to phenotype development and observed evidence of cross-phenotype associations influencing cluster membership and comorbidity patterns. However, the role of genetics recedes during aging. We also identified several high-risk clusters with elevated Charlson Comorbidity Index (CCI) scores and validated these findings in an independent cohort. Further analysis of these clusters revealed phenotypes linked to premature aging and highlighted a survival selection among older participants in observational studies. Overall, this study enables phenome-wide unsupervised patient stratification for multimorbidity discovery in largely unannotated clinical data, offering valuable insights into patient stratification, comorbidity analysis, aging, and health outcomes.

Journal Article

Association of Genetic Liability to Psychiatric Disorders with Peripheral Metabolic Dysregulation.

IMPORTANCE: Individuals with psychiatric disorders face elevated cardiometabolic risk which is linked to increased mortality. The extent to which this reflects shared pathogenesis or the downstream effects of illness and treatment remains poorly understood. OBJECTIVE: To characterize the direct pleiotropic effects of psychiatric genetic liability on circulating metabolites and aggregate cardiometabolic risk, independent of psychiatric diagnosis and psychotropic medication use. DESIGN SETTING AND PARTICIPANTS: Cross-sectional analysis of Mass General Brigham Biobank participants with metabolomic profiling, genomic data, and linked electronic health records. EXPOSURES: Genetic liability to nine psychiatric disorders quantified using polygenic risk scores (PRS): attention deficit/hyperactivity disorder (ADHD), anorexia nervosa (ANO), anxiety disorder (ANX), autism spectrum disorder (ASD), bipolar disorder (BD), major depressive disorder (MDD), PTSD, schizophrenia (SCZ), and substance use disorder (SUD). MAIN OUTCOMES AND MEASURES: 249 circulating metabolites and four metabolomic risk scores (MRS) for type 2 diabetes, myocardial infarction, ischemic stroke, and vascular dementia. PRS-metabolite associations were estimated using nested models adjusting for lifetime psychiatric diagnosis and psychotropic medication use. RESULTS: Across 25,290 participants, we identified 604 significant PRS-metabolite associations (Bonferroni p< 1.36 x 10-4), of which 89% persisted after adjustment for lifetime diagnosis and medication use, suggesting that the direct genetic effects on metabolism are largely independent of illness or treatment. PRS for MDD, PTSD, and ADHD showed the most extensive dysregulation, with a transdiagnostic pattern of elevated lipids and systemic inflammation, specifically triglycerides (&#x3b2; = 0.04 to 0.05, all p< 4.4 x10-13) and glycoprotein acetyls (&#x3b2; = 0.05, all p< 2.2 x10-16). Notably, PRS for SCZ and BD showed minimal metabolite dysregulation despite having the strongest association with their target diagnoses. PRS for MDD, PTSD, ADHD, and SUD were associated with increased MRS across cardiometabolic conditions (&#x3b2; = 0.03 to 0.08, all p< 2.1 x10-4). Sensitivity analyses controlling for BMI or excluding participants without any psychiatric history (N: 21,305 and 11,150, respectively) showed a similar pattern. CONCLUSIONS AND RELEVANCE: Psychiatric genetic liability is associated with systemic metabolic dysregulation independent of illness onset or treatment, supporting a partially pleiotropic basis for psychiatric-cardiometabolic comorbidity.

Journal Article

Systematic common and rare variant association testing in 392,030 whole genomes in All of Us.

Large-scale genome-wide association studies (GWAS) and rare variant association studies (RVAS) from population biobanks provide valuable resources for gene discovery in complex human traits. We present an analysis of the All of Us Research Program v8 release, which includes whole genome sequencing data and harmonized phenotypic information of 392,030 participants after quality control, enabling a unified investigation of rare and common variants across a spectrum of human traits and diseases. We build an extensive phenome- and genome-wide ("All by All") computational framework to perform GWAS and RVAS on 3,602 phenotypes and identify 49,863 approximately independent, high-quality single-variant and gene-level associations. Meta-analyses of All of Us and UK Biobank, with sample sizes as large as 786,871 participants, further enhance statistical power and find 193 pLoF gene-phenotype associations that are not significant in either cohort alone, including 22 associations not highlighted by previous studies. We also present a public interactive browser that integrates association results for common and rare variants to facilitate interpretation and rapid querying of summary statistics, along with supporting documentation, and a Featured Workspace in the All of Us Researcher Workbench. Our framework will apply to iterative data releases as All of Us grows, empowering researchers worldwide to uncover insights into the functional effects of genetic components on complex traits and diseases.

Journal Article

Optimizing Control Definitions in Opioid Use Disorder Genetic Research Using Electronic Health Records.

Amidst the opioid crisis, understanding the genetic basis of opioid use disorder (OUD) is crucial for identifying biological mechanisms and intervention points. However, genome-wide association studies (GWASs) have been hampered by inadequate sample sizes and often the use of control populations not assessed for prior opioid exposure. Because opioid exposure is a prerequisite for the development of OUD, consideration of exposure history in controls is important. Electronic health record data (EHR) paired with genomic information allow a broader sampling of patients with OUD and exposed controls. We leveraged data across two healthcare systems to evaluate the impact of using controls not screened for opioid exposure ('generic') versus minimally opioid-exposed control ('exposed'). First, at the phenotypic level, we conducted phenome-wide association studies (PheWAS) to compare the medical comorbidity profiles of OUD cases when using generic versus exposed controls. While PheWAS results for OUD-related comorbidities were more pronounced when using the generic group, 83% of the disease associations were overlapping and of similar effect sizes. Second, at the genetic level, we conducted GWAS (cases vs. generic; cases vs. exposed) and assessed differences in genetic correlations and degrees of phenotypic misclassification. Genetic results were concordant across control groups based on heritability (generic: 0.16&#x2009;&#xb1;&#x2009;0.07 vs. 0.10&#x2009;&#xb1;&#x2009;0.07), associations with the coding OPRM1 variant rs1799971 (pgeneric&#x2009;=&#x2009;8.83E-03 vs. pexposed&#x2009;=&#x2009;1.83E-02) and genetic correlations with prior OUD GWAS (rg-generic&#x2009;=&#x2009;0.83&#x2009;&#xb1;&#x2009;0.26 vs. rg-exposed&#x2009;=&#x2009;0.78&#x2009;&#xb1;&#x2009;0.27). Although GWASs were limited by sample size (Ngeneric&#x2009;=&#x2009;6269, Nexposed&#x2009;=&#x2009;6365), compared to an independent OUD GWAS (N&#x2009;=&#x2009;425&#x2009;944), the dilution value for the two GWAS was not different from 1, suggesting no major impact of phenotypic misclassification. This study represents the first effort to enhance OUD genetic research through optimization of control definitions using EHR data. Generic controls ascertained within the US health systems, where exposure to prescription opioids is high, offer a practical alternative for genetic studies of OUD.

Humans

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Multi-ancestry meta-analysis of tobacco use disorder identifies 461 potential risk genes and reveals associations with multiple health outcomes.

Tobacco use disorder (TUD) is the most prevalent substance use disorder in the world. Genetic factors influence smoking behaviours and although strides have been made using genome-wide association studies to identify risk variants, most variants identified have been for nicotine consumption, rather than TUD. Here we leveraged four US biobanks to perform a multi-ancestral meta-analysis of TUD (derived via electronic health records) in 653,790 individuals (495,005 European, 114,420 African American and 44,365 Latin American) and data from UK Biobank (ncombined&#x2009;=&#x2009;898,680). We identified 88 independent risk loci; integration with functional genomic tools uncovered 461 potential risk genes, primarily expressed in the brain. TUD was genetically correlated with smoking and psychiatric traits from traditionally ascertained cohorts, externalizing behaviours in children and hundreds of medical outcomes, including HIV infection, heart disease and pain. This work furthers our biological understanding of TUD and establishes electronic health records as a source of phenotypic information for studying the genetics of TUD.

Humans

Distinguishing different psychiatric disorders using DDx-PRS.

Despite great progress on methods for case-control polygenic prediction (e.g. schizophrenia vs. control), there remains an unmet need for a method that genetically distinguishes clinically related disorders (e.g. schizophrenia (SCZ) vs. bipolar disorder (BIP) vs. depression (MDD) vs. control); such a method could have important clinical value, especially at disorder onset when differential diagnosis can be challenging. Here, we introduce a method, Differential Diagnosis-Polygenic Risk Score (DDx-PRS), that jointly estimates posterior probabilities of each possible diagnostic category (e.g. SCZ=50%, BIP=25%, MDD=15%, control=10%) by modeling variance/covariance structure across disorders, leveraging case-control polygenic risk scores (PRS) for each disorder (computed using existing methods) and prior clinical probabilities for each diagnostic category. DDx-PRS uses only summary-level training data and does not use tuning data, facilitating implementation in clinical settings. In simulations, DDx-PRS was well-calibrated (whereas a simpler approach that analyzes each disorder marginally was poorly calibrated), and effective in distinguishing each diagnostic category vs. the rest. We then applied DDx-PRS to Psychiatric Genomics Consortium SCZ/BIP/MDD/control data, including summary-level training data from 3 case-control GWAS ( N =41,917-173,140 cases; total N =1,048,683) and held-out test data from different cohorts with equal numbers of each diagnostic category (total N =11,460). DDx-PRS was well-calibrated and well-powered relative to these training sample sizes, attaining AUCs of 0.66 for SCZ vs. rest, 0.64 for BIP vs. rest, 0.59 for MDD vs. rest, and 0.68 for control vs. rest. DDx-PRS produced comparable results to methods that leverage tuning data, confirming that DDx-PRS is an effective method. True diagnosis probabilities in top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, particularly in projections to larger training sample sizes, implying considerable potential for clinical utility under certain circumstances. In conclusion, DDx-PRS is an effective method for distinguishing clinically related disorders.

Journal Article

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results across studies. Here, we performed multiple modeling experiments integrating clinical and demographic data from electronic health records (EHR) and genetic data to understand which decision points may affect performance. Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from two large independent health systems and polygenic risk scores (PRS) were generated across all patients with genetic data in the corresponding biobanks. Crohn's disease was used as the model phenotype based on its substantial genetic component, established EHR-based definition, and sufficient prevalence for model training and testing. We investigated the impact of PRS integration method, as well as choices regarding training sample, model complexity, and performance metrics. Overall, our results show that including PRS resulted in higher performance by some metrics but the gain in performance was only robust when combined with demographic data alone. Improvements were inconsistent or negligible after including additional clinical information. The impact of genetic information on performance also varied by PRS integration method, with a small improvement in some cases from combining PRS with the output of a clinical model (late-fusion) compared to its inclusion an additional feature (early-fusion). The effects of other modeling decisions varied between institutions though performance increased with more compute-intensive models such as random forest. This work highlights the importance of considering methodological decision points in interpreting the impact on prediction performance when including PRS information in clinical models.

Preprint

Patient and provider perspectives on polygenic risk scores: implications for clinical reporting and utilization.

BACKGROUND: Polygenic risk scores (PRS), which offer information about genomic risk for common diseases, have been proposed for clinical implementation. The ways in which PRS information may influence a patient's health trajectory depend on how both the patient and their primary care provider (PCP) interpret and act on PRS information. We aimed to probe patient and PCP responses to PRS clinical reporting choices METHODS: Qualitative semi-structured interviews of both patients (N=25) and PCPs (N=21) exploring responses to mock PRS clinical reports of two different designs: binary and continuous representations of PRS. RESULTS: Many patients did not understand the numbers representing risk, with high numeracy patients being the exception. However, all the patients still understood a key takeaway that they should ask their PCP about actions to lower their disease risk. PCPs described a diverse range of heuristics they would use to interpret and act on PRS information. Three separate use cases for PRS emerged: to aid in gray-area clinical decision-making, to encourage patients to do what PCPs think patients should be doing anyway (such as exercising regularly), and to identify previously unrecognized high-risk patients. PCPs indicated that receiving "below average risk" information could be both beneficial and potentially harmful, depending on the use case. For "increased risk" patients, PCPs were favorable towards integrating PRS information into their practice, though some would only act in the presence of evidence-based guidelines. PCPs describe the report as more than a way to convey information, viewing it as something to structure the whole interaction with the patient. Both patients and PCPs preferred the continuous over the binary representation of PRS (23/25 and 17/21, respectively). We offer recommendations for the developers of PRS to consider for PRS clinical report design in the light of these patient and PCP viewpoints. CONCLUSIONS: PCPs saw PRS information as a natural extension of their current practice. The most pressing gap for PRS implementation is evidence for clinical utility. Careful clinical report design can help ensure that benefits are realized and harms are minimized.

Clinical Decision-Making