PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

The Progress of Gout Prediction Models Based on Multi-source Data.

INTRODUCTION: Gout, a highly serious inflammatory disease that is caused by monosodium urate crystals, is becoming an increasingly significant health concern. Artificial Intelligence and multi-omics-based research have made significant gains for the early detection and prevention of gout based on diverse approaches. This review intends to summarize current advances in forecasting gout susceptibility and gout-related symptoms, evaluate the predictive efficacy of different features, and ascertain which clinical and omics characteristics are most effective in these prediction models. METHODS: We explored the PubMed database after 2010 using keywords such as "gout", "predictive model", "risk prediction", and "machine learning", and confined our search to Englishlanguage articles. The original peer-reviewed research articles that developed gout models were selected. Research that was not original or lacked internal validation was excluded. RESULTS: Clinical features, genomics, microbiomics, radiomics, and metabolomics have been utilized to construct models related to gout and have demonstrated excellent predictive performance. Multisource data prediction models usually exhibit better effectiveness. DISCUSSION: Gout-oriented models performed excellently in predictive performance but present limitations in certain clinical and omics domains. However, if they are to affect actual patient care, they must overcome some external confirmation roadblocks and the fiscal and practical implications they will face ahead of time. CONCLUSION: This review indicates that clinical and multi-omics models of gout are significant instruments for clinical decision-making. The models constructed in these studies may be crucial for the treatment of gout and its practical benefits.

Gout

A comparison between predicted VO2 max from the Astrand procedure and the Canadian Home Fitness Test.

The purpose of this study was to compare the predicted maximal oxygen concumption derived from the Canadian Home Fitness Test (CHFT) and the Astrand ergometer test to the observed VO2 max determined from a progressive multi-stage treadmill test. Sixty-four sedentary subjects (35 males and 29 females) ranging in age from 20 to 54 years participated in the study. The mean VO2 max measured on the treadmill for males and females was 34.6 +/- 6.0 ml/kg/min while the Astrand procedure predicted a mean VO2max of 29.6 +/- 6.5 ml/kg/min and the CHFT predicted a mean VO2 max of 34.8 +/- 5.0 ml/kg/min. Statistical analysis revealed a significant under-prediction (P less than 0.001) of the VO2 predicted by the Astrand test to the VO2 max derived from the treadmill test while there were no differences between the treadmill VO2max and that predicted by the CHFT. When the male and female values were analyzed separately, the same results were seen in the males. For the females, however, there were no significant differences among predicted and observed values. It concluded that the CHFT provided an adequate prediction of cardio-respiratory fitness as well as, if not superior to, the Astrand procedure.

Adult

Prediction of bacterial protein-compound interactions with only positive samples.

MOTIVATION: Prediction of Compound-Protein Interactions (CPI) in bacteria is crucial to advance various pharmaceutical and chemical engineering fields, including biocatalysis, drug discovery, and industrial processing. However, current CPI models cannot be applied for bacterial CPI prediction due to the lack of curated negative interaction samples. RESULTS: We propose a novel Positive-Unlabeled (PU) learning framework, named BIN-PU, to address this limitation. BIN-PU generates pseudo positive and negative labels from known positive interaction data, enabling effective training of deep learning models for CPI prediction. We also propose a weighted positive loss function that weights to truly positive samples. We have validated BIN-PU coupled with multiple CPI backbone models, comparing the performance with the existing PU models using bacterial cytochrome P450 (CYP) data. Extensive experiments demonstrate the superiority of BIN-PU over the benchmark models in predicting CPIs with only truly positive samples. Furthermore, we have validated BIN-PU on additional bacterial proteins obtained from literature review, human CYP datasets, and uncurated data for its reproducibility. We have also validated the CPI prediction for the uncurated CYP data with biological and biophysical experiments. BIN-PU represents a significant advancement in CPI prediction for bacterial proteins, opening new possibilities for improving predictive models in related biological interaction tasks. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/datax-lab/CYP.

Bacterial Proteins

CROP: a feature-independent context-aware method for CRISPR-Cas9 frameshift prediction.

MOTIVATION: The CRISPR-Cas9 complex has revolutionized genome-editing technologies. By designing a 20 nt-long guide RNA, a Cas9 nuclease can be guided to cleave almost any genomic target site (followed by NGG). The cleavage induces double-stranded DNA breaks, which are then repaired by cellular pathways. Accurate CRISPR-Cas9 repair-outcome prediction is essential for designing guide RNAs with desired genomic effects, such as gene knockout. A central challenge is quantifying the rate of frameshifts, i.e. repair-outcomes that lead to a change in the local length that is not a multiple of three. Previous methods for frameshift-rate prediction were trained on only a few experimental or cellular contexts, mostly relied on manually defined microhomology features, and were limited by sparse features and class labels. RESULTS: We developed CROP, a feature-independent context-aware repair-outcome prediction method. By aggregating specific repair outcomes as Δlength classes, CROP overcomes class sparsity. We designed CROP to work with variable input sequence lengths and output classes to utilize multiple datasets simultaneously. We benchmarked CROP against state-of-the-art repair-outcome prediction methods over 18 datasets, which we curated and standardized from various studies. Across all datasets, CROP outperformed all competing methods in frameshift-rate prediction. We performed cross-experiment and cross-cellular frameshift-rate predictions to investigate the generalizability of repair mechanisms. Finally, we show that CROP learned microhomology principles from raw sequences without explicit feature engineering, establishing an end-to-end architecture for CRISPR-Cas9 repair-outcome prediction that learns from multiple datasets. AVAILABILITY AND IMPLEMENTATION: CROP is available at https://github.com/OrensteinLab/CROP.

CRISPR-Cas Systems

Addition of CAD polygenic risk score to coronary artery calcium score enhances prediction of MACE.

BACKGROUND: Coronary heart disease (CHD) is prevalent in the United States, highlighting the need for accurate risk prediction to inform primary prevention strategies. While multivariate risk models like the Framingham Risk Score and ACC/AHA Pooled Cohort Equations are commonly utilized, novel risk markers, such as the coronary artery calcium score (CACS) and polygenic risk score (PRS), are increasingly gaining recognition. OBJECTIVES: This study aimed to compare the diagnostic utility of CACS and CAD PRS, both individually and in combination, for predicting major adverse cardiovascular events (MACE). METHODS: We conducted a retrospective analysis of a cohort comprising 1,380 predominantly Caucasian participants from the Sanford Health System. CAD PRS was constructed using genetic variants, while CACS was assessed via cardiac computed tomography (CT). Statistical analyses evaluated the relationship between each modality and MACE. RESULTS: Both CAD PRS and CACS were significantly associated with future MACE. Following the adjustment for covariates, the area under the curve (AUC) for both the CACS and PRS models was comparable, indicating similar predictive capabilities for MACE. However, the combination of CAD PRS with CACS significantly enhanced predictive accuracy, outperforming either modality alone. CONCLUSIONS: This study underscores the value of integrating CACS and CAD PRS in predicting MACE. The synergistic effect of CAD PRS combined with CACS markedly improves predictive power. Further research and prospective studies are necessary to validate these findings and assess their clinical implications. Investigating the interactions between PRS and CACS will be crucial for refining cardiovascular risk prediction and optimizing prevention strategies.

cardiac genetics

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum

A comparison of four methods of predicting arch length.

1. Four arch length prediction equations (Nance, Johnston-Tanaka, Moyers, and Hixon-Oldfather) were compared by examining pretreatment casts, pretreatment intraoral radiographs, and posttreatment casts of forty-one patients of mixed-dentition age. 2. A comparison of correlation coefficients and slopes of the predicted arch length versus the actual arch lengths revealed that the Hixon-Oldfather method conformed closest to the ideal. 3. No combination of the four methods produced a more accurate equation than the single most accurate method. 4. Neither the sex of the patient nor the type of occlusion affected the prediction accuracy of any of the four equations. 5. All methods tend to overpredict the arch length size by 1 to 3 mm., with the exception of the Hixon-Oldfather equation, which underpredicted by approximately 0.5 mm. 6. An analysis of the intrainvestigator error showed a very low standard error of estimate for individual tooth measurements and for the prediction values. 7. A variance analysis showed that most of the variation was due to arch length (85%), a slight amount was due to the prediction method (8%), and 6% of the variation was due to the rater. 8. A low correlation was found between space available versus actual discrepancy and space available versus actual arch length. 9. High correlation coefficients were found for the predicted arch lengths when compared with the actual arch lengths. As expected, the correlation coefficients for the predicted widths of only the canines and premolars compared with the actual widths were not quite as high.

Bicuspid

Risk Factors and Predictive Model for Postoperative High Myopia in Children Undergoing Congenital Cataract Surgery With Intraocular Lens Implantation.

PURPOSE: To identify risk factors associated with the development of high myopia following congenital cataract surgery and to establish a robust predictive model. DESIGN: Retrospective clinical cohort study. SUBJECTS: This retrospective study included 106 pediatric patients who underwent congenital cataract surgery with primary IOL implantation (mean follow-up 8.19 years). The model was externally validated in an independent cohort of 72 patients with a mean follow-up of 7.83 years. METHODS: Preoperative and postoperative ocular biometric parameters were collected. Risk factors for postoperative high myopia were analyzed using Cox proportional hazards regression, which served as the basis for model construction. The predictive performance of the model was rigorously evaluated for discrimination and calibration. Discriminative ability was quantified using Harrell's C-index and the area under the receiver operating characteristic curve (AUC). Model calibration was assessed via calibration plots by comparing predicted probabilities with actual observed outcomes. Internal validation was performed using a bootstrapping method (500 iterations) to ensure model stability and adjust for potential overfitting. RESULTS: An initial postoperative refraction of <+0.75D, and a higher IOL Power to Axial length Ratio (IOL/AL ratio) were identified as significant risk factors for the development of postoperative high myopia. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. The predictive model demonstrated robust performance, achieving a C-index of 0.711 (internal validation C-index: 0.713). The area under the receiver operating characteristic curve (AUC) values for predicting high myopia at 5 and 10 years were 0.858 and 0.745, respectively. Furthermore, calibration curves demonstrated excellent agreement between the predicted and observed outcomes throughout the follow-up period. In external validation, the model achieved a C-index of 0.825, 5-year AUC of 0.833, and 10-year AUC of 0.713. CONCLUSIONS: Our analysis established that initial postoperative refraction <+0.75D, and an elevated IOL/AL ratio are key determinants of high myopia risk following surgery. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. This predictive framework provides clinicians with a practical tool to optimize preoperative IOL selection and identify high-risk infants who require vigilant myopia prevention and balanced amblyopia management.

Humans

Gastrointestinal digestion governs insect protein hydrolysis and predicted bioactive peptide release: Species-dependent implications for functional food applications.

This study investigates the digestion of insect proteins and the release of predicted bioactive peptides during human gastrointestinal digestion. Using the Infogest in vitro model, mealworm, cricket, and black soldier fly larvae (BSFL) proteins were digested and analyzed through discovery proteomics and bioinformatics to identify predicted bioactive peptides. Sequential windowed acquisition of all theoretical fragment ion mass spectra (SWATH-MS) quantified insect proteins including predicted bioactive peptide precursor proteins, the precursors of predicted bioactive peptides. Results indicated that gastrointestinal digestion strongly influences peptide release, with the gastric phase exhibiting a richer predicted bioactive peptide profile than the small intestinal phase. Many predicted bioactive peptides were rapidly hydrolysed under small intestine conditions, which may lead to reduced stability or diminished activity in vivo, potentially explaining why certain peptides show strong bioactivity in vitro but limited effects in vivo. Additionally, predicted bioactive peptide release varied by insect species, influenced by genetic factors and peptide abundance. These findings highlight the importance of species selection and consideration of proteolytic digestion patterns in optimizing insect-derived bioactive peptides for functional foods and nutraceutical applications.

Animals

CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity.

Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network to learn more robust protein representations alongside a standard graph neural network to learn molecular representations to predict drug-target affinity. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting drug-target affinity and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every FDA-approved drug against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.

deep learning

Prediction of survival in patients with acute myocardial infarction. A clinical study on 100 consecutive patients.

Expected survival after acute myocardial infarction (AMI) in 100 consecutive patients was predicted by three doctors and two nurses at the time of discharge from a CCU. Predictions were compared with various coronary prognostic indices (CPI) and were found to be too optimistic for the first 9 months. Experienced physicians made more reliable predictions than junior physicians and nurses. All patients with a predicted survival of more than 10 years were alive after 1 year and all with predicted death within one month died during the first year. Intermediate predictions were unreliable with reference to the one-year survival. Regardless of which CPI was used, a low index score carried a very low one-year mortality and high index a high mortality. Intermediate index scores were unreliable. A comparison between the predictions and index scores showed that there was no difference in sensitivity and specificity between the methods. Our study thus shows that patients with either a very good or a very poor prognosis will be identified regardless of the method used. The problem of identifying the individual with an intermediate risk remains to be solved.

Acute Disease

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly popular approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n&#x2009;=&#x2009;1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy&#x2009;=&#x2009;89%) using 3728 features and MoRSAE (accuracy&#x2009;=&#x2009;84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta&#x2009;=&#x2009;0.6839, p=0.006), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta&#x2009;=&#x2009;1.92; MoRS, beta&#x2009;=&#x2009;1.99 and MoRSAE, beta&#x2009;=&#x2009;1.77) displayed a significant (p&#x2009;<&#x2009;0.001) predictive power for post-deployment PTSD. CONCLUSION: The inclusion of exposure variables adds to the predictive power of MRS. Classification-based MRS may be useful in predicting risk of future PTSD in populations with anticipated trauma exposure. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting PTSD and, relatedly, improve their performance in independent cohorts.

Humans

SCMO: a deep learning model integrating the single-cell resolution TME ecosystem and multi-omics for survival prediction in CRC patients.

BACKGROUND: Colorectal cancer (CRC) remains a leading cause of global cancer mortality, highlighting the need for precise survival prediction to guide clinical decisions. Although tissue-level multi-omics is widely utilized for survival prediction, its limited resolution cannot capture tumor heterogeneity. Single-cell RNA sequencing (scRNA-seq) enables dissection of the tumor microenvironment (TME) at cellular resolution, supporting personalized prognostic assessment. METHODS: We collected 213 CRC scRNA-seq samples and established a CRC-specific TME atlas comprising 339,060 cells. Using this atlas as a reference, we deconvolved bulk RNA-seq data from TCGA-CRC cohort with the EcoTyper algorithm to reconstruct TME features. Clinical, genomic, and transcriptomic data were obtained from the Xena platform; microbial data were sourced from the BIC database. We integrated TME and multi-omics features through a self-normalizing neural network to construct a deep learning model (single-cell resolution TME ecosystem with multi-omics data [SCMO]) for survival prediction. To enhance interpretability, we utilized the Integrated Gradients algorithm and spatial transcriptomic data to analyze multi-omics and TME features. We performed anticancer drug screening with tumor necrosis factor receptor-associated protein 1 (TRAP1), a critical feature according to the Integrated Gradients algorithm, as a potential target. RESULTS: We identified 13 survival-related TME features from the CRC-specific atlas: 12 cell states and one multi-cellular ecosystem. SCMO, which combined TME and multi-omics features, improved survival prediction and outperformed existing methods, achieving a concordance index of 0.762. The SCMO demonstrated robust performance for long-term predictions, achieving areas under the curve (AUCs) of 0.752, 0.772, and 0.869 for 1-, 3-, and 5-year predictions in the training set, with corresponding test set AUCs of 0.639, 0.756, and 0.772. TME features from the SCMO model revealed that ecosystem density increased with CRC malignancy. Multi-omics features included TRAP1 as a potential drug target. Drug screening identified saikosaponin A as a novel TRAP1 inhibitor, and its anticancer activity was validated in vitro. We developed SCMO-Lite, a simplified model incorporating 12 high-attribution-weight multi-omics features, which demonstrated robust risk stratification. CONCLUSIONS: SCMO combines analytical precision with biological interpretability, offering novel insights for oncology survival prediction.

Humans