PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity.

Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network to learn more robust protein representations alongside a standard graph neural network to learn molecular representations to predict drug-target affinity. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting drug-target affinity and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every FDA-approved drug against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.

deep learning

Prediction of survival in patients with acute myocardial infarction. A clinical study on 100 consecutive patients.

Expected survival after acute myocardial infarction (AMI) in 100 consecutive patients was predicted by three doctors and two nurses at the time of discharge from a CCU. Predictions were compared with various coronary prognostic indices (CPI) and were found to be too optimistic for the first 9 months. Experienced physicians made more reliable predictions than junior physicians and nurses. All patients with a predicted survival of more than 10 years were alive after 1 year and all with predicted death within one month died during the first year. Intermediate predictions were unreliable with reference to the one-year survival. Regardless of which CPI was used, a low index score carried a very low one-year mortality and high index a high mortality. Intermediate index scores were unreliable. A comparison between the predictions and index scores showed that there was no difference in sensitivity and specificity between the methods. Our study thus shows that patients with either a very good or a very poor prognosis will be identified regardless of the method used. The problem of identifying the individual with an intermediate risk remains to be solved.

Acute Disease

The reliability of prediction of outcome in spina bifida.

The clinical findings in 85 neonates with spina bifida were given to two neurosurgeons and two paediatricians, who were asked to predict from them the length of survival and quality of survival with regard to intellect, locomotion and continence, without their knowing the actual outcome. All four clinicians correctly predicted the survival of infants with meningocele, closed myelocele and encephalocele. The paediatricians correctly predicted the survival of all infants with open myelocele who actually survived, but also included some who had died. The surgeons correctly predicted the deaths of all those with open myelocele who actually died, but expected a considerable number to die who in fact survived. All four clinicians were similar in their predictions of intellect: they underestimated the outcome in patients with successfully shunted hydrocephalus, they overestimated the intellect in patients who had developed intracranial infection and shunt blockage, and they largely underestimated the outcome in the patients who did not require shunts. They made correct predictions for limb and sphincter function in nearly all the survivors. This investigation underlines the problem of selection for treatment caused by the inability to predict the complications of hydrocephalus and infection. Reasons for the differences between the expectations of the paediatricians and surgeons, and the implications of the results of this study for selection for surgery are discussed. It is suggested that limb paralysis and incontinence ought not to be considered as factors excluding infants from treatment.

Child

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly popular approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n&#x2009;=&#x2009;1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy&#x2009;=&#x2009;89%) using 3728 features and MoRSAE (accuracy&#x2009;=&#x2009;84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta&#x2009;=&#x2009;0.6839, p=0.006), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta&#x2009;=&#x2009;1.92; MoRS, beta&#x2009;=&#x2009;1.99 and MoRSAE, beta&#x2009;=&#x2009;1.77) displayed a significant (p&#x2009;<&#x2009;0.001) predictive power for post-deployment PTSD. CONCLUSION: The inclusion of exposure variables adds to the predictive power of MRS. Classification-based MRS may be useful in predicting risk of future PTSD in populations with anticipated trauma exposure. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting PTSD and, relatedly, improve their performance in independent cohorts.

Humans

SCMO: a deep learning model integrating the single-cell resolution TME ecosystem and multi-omics for survival prediction in CRC patients.

BACKGROUND: Colorectal cancer (CRC) remains a leading cause of global cancer mortality, highlighting the need for precise survival prediction to guide clinical decisions. Although tissue-level multi-omics is widely utilized for survival prediction, its limited resolution cannot capture tumor heterogeneity. Single-cell RNA sequencing (scRNA-seq) enables dissection of the tumor microenvironment (TME) at cellular resolution, supporting personalized prognostic assessment. METHODS: We collected 213 CRC scRNA-seq samples and established a CRC-specific TME atlas comprising 339,060 cells. Using this atlas as a reference, we deconvolved bulk RNA-seq data from TCGA-CRC cohort with the EcoTyper algorithm to reconstruct TME features. Clinical, genomic, and transcriptomic data were obtained from the Xena platform; microbial data were sourced from the BIC database. We integrated TME and multi-omics features through a self-normalizing neural network to construct a deep learning model (single-cell resolution TME ecosystem with multi-omics data [SCMO]) for survival prediction. To enhance interpretability, we utilized the Integrated Gradients algorithm and spatial transcriptomic data to analyze multi-omics and TME features. We performed anticancer drug screening with tumor necrosis factor receptor-associated protein 1 (TRAP1), a critical feature according to the Integrated Gradients algorithm, as a potential target. RESULTS: We identified 13 survival-related TME features from the CRC-specific atlas: 12 cell states and one multi-cellular ecosystem. SCMO, which combined TME and multi-omics features, improved survival prediction and outperformed existing methods, achieving a concordance index of 0.762. The SCMO demonstrated robust performance for long-term predictions, achieving areas under the curve (AUCs) of 0.752, 0.772, and 0.869 for 1-, 3-, and 5-year predictions in the training set, with corresponding test set AUCs of 0.639, 0.756, and 0.772. TME features from the SCMO model revealed that ecosystem density increased with CRC malignancy. Multi-omics features included TRAP1 as a potential drug target. Drug screening identified saikosaponin A as a novel TRAP1 inhibitor, and its anticancer activity was validated in vitro. We developed SCMO-Lite, a simplified model incorporating 12 high-attribution-weight multi-omics features, which demonstrated robust risk stratification. CONCLUSIONS: SCMO combines analytical precision with biological interpretability, offering novel insights for oncology survival prediction.

Humans

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

Correlation of predicted versus measured creatinine clearance values in burn patients.

Four methods for predicting creatinine clearance (Ccr) from serum creatinine concentration (Scr) were evaluated in 19 male burn patients with burn wound sepsis. Measured Ccr values were calculated from 24-hour urinary catheter collections. Steady state Scr values were obtained during the same collection interval. Predicted Ccr values were derived from Scr using the methods of Cockcroft and Gault (Method II), Siersbaek-Nielsen, Kampmann and others (Method III) and Jeliffe (Methods I and IV). Wide differences between measured and predicted values were observed but were statistically significant (p less than 0.05) for Method I only. The smallest mean difference (+/-0.02 ml/min/1.73 m2) occurred with Method II measured-predicted data pairs. Method III predicted Ccr values which correlated best with measured values (r=0.770) and showed the least variability (+/-7.6 ml/min/1.73 m2). All methods appeared to overestimate when measured Ccr was less than 60 ml/min/1.73 m2. Use of estimated lean body weights did not improve correlations between predicted and measured Ccr values. While Methods II and III may provide useful initial approximations of Ccr in burn patients, reliance upon predicted Ccr values for dosage modification in burn patients may result in an insufficient reduction in dosage. Whenever possible, dosage regimens for drugs with narrow therapeutic margins should be developed or adjusted using pharmacokinetic values determined in the individual patient.

Adult

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Polygenic Risk Scores for Preeclampsia Prediction Beyond Gold-Standard Clinical Models in Multiethnic Populations.

BACKGROUND: Preeclampsia is a major cause of maternal and fetal mortality and morbidity. Early risk stratification enables timely preventative therapy in high-risk women. Polygenic risk scores (PGS) improve prediction in complex diseases, but their added value for preeclampsia remains unclear, particularly in comparison to gold-standard first-trimester prediction models and across non-European ancestries. METHODS: We evaluated the performance of both a preeclampsia and systolic blood pressure PGS in 2 prospective pregnancy cohorts with detailed phenotyping: the Fetal Medicine Foundation study (n=5207; 2127 cases) and the Pregnancy Outcome Prediction study (n=3659; 228 cases). Risk models included (1) clinical factors; (2) clinical factors plus PGS; (3) advanced model including first-trimester mean arterial pressure, PAPP-A (pregnancy-associated plasma protein-A), and uterine artery pulsatility index; and (4) advanced model plus PGS. Discriminative performance, measured by the area under the receiver operating characteristic curve, was assessed overall and by ancestry. RESULTS: The preeclampsia PGS was independently associated with preeclampsia (odds ratio per SD, 1.24 [95% CI, 1.17-1.31]; P<0.001). It modestly improved prediction over clinical models (area under the receiver operating characteristic curve 0.746 versus 0.750; P=0.017) but not over the advanced model (area under the receiver operating characteristic curve 0.817 versus 0.818; P=0.326). The systolic blood pressure PGS showed stronger performance, improving prediction over both models in women of European ancestry. No improvement was observed with either score in women of African ancestry. CONCLUSIONS: PGSs for preeclampsia and SBP provide modest added predictive value beyond clinical risk factors in European ancestry women. Limited utility in African ancestry women reflects underrepresentation in the genome-wide association studies used to develop current scores. As cohort sizes grow and models are refined, PGSs may become important tools for equitable risk stratification in maternal health.

Adult

Enhanced Prediction of Peripheral Artery Disease Using Plasma Proteomics Among Individuals Without Diabetes.

BACKGROUND: Although peripheral artery disease (PAD) is an important diabetes complication, a substantial proportion of cases occur among individuals without diabetes. This study aimed to assess the predictive value of plasma proteomics in the long-term risk of PAD among individuals initially free of diabetes. METHODS: Included were 46&#x2009;508 participants (6046 with prediabetes) without diabetes or major cardiovascular disease at recruitment of the UK Biobank. Using multivariable Cox regression models, a total of 2923 unique plasma proteins were assessed for the associations with incident PAD. Significant proteins were subsequently processed by a trained light gradient boosting machine classifier to determine important proteins. Using receiver operating characteristic analyses, the performance of these important proteins in predicting incident PAD were evaluated, in the whole sample and by glycemic status (normoglycemia and prediabetes). RESULTS: During a median follow-up of 12.7&#x2009;years, 461 participants developed PAD. There were 107 proteins associated with incident PAD, with 103 positive associations. The LGBM approach identified 9 proteins (eg, WFDC2 [WAP 4-disulfide core domain protein 2], MMP12 [macrophage metalloelastase], and GDF15 [growth differentiation factor 15]) as the top-ranked proteins based on their importance ordering. Whereas glycated hemoglobin showed very modest predictive accuracy, a panel incorporating these top proteins showed good performance in the prediction of PAD risk (area under the curve 0.820), and it significantly enhanced the prediction beyond traditional risk factors (raising area under the curve from 0.803 to 0.837, DeLong test P=5.21&#xd7;10-3). These observations were consistent for participants with normoglycemia or prediabetes. CONCLUSIONS: Plasma protein biomarkers enhance the prediction of long-term risk for PAD among individuals without diabetes, regardless of glycemic status.

Humans

Beyond Morphology: Reframing Lymph-Node Metastasis Prediction Through Clonal Ecology-Decades-Long Genomic Instability and Polyclonal-to-Monoclonal Transitions as the Missing Dimension in Cancer.

Recent whole-genome, lineage-tracing, single-cell, and spatial studies have reshaped our understanding of tumor evolution, revealing that cancers can arise from polyclonal populations, undergo decades-long genomic instability before clinical detection, and progress through dynamic changes in subclonal composition, cellular state, and ecological organization. These findings challenge the assumption underlying morphology-based prediction models that metastatic risk can be inferred from static histological features alone. Here, we revisit lymph-node metastasis prediction in colorectal cancer through clonal ecology, integrating computational pathology with evolutionary oncology. Drawing on the subclonal switchboard model proposed in 2012 and subsequent artificial intelligence (AI)-enabled approaches for tracking dominant and dormant subclones, we synthesize evidence that metastatic potential reflects clonal ancestry, evolutionary timing, spatial niche architecture, cellular plasticity, intercellular interactions, dormancy, and treatment-driven shifts in subclonal fitness. We define five complementary methodological pillars for operationalizing clonal ecology: single-cell transcriptomics for resolving rare subclones, evolutionary trajectories, and adaptive cell states; lineage tracing and phylogenetics for reconstructing clonal ancestry and divergence; spatial transcriptomics and genomics for mapping subclonal geography and tumor-stromal-immune interactions; longitudinal liquid biopsy surveillance for monitoring residual disease, clonal turnover, and emerging resistance; and AI-enabled multimodal integration for connecting histopathology, genomics, spatial biology, and longitudinal data into predictive ecological-state models. Multiple-instance learning and pathology foundation models provide scalable computational foundations for evolution-aware prediction. Translationally, dormant subclones represent actionable reservoirs of recurrence. A longitudinal clinical and experimental study of KMT2A-rearranged acute myeloid leukemia further supports central predictions of the subclonal switchboard framework by demonstrating treatment-associated shifts in subclonal dominance, persistence of cryptic adaptive programs, and ecological rewiring during resistance and relapse. We propose clonal ecology as a measurable dimension for extending morphology-driven prediction toward integrative models that anticipate evolutionary transitions, identify therapeutic windows, and proactively constrain adaptive tumor ecosystems before resistant or metastatic subclones achieve clinical dominance.

Humans

DPAS-Graph: adaptive spatial-feature relation learning for spatial RNA-to-protein prediction and virtual protein profiling.

Paired spatial multi-omics provides a supervised basis for learning RNA-protein correspondence in situ, but predicting protein abundance from spatial transcriptomic data alone remains challenging across tissue contexts and protein panels. Here, we present DPAS-Graph, an adaptive relation-learning framework for spatial RNA-to-protein prediction. Rather than directly merging spatial proximity and transcriptomic similarity as fixed graph priors, DPAS-Graph represents them as two relation channels on a shared edge support and updates their contributions during representation learning for protein prediction. Its Niche-Coupled Field Encoder combines layer-wise edge-relation modeling, intra-branch relation refinement, and cross-branch residual correction to learn spot representations for protein abundance prediction. In a leave-one-dataset-out benchmark across seven paired spatial multi-omics datasets, DPAS-Graph achieved lower aggregate prediction errors and improved spot-level agreement of protein expression profiles, with gains mainly reflected in error-based metrics and PCC-Spot. Spatial autocorrelation and protein-derived domain agreement analyses were further used to characterize the spatial behavior of the predicted protein maps. When applied to external RNA-only spatial sections, DPAS-Graph generated qualitatively interpretable marker-level virtual protein maps, illustrating its use as a complementary tool for protein-level interpretation of transcriptomics-only spatial data.

RNA

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core C&#x3b1; RMSDs of 0.62&#xa0;&#xc5; for MHC-I and 0.33&#xa0;&#xc5; for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63&#xa0;817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides

Proteomics-enabled learning machine algorithms enhance the prediction of cardiovascular diseases in patients with type 2 diabetes mellitus.

BACKGROUND AND AIMS: Estimating the risk of cardiovascular disease (CVD) complications in type 2 diabetes mellitus (T2DM) patients is critical in the medical decision-making process. This study aimed to use a machine learning technique combined with proteomics to develop personalized models for predicting CVD in patients with T2DM. METHODS AND RESULTS: In total, 874 patients with T2DM and 2,920 Olink proteins obtained from the UK Biobank were used in this study. Proteins were screened using Cox regression and LASSO regression. A basic model containing clinical features and a full model combining proteome and clinical features were constructed using the random survival forest algorithm. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictive performance of the models and compare them with other CVD predictive models. Compared with the basic model, the full model performed better in predicting CVD, with time-dependent AUCs of 0.81 (3&#x2009;years), 0.74 (5&#x2009;years) and 0.74 (10&#x2009;years) (0.77, 0.69 and 0.67). We calculated the risk scores of the Framingham, ASCVD and Score2-Diabetes models. The results revealed that the prediction performance of the full model was also better than that of the abovementioned models. In terms of differentiation accuracy, the results of the net reclassification improvement index and integrated discrimination improvement index showed that the full model can identify high-risk individuals more accurately (accuracy rate: 79% vs. 69%). CONCLUSIONS: Proteomics can be used to predict cardiovascular complications in diabetic patients. It is also necessary to consider the applicability of the model due to the limitations of the sample size and the constraints of proteomics in clinical applications.

Humans

Predictive evolutionary genomics: principles, validation, and practice.

Climate change and habitat loss are driving rapid evolutionary responses in populations world-wide, which creates an urgent need for evolutionary forecasting in conservation and agriculture. Such forecasting can be categorized into three time scales: trait-based models that use multivariate quantitative genetic equations to project correlated phenotypic responses up to c.&#xa0;20 generations, allele-based analyses that model allele frequency dynamics up to 100 generations, and composite adaptation scores that aggregate many small effects to yield predictions across longer horizons. However, these approaches have remained largely disconnected. Here, we present a Bayesian framework that integrates these three complementary approaches for evolutionary prediction. Our framework combines genomic, phenotypic, and environmental data to yield probabilistic predictions with explicit uncertainty. We show how predictive evolutionary forecasts can be validated with experimental evolution, field experimentation, historical specimens, and reciprocal transplants. These validated forecasts can help advance conservation and agricultural programmes by helping predict which populations are at risk of future extinction, optimizing breeding programmes for future climates, and planning ecosystem management under environmental change. By supporting a shift towards more predictive approaches in evolutionary biology, this framework may help improve our ability to manage biodiversity and food security in a changing world.

Genomics

Utilizing evolutionary conservation to detect deleterious mutations and improve genomic prediction in cassava.

INTRODUCTION: Cassava (Manihot esculenta) is an annual root crop which provides the major source of calories for over half a billion people around the world. Since its domestication ~10,000 years ago, cassava has been largely clonally propagated through stem cuttings. Minimal sexual recombination has led to an accumulation of deleterious mutations made evident by heavy inbreeding depression. METHODS: To locate and characterize these deleterious mutations, and to measure selection pressure across the cassava genome, we aligned 52 related Euphorbiaceae and other related species representing millions of years of evolution. With single base-pair resolution of genetic conservation, we used protein structure models, amino acid impact, and evolutionary conservation across the Euphorbiaceae to estimate evolutionary constraint. With known deleterious mutations, we aimed to improve genomic evaluations of plant performance through genomic prediction. We first tested this hypothesis through simulation utilizing multi-kernel GBLUP to predict simulated phenotypes across separate populations of cassava. RESULTS: Simulations showed a sizable increase of prediction accuracy when incorporating functional variants in the model when the trait was determined by<100 quantitative trait loci (QTL). Utilizing deleterious mutations and functional weights informed through evolutionary conservation, we saw improvements in genomic prediction accuracy that were dependent on trait and prediction. CONCLUSION: We showed the potential for using evolutionary information to track functional variation across the genome, in order to improve whole genome trait prediction. We anticipate that continued work to improve genotype accuracy and deleterious mutation assessment will lead to improved genomic assessments of cassava clones.

cassava (Manihot esculenta)