PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n = 5), support vector machines (n = 4), k-nearest neighbor (n = 3), decision trees (n = 3), random forests (n = 5), neural networks (n = 2), linear discriminant analysis (n = 1), and pre-trained CNNs (n = 1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n = 12 to n = 39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans↗

Symbolic, neural, and Bayesian machine learning models for predicting carcinogenicity of chemical compounds.

Experimental programs have been underway for several years to determine the environmental effects of chemical compounds, mixtures, and the like. Among these programs is the National Toxicology Program (NTP) on rodent carcinogenicity. Because these experiments are costly and time-consuming, the rate at which test articles (i.e., chemicals) can be tested is limited. The ability to predict the outcome of the analysis at various points in the process would facilitate informed decisions about the allocation of testing resources. To assist human experts in organizing an empirical testing regime, and to try to shed light on mechanisms of toxicity, we constructed toxicity models using various machine learning and data mining methods, both existing and those of our own devising. These models took the form of decision trees, rule sets, neural networks, rules extracted from trained neural networks, and Bayesian classifiers. As a training set, we used recent results from rodent carcinogenicity bioassays conducted by the NTP on 226 test articles. We performed 10-way cross-validation on each of our models to approximate their expected error rates on unseen data. The data set consists of physical-chemical parameters of test articles, alerting chemical substructures, salmonella mutagenicity assay results, subchronic histopathology data, and information on route, strain, and sex/species for 744 individual experiments. These results contribute to the ongoing process of evaluating and interpreting the data collected from chemical toxicity studies.

Animals↗

Predicting the phosphorylation sites using hidden Markov models and machine learning methods.

Accurately predicting phosphorylation sites in proteins is an important issue in postgenomics, for which how to efficiently extract the most predictive features from amino acid sequences for modeling is still challenging. Although both the distributed encoding method and the bio-basis function method work well, they still have some limits in use. The distributed encoding method is unable to code the biological content in sequences efficiently, whereas the bio-basis function method is a nonparametric method, which is often computationally expensive. As hidden Markov models (HMMs) can be used to generate one model for one cluster of aligned protein sequences, the aim in this study is to use HMMs to extract features from amino acid sequences, where sequence clusters are determined using available biological knowledge. In this novel method, HMMs are first constructed using functional sequences only. Both functional and nonfunctional training sequences are then inputted into the trained HMMs to generate functional and nonfunctional feature vectors. From this, a machine learning algorithm is used to construct a classifier based on these feature vectors. It is found in this work that (1) this method provides much better prediction accuracy than the use of HMMs only for prediction, and (2) the support vector machines (SVMs) algorithm outperforms decision trees and neural network algorithms when they are constructed on the features extracted using the trained HMMs.

Algorithms↗

Comparison of machine learning techniques with classical statistical models in predicting health outcomes.

Several machine learning techniques (multilayer and single layer perceptron, logistic regression, least square linear separation and support vector machines) are applied to calculate the risk of death from two biomedical data sets, one from patient care records, and another from a population survey. Each dataset contained multiple sources of information: history of related symptoms and other illnesses, physical examination findings, laboratory tests, medications (patient records dataset), health attitudes, and disabilities in activities of daily living (survey dataset). Each technique showed very good mortality prediction in the acute patients data sample (AUC up to 0.89) and fair prediction accuracy for six year mortality (AUC from 0.70 to 0.76) in individuals from epidemiological database surveys. The results suggest that the nature of data is of primary importance rather than the learning technique. However, the consistently superior performance of the artificial neural network (multi-layer perceptron) indicates that nonlinear relationships (which cannot be discerned by linear separation techniques) can provide additional improvement in correctly predicting health outcomes.

Aged↗

Human pol II promoter prediction: time series descriptors and machine learning.

Although several in silico promoter prediction methods have been developed to date, they are still limited in predictive performance. The limitations are due to the challenge of selecting appropriate features of promoters that distinguish them from non-promoters and the generalization or predictive ability of the machine-learning algorithms. In this paper we attempt to define a novel approach by using unique descriptors and machine-learning methods for the recognition of eukaryotic polymerase II promoters. In this study, non-linear time series descriptors along with non-linear machine-learning algorithms, such as support vector machine (SVM), are used to discriminate between promoter and non-promoter regions. The basic idea here is to use descriptors that do not depend on the primary DNA sequence and provide a clear distinction between promoter and non-promoter regions. The classification model built on a set of 1000 promoter and 1500 non-promoter sequences, showed a 10-fold cross-validation accuracy of 87% and an independent test set had an accuracy >85% in both promoter and non-promoter identification. This approach correctly identified all 20 experimentally verified promoters of human chromosome 22. The high sensitivity and selectivity indicates that n-mer frequencies along with non-linear time series descriptors, such as Lyapunov component stability and Tsallis entropy, and supervised machine-learning methods, such as SVMs, can be useful in the identification of pol II promoters.

Algorithms↗

Implementation of a fuzzy prototype-based machine learning method to predict myocardial infarction from coronary angiography.

Formal knowledge on the predictive value of morphological angiographic factors is lacking to estimate the risk of myocardial infarction. This article presents a computer system for predicting the incidence of myocardial infarction from angiographic morphological descriptions of coronary lesions. The system includes two phases. The learning phase consists in extracting from a large database of described stenoses two classes represented by one or several fuzzy prototypes. One class corresponds to stenoses leading to infarction and the other to stenoses not leading to that event. The evaluation phase consists in classifying a stenosis according to its morphological characteristics in one of these two classes. The learning method is based on a fuzzy supervised Machine Learning algorithm that combines some aspects of the K-nearest neighbours clustering approach with a defined measure of similarity, and a prototype induction function from the most similar stenoses, taking into account their degree of typicality. The current results of the evaluation phase to correctly predicted X% stenoses for their risk of myocardial infarction. This article emphasizes the feasibility of the approach, however, the learning phase relies on some heuristics that should be validated to get a formal evaluation of the system.

Artificial Intelligence↗

A machine learning approach to predicting peptide fragmentation spectra.

Accurate peptide identification from tandem mass spectrometry experiments is the cornerstone of proteomics. Although various approaches for matching database sequences with experimental spectra have been developed to date (e.g. Sequest, Mascot) the sensitivity and specificity of peptide identification have not yet reached their full potential. This is in part due to the tradeoffs between robustness and accuracy of the existing methods with respect to the non-uniform nature of peptide fragmentation and bond cleavages induced by different mass spectrometers. Accordingly, it is expected that new approaches to de novo predicting peptide fragmentation spectra will enable more accurate peptide identification. To address this problem, here we used a data-driven approach to learn peptide fragmentation rules in mass spectrometry, in the form of posterior probabilities, for various fragment-ion types of doubly and triply charged precursor ions. We show that the accuracy of our neural-network based methodology is useful for subsequent peptide database searches and that the most useful rules of fragmentation significantly differ across ion and precursor types.

Amino Acids↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve–based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC = 0.844). The NODM cohort was stratified into high- (n = 2,362) and low-risk (n = 5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Proteomics-enabled learning machine algorithms enhance the prediction of cardiovascular diseases in patients with type 2 diabetes mellitus.

BACKGROUND AND AIMS: Estimating the risk of cardiovascular disease (CVD) complications in type 2 diabetes mellitus (T2DM) patients is critical in the medical decision-making process. This study aimed to use a machine learning technique combined with proteomics to develop personalized models for predicting CVD in patients with T2DM. METHODS AND RESULTS: In total, 874 patients with T2DM and 2,920 Olink proteins obtained from the UK Biobank were used in this study. Proteins were screened using Cox regression and LASSO regression. A basic model containing clinical features and a full model combining proteome and clinical features were constructed using the random survival forest algorithm. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictive performance of the models and compare them with other CVD predictive models. Compared with the basic model, the full model performed better in predicting CVD, with time-dependent AUCs of 0.81 (3 years), 0.74 (5 years) and 0.74 (10 years) (0.77, 0.69 and 0.67). We calculated the risk scores of the Framingham, ASCVD and Score2-Diabetes models. The results revealed that the prediction performance of the full model was also better than that of the abovementioned models. In terms of differentiation accuracy, the results of the net reclassification improvement index and integrated discrimination improvement index showed that the full model can identify high-risk individuals more accurately (accuracy rate: 79% vs. 69%). CONCLUSIONS: Proteomics can be used to predict cardiovascular complications in diabetic patients. It is also necessary to consider the applicability of the model due to the limitations of the sample size and the constraints of proteomics in clinical applications.

Humans↗

Feature selection and the class imbalance problem in predicting protein function from sequence.

When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence features. However, there are two issues to consider before successful functional prediction can take place: identifying discriminatory features, and overcoming the challenge of a large imbalance in the training data. We show that by applying feature subset selection followed by undersampling of the majority class, significantly better support vector machine (SVM) classifiers are generated compared with standard machine learning approaches. As well as revealing that the features selected could have the potential to advance our understanding of the relationship between sequence and function, we also show that undersampling to produce fully balanced data significantly improves performance. The best discriminating ability is achieved using SVMs together with feature selection and full undersampling; this approach strongly outperforms other competitive learning algorithms. We conclude that this combined approach can generate powerful machine learning classifiers for predicting protein function directly from sequence.

Algorithms↗

Comparison of Cox regression with other methods for determining prediction models and nomograms.

PURPOSE: There is controversy as to whether artificial neural networks and other machine learning methods provide predictions that are more accurate than those provided by traditional statistical models when applied to censored data. MATERIALS AND METHODS: Several machine learning prediction methods are compared with Cox proportional hazards regression using 3 large urological datasets. As a measure of predictive ability, discrimination that is similar to an area under the receiver operating characteristic curve is computed for each. RESULTS: In all 3 datasets Cox regression provided comparable or superior predictions compared with neural networks and other machine learning techniques. In general, this finding is consistent with the literature. CONCLUSIONS: Although theoretically attractive, artificial neural networks and other machine learning techniques do not often provide an improvement in predictive accuracy over Cox regression.

Brachytherapy↗

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning↗

Protein secondary structure prediction using logic-based machine learning.

Many attempts have been made to solve the problem of predicting protein secondary structure from the primary sequence but the best performance results are still disappointing. In this paper, the use of a machine learning algorithm which allows relational descriptions is shown to lead to improved performance. The Inductive Logic Programming computer program, Golem, was applied to learning secondary structure prediction rules for alpha/alpha domain type proteins. The input to the program consisted of 12 non-homologous proteins (1612 residues) of known structure, together with a background knowledge describing the chemical and physical properties of the residues. Golem learned a small set of rules that predict which residues are part of the alpha-helices--based on their positional relationships and chemical and physical properties. The rules were tested on four independent non-homologous proteins (416 residues) giving an accuracy of 81% (+/- 2%). This is an improvement, on identical data, over the previously reported result of 73% by King and Sternberg (1990, J. Mol. Biol., 216, 441-457) using the machine learning program PROMIS, and of 72% using the standard Garnier-Osguthorpe-Robson method. The best previously reported result in the literature for the alpha/alpha domain type is 76%, achieved using a neural net approach. Machine learning also has the advantage over neural network and statistical methods in producing more understandable results.

Amino Acid Sequence↗

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article↗

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans↗

Comparing statistical and machine learning classifiers: alternatives for predictive modeling in human factors research.

Multivariate classification models play an increasingly important role in human factors research. In the past, these models have been based primarily on discriminant analysis and logistic regression. Models developed from machine learning research offer the human factors professional a viable alternative to these traditional statistical classification methods. To illustrate this point, two machine learning approaches--genetic programming and decision tree induction--were used to construct classification models designed to predict whether or not a student truck driver would pass his or her commercial driver license (CDL) examination. The models were developed and validated using the curriculum scores and CDL exam performances of 37 student truck drivers who had completed a 320-hr driver training course. Results indicated that the machine learning classification models were superior to discriminant analysis and logistic regression in terms of predictive accuracy. Actual or potential applications of this research include the creation of models that more accurately predict human performance outcomes.

Adolescent↗

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans↗