PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Analysing and improving the diagnosis of ischaemic heart disease with machine learning.

Ischaemic heart disease is one of the world's most important causes of mortality, so improvements and rationalization of diagnostic procedures would be very useful. The four diagnostic levels consist of evaluation of signs and symptoms of the disease and ECG (electrocardiogram) at rest, sequential ECG testing during the controlled exercise, myocardial scintigraphy, and finally coronary angiography (which is considered to be the reference method). Machine learning methods may enable objective interpretation of all available results for the same patient and in this way may increase the diagnostic accuracy of each step. We conducted many experiments with various learning algorithms and achieved the performance level comparable to that of clinicians. We also extended the algorithms to deal with non-uniform misclassification costs in order to perform ROC analysis and control the trade-off between sensitivity and specificity. The ROC analysis shows significant improvements of sensitivity and specificity compared to the performance of the clinicians. We further compare the predictive power of standard tests with that of machine learning techniques and show that it can be significantly improved in this way.

Algorithms↗

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies↗

Machine learning for survival analysis: a case study on recurrence of prostate cancer.

Machine learning techniques have recently received considerable attention, especially when used for the construction of prediction models from data. Despite their potential advantages over standard statistical methods, like their ability to model non-linear relationships and construct symbolic and interpretable models, their applications to survival analysis are at best rare, primarily because of the difficulty to appropriately handle censored data. In this paper we propose a schema that enables the use of classification methods--including machine learning classifiers--for survival analysis. To appropriately consider the follow-up time and censoring, we propose a technique that, for the patients for which the event did not occur and have short follow-up times, estimates their probability of event and assigns them a distribution of outcome accordingly. Since most machine learning techniques do not deal with outcome distributions, the schema is implemented using weighted examples. To show the utility of the proposed technique, we investigate a particular problem of building prognostic models for prostate cancer recurrence, where the sole prediction of the probability of event (and not its probability dependency on time) is of interest. A case study on preoperative and postoperative prostate cancer recurrence prediction shows that by incorporating this weighting technique the machine learning tools stand beside modern statistical methods and may, by inducing symbolic recurrence models, provide further insight to relationships within the modeled data.

Artificial Intelligence↗

dsRNAscan maps human dsRNAome, revealing conservation, intermolecular dsRNA, and correlates of ADAR dependency.

The human transcriptome contains millions of A-to-I editing sites arising from an unclear number of poorly characterized dsRNAs. Editing sites reveal the presence of dsRNA, but this method is limited by transcription levels, read depth, and ADAR expression and cannot identify unedited dsRNA. To address these limitations, we developed dsRNAscan. Applying dsRNAscan to the human genome predicted 5 million dsRNAs, mostly in repetitive and intergenic regions. Machine learning models trained on A-to-I editing and RNA structure-probing data identified ∼2.4 million high-confidence predictions, which were enriched at dsRNA-binding protein binding sites. Additionally, we predicted hundreds of dsRNAs conserved across vertebrates and observed thousands of editing-enriched regions suspected to arise from intermolecular dsRNAs formed with sense-antisense transcripts. Quantifying expression of intramolecular and intermolecular dsRNAs accessible to cytoplasmic immune sensors revealed that their ratio correlated with ADAR dependency across cancer cell lines. The human dsRNAome is available as a resource at https://dsrna.chpc.utah.edu/.

A-to-I RNA editing↗

Accurate prediction of solvent accessibility using neural networks-based regression.

Accurate prediction of relative solvent accessibilities (RSAs) of amino acid residues in proteins may be used to facilitate protein structure prediction and functional annotation. Toward that goal we developed a novel method for improved prediction of RSAs. Contrary to other machine learning-based methods from the literature, we do not impose a classification problem with arbitrary boundaries between the classes. Instead, we seek a continuous approximation of the real-value RSA using nonlinear regression, with several feed forward and recurrent neural networks, which are then combined into a consensus predictor. A set of 860 protein structures derived from the PFAM database was used for training, whereas validation of the results was carefully performed on several nonredundant control sets comprising a total of 603 structures derived from new Protein Data Bank structures and had no homology to proteins included in the training. Two classes of alternative predictors were developed for comparison with the regression-based approach: one based on the standard classification approach and the other based on a semicontinuous approximation with the so-called thermometer encoding. Furthermore, a weighted approximation, with errors being scaled by the observed levels of variability in RSA for equivalent residues in families of homologous structures, was applied in order to improve the results. The effects of including evolutionary profiles and the growth of sequence databases were assessed. In accord with the observed levels of variability in RSA for different ranges of RSA values, the regression accuracy is higher for buried than for exposed residues, with overall 15.3-15.8% mean absolute errors and correlation coefficients between the predicted and experimental values of 0.64-0.67 on different control sets. The new method outperforms classification-based algorithms when the real value predictions are projected onto two-class classification problems with several commonly used thresholds to separate exposed and buried residues. For example, classification accuracy of about 77% is consistently achieved on all control sets with a threshold of 25% RSA. A web server that enables RSA prediction using the new method and provides customizable graphical representation of the results is available at http://sable.cchmc.org.

Artificial Intelligence↗

Plasma signals of lung tumor promotion for molecular cancer prevention.

Predicting lung cancer risk would enhance prevention trials. Although the Canakinumab Anti-inflammatory Thrombosis Outcome Study (CANTOS) trial demonstrated reduced lung cancer incidence with interleukin (IL)-1β inhibition, the high number needed to treat (NNT) to prevent lung cancer limits its use in unselected populations. Using machine learning, we identified a 14-protein plasma signature predicting lung cancer more than 5 years before diagnosis. The signature, validated across eight cohorts, was elevated in current smokers and individuals exposed to particulate matter (PM) and linked to lung myeloid and alveolar cells. In epidermal growth factor receptor (EGFR)-driven lung adenocarcinoma, diverse epithelial lineages converged on a keratin8+/claudin4+ alveolar transitional state (KAC), whose transcriptional programs correlated with signature emergence. Components of the signature were induced by PM, oncogenic EGFR, or IL-1β, whereas IL-1β inhibition restrained PM-driven KAC expansion and early tumorigenesis. In CANTOS, the signature identified individuals who seemed to benefit more from anti-IL-1β therapy, lowering the NNT threshold and nominating circulating signals of tumor promotion for prevention.

Humans↗

Predicting protein secondary structure by a support vector machine based on a new coding scheme.

Protein structure prediction is one of the most important problems in modern computational biology. Protein secondary structure prediction is a key step in prediction of protein tertiary structure. There have emerged many methods based on machine learning techniques, such as neural networks (NN) and support vector machine (SVM) etc., to focus on the prediction of the secondary structures. In this paper, a new method was proposed based on SVM. Different from the existing methods, this method takes into account of the physical-chemical properties and structure properties of amino acids. When tested on the most popular dataset CB513, it achieved a Q(3) accuracy of 0.7844, which illustrates that it is one of the top range methods for protein of secondary structure prediction.

Algorithms↗

Computer prediction of allergen proteins from sequence-derived protein structural and physicochemical properties.

BACKGROUND: Computational methods have been developed for predicting allergen proteins from sequence segments that show identity, homology, or motif match to a known allergen. These methods achieve good prediction accuracies, but are less effective for novel proteins with no similarity to any known allergen. METHODS: This work tests the feasibility of using a statistical learning method, support vector machines, as such a method. The prediction system is trained and tested by using 1005 allergen proteins from the Allergome database and 22,469 non-allergen proteins from 7871 Pfam families. RESULTS: Testing results by an independent set of 229 allergen and 6717 non-allergen proteins from 7871 Pfam families show that 93.0% and 99.9% of these are correctly predicted, which are comparable to the best results of other methods. Of the 18 novel allergen proteins non-homologous to any other proteins in the Swissprot database, 88.9% is correctly predicted. A further screening of 168,128 proteins in the Swissprot database finds that 2.9% of the proteins are predicted as allergen proteins, which is consistent with the estimated numbers from motif-based methods. CONCLUSIONS: Our study suggests that SVM is a potentially useful method for predicting allergen proteins and it has certain capability for predicting novel allergen proteins. Our software can be accessed at .

Allergens↗

A multi-scale fusion model based on multi-phase contrast-enhanced CT for predicting pancreatic cancer resectability.

Purpose.Develop a multi-scale fusion model (MSFM) based on multi-phase contrast-enhanced computed tomography (CECT) to predict pancreatic cancer (PC) resectability, thereby assisting expert decision-making.Methods.This retrospective study enrolled 280 patients with PC from four institutions, which were randomly divided into a training cohort (202 patients) and an independent test cohort (78 patients). Three-phase CECT images (arterial, venous, and delayed phases) were used for modeling. The MSFM comprises two sub-networks: (1) a multi-phase fusion network for extracting cross-phase shared fusion features, (2) a phase-specific branch network for capturing phase-specific features; and a post-fusion strategy to generate the final predictive score by integrating the shared fusion features and three groups of phase-specific features. Additionally, a human-machine fusion deep learning model (HMfDL) was constructed by fusing the predictive score of the MSFM with expert assessments.Results.In the independent test, the MSFM achieved an AUC (area under the receiver operating characteristic curve) of 0.8385 (95% CI: 0.7521-0.9249), accuracy of 84.62%, sensitivity of 72.00%, and specificity of 90.57%. This performance outperformed single-phase models (AUC range: 0.7638-0.7781), two-phase models (AUC range: 0.7826-0.7864), and ten states-of-the-art classifiers (AUC range: 0.7404-0.7796). The HMfDL further improved the performance, reaching an AUC of 0.8626 (95% CI: 0.7853-0.9400), accuracy of 91.03%, sensitivity of 80.00%, and specificity of 96.23%. Notably, the HMfDL corrected 58.82% of misdiagnosis made by experts.Conclusions. The MSFM effectively fuses multi-phase CECT to enable highly accurate predictions of PC resectability, and provides valuable support for expert decision-making through HMfDL.

Humans↗

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

Immunoinformatics Approach for Optimization of Targeted Vaccine Design: New Paradigm in Clinical Trials and Healthcare Management.

INTRODUCTION: The immunoinformatics approach combines bioinformatics and computational tools, offering a revolutionary method for improving vaccine development by analyzing immune responses at the molecular level. Immunoinformatics enables the creation of customized vaccines designed for specific infections or cancer cells. OBJECTIVE: The primary objective of immunoinformatics is to enhance the vaccine development process by predicting and boosting the body's immune response. It aims to identify potential immunogenic epitopes and biomarkers that are important for creating vaccines with greater specificity and efficacy, especially when dealing with large-scale data. METHODS: Immunoinformatics utilizes a combination of proteomic, genomic, and epigenomic data, as well as machine learning algorithms and artificial intelligence techniques. These tools predict how various immunological components, e.g., T-cell and B-cell epitopes, interact with the immune system. This approach allows researchers to avoid traditional trial-and-error methods, enabling the efficient identification of potential vaccine candidates. Additionally, personalized vaccines can be developed by considering individual genetic and immunological characteristics. RESULTS: The use of immunoinformatics techniques accelerates the screening of vaccine candidates, enhances patient stratification, and optimizes formulations for clinical trials. This approach has been shown to improve vaccine safety, efficacy, and development speed. It also holds promise for managing healthcare on a large scale by producing vaccines tailored to specific populations, thereby improving the overall effectiveness of vaccination programs. CONCLUSION: Immunoinformatics represents a transformative approach to vaccine research, improving clinical trial efficiency and enabling the development of more reliable, flexible, and personalized vaccines. This approach has the potential to significantly enhance global healthcare outcomes by accelerating the vaccine development process and optimizing vaccination strategies.

Immunoinformatics↗

Neural pattern formation via a competitive Hebbian mechanism.

In this contribution we investigate a simple pattern formation process [9,10] based on Hebbian learning and competitive interactions within cortex. This process generates spatial representations of afferent (sensory) information which strongly resemble patterns of response properties of neurons commonly called brain maps. For one of the most thoroughly studied phenomena in cortical development, the formation of topographic maps, orientation and ocular dominance columns in macaque striate cortex, the process, for example, generates the observed patterns of receptive field properties including the recently described correlations between orientation preference and ocular dominance. Competitive Hebbian learning has not only proven to be a useful concept in the understanding of development and plasticity in several brain areas, but the underlying principles have have been successfully applied to problems in machine learning [22]. The model's universality, simplicity, predictive power, and usefulness warrants a closer investigation.

Animals↗

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans↗

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames↗

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease↗

Ecologic niche modeling and differentiation of populations of Triatoma brasiliensis neiva, 1911, the most important Chagas' disease vector in northeastern Brazil (hemiptera, reduviidae, triatominae).

Ecologic niche modeling has allowed numerous advances in understanding the geographic ecology of species, including distributional predictions, distributional change and invasion, and assessment of ecologic differences. We used this tool to characterize ecologic differentiation of Triatoma brasiliensis populations, the most important Chagas' disease vector in northeastern Brazil. The species' ecologic niche was modeled based on data from the Fundação Nacional de Saúde of Brazil (1997-1999) with the Genetic Algorithm for Rule-Set Prediction (GARP). This method involves a machine-learning approach to detecting associations between occurrence points and ecologic characteristics of regions. Four independent "ecologic niche models" were developed and used to test for ecologic differences among T. brasiliensis populations. These models confirmed four ecologically distinct and differentiated populations, and allowed characterization of dimensions of niche differentiation. Patterns of ecologic similarity matched patterns of molecular differentiation, suggesting that T. brasiliensis is a complex of distinct populations at various points in the process of speciation.

Animals↗

Machine learning-based analysis of the impact of 5' untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions↗

Proteo-metabolomic integration identifies stage-specific candidate biomarkers for Parkinson's disease.

Parkinson's disease (PD) is a progressive neurodegenerative disorder with a prolonged prodromal phase and complex motor symptoms. Despite improved clinical criteria, early diagnosis and longitudinal monitoring remain challenging. While cerebrospinal fluid (CSF) and plasma metabolites and proteins show biomarker potential, their utility in predictive models is insufficiently characterized. We employed a secondary computational approach to integrate proteometabolomic profiles from CSF and plasma samples of >1100 Parkinson's Progression Markers Initiative (PPMI) participants. Using multi-omics machine learning, we identified biofluid-specific signatures and evaluated predictive performance. Twenty-one biomarker candidates were validated across three models (SVM, GLMNET, RF); SVM and GLMNET achieved the highest recall (83-86%) and AUCs of 0.84-0.89. Longitudinal mixed-effects modeling revealed eight candidates associated with progression across diagnostic stages. We identified a three-part molecular framework characterizing neurodegeneration: a diagnostic subpanel reflecting early microbiome dysregulation (secretory granins and metabolites) and synaptic breakdown; a second subpanel monitoring phenoconversion via neurogenesis precursors and extracellular matrix proteins; and a third subpanel tracking progression through chronic neuroinflammation and immune activation. This integrated multi-omics approach provides a robust framework for stage-specific PD monitoring and potential clinical deployment.

Journal Article↗