PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Knowledge discovery in biomedical databases: a machine induction approach.

The increase in the number and size of available databases by far exceeds the growth of the corresponding knowledge. Furthermore, many databases contain information which is not possessed by an existing human expert. This creates both a need and an opportunity for extracting knowledge from databases. An unsolved problem in molecular biology is the problem of predicting a protein's secondary structure from its primary structure. Inductive machine learning is a search for a plausible general description which can explain the given input data, and is useful for predicting new data. In this paper we present a statistical inductive algorithm which can be used to produce new rules for predicting multiple protein secondary structures from protein primary structure databases.

Algorithms↗

Automated epiluminescence microscopy--tissue counter analysis using CART and 1-NN in the diagnosis of Melanoma.

BACKGROUND/PURPOSE: In tissue counter analysis, digital images are overlayed with regularly distributed measuring masks (elements) of equal size and shape, and the digital contents (grey level, colour and texture parameters) of each element are used for statistical analysis. In this study we assessed the applicability of tissue counter analysis and machine learning algorithms on tumour segmentation and diagnostic discrimination of benign and malignant melanocytic skin lesions. METHODS: A total of 369 standardised dermatoscopic images (93 melanomas, 276 benign nevi) were evaluated. The Classification and Regression Tree (CART) analysis was performed in order to differentiate between melanocytic skin lesions and surrounding skin. Instance-based learning (1-NN) was tested for differentiating between benign and malignant tumour elements. For diagnostic assessment, only the percentage of elements suggestive for malignancy in each lesion was used. RESULTS: Evaluation of a total of 369 melanocytic skin lesions showed a suitable segmentation of the tumour portion in 97.6%. When instance-based learning was applied to an independent test set, a threshold value of 27.4% of elements suggestive for malignancy recognised 35 out of 35 melanomas and 100 out of 101 nevi (sensitivity 100%, specificity 99%, positive predictive value 97.2%, negative predictive value 100%). CONCLUSION: Tissue counter analysis combined with machine learning algorithms turned out to be a useful method for diagnostic purposes in epiluminescence microscopy.

Algorithms↗

Using MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles.

Biomedical abbreviations and acronyms are widely used in biomedical literature. Since many of them represent important content in biomedical literature, information retrieval and extraction benefits from identifying the meanings of those terms. On the other hand, many abbreviations and acronyms are ambiguous, it would be important to map them to their full forms, which ultimately represent the meanings of the abbreviations. In this study, we present a semi-supervised method that applies MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles. We first automatically generated from the MEDLINE abstracts a dictionary of abbreviation-full pairs based on a rule-based system that maps abbreviations to full forms when full forms are defined in the abstracts. We then trained on the MEDLINE abstracts and predicted the full forms of abbreviations in full-text journal articles by applying supervised machine-learning algorithms in a semi-supervised fashion. We report up to 92% prediction precision and up to 91% coverage.

Artificial Intelligence↗

Why neural networks should not be used for HIV-1 protease cleavage site prediction.

UNLABELLED: Several papers have been published where nonlinear machine learning algorithms, e.g. artificial neural networks, support vector machines and decision trees, have been used to model the specificity of the HIV-1 protease and extract specificity rules. We show that the dataset used in these studies is linearly separable and that it is a misuse of nonlinear classifiers to apply them to this problem. The best solution on this dataset is achieved using a linear classifier like the simple perceptron or the linear support vector machine, and it is straightforward to extract rules from these linear models. We identify key residues in peptides that are efficiently cleaved by the HIV-1 protease and list the most prominent rules, relating them to experimental results for the HIV-1 protease. MOTIVATION: Understanding HIV-1 protease specificity is important when designing HIV inhibitors and several different machine learning algorithms have been applied to the problem. However, little progress has been made in understanding the specificity because nonlinear and overly complex models have been used. RESULTS: We show that the problem is much easier than what has previously been reported and that linear classifiers like the simple perceptron or linear support vector machines are at least as good predictors as nonlinear algorithms. We also show how sets of specificity rules can be generated from the resulting linear classifiers. AVAILABILITY: The datasets used are available at http://www.hh.se/staff/bioinf/

Algorithms↗

Proteomic mass spectra classification using decision tree based ensemble methods.

MOTIVATION: Modern mass spectrometry allows the determination of proteomic fingerprints of body fluids like serum, saliva or urine. These measurements can be used in many medical applications in order to diagnose the current state or predict the evolution of a disease. Recent developments in machine learning allow one to exploit such datasets, characterized by small numbers of very high-dimensional samples. RESULTS: We propose a systematic approach based on decision tree ensemble methods, which is used to automatically determine proteomic biomarkers and predictive models. The approach is validated on two datasets of surface-enhanced laser desorption/ionization time of flight measurements, for the diagnosis of rheumatoid arthritis and inflammatory bowel diseases. The results suggest that the methodology can handle a broad class of similar problems.

Algorithms↗

A multimarker model to predict outcome in tamoxifen-treated breast cancer patients.

PURPOSE: This study was designed to produce a model to predict outcome in tamoxifen-treated breast cancer patients based on clinicopathologic features and multiple molecular markers. EXPERIMENTAL DESIGN: This was a retrospective study of 324 stage I to III female breast cancer patients treated with tamoxifen for whom standard clinicopathologic data and tumor tissue microarrays were available. Nine molecular markers were studied by semiquantitative immunohistochemistry and/or fluorescence in situ hybridization. Cox proportional hazards analysis was used to determine the contributions of each variable to disease-specific and overall survival, and machine learning was used to produce a model to predict patient outcome. RESULTS: On a univariate basis, the following features were significantly associated with worse survival: high pathologic tumor or nodal class, histologic grade, epidermal growth factor receptor, ERBB2, MYC, or TP53; absent estrogen receptor (ER) or progesterone receptor; and low BCL2. CCND1 and CDKN1B did not reach statistical significance. On a multivariate basis, nodal class, ER, and MYC were statistically significant as independent factors for survival. However, the benefit of ER-positive status was moderated by BCL2, ERBB2, and progesterone receptor. BCL2 and TP53 also interacted as an independent risk factor. A kernel partial least squares polynomial model was developed with an area under the receiver operating characteristic curve of 0.90. CONCLUSIONS: Our data show the predictive value of BCL2, ERBB2, MYC, and TP53 in addition to the standard hormone receptors and clinicopathologic features, and they show the importance of conditional interpretation of certain molecular markers. Our multimarker predictive model performed significantly better than standard guidelines.

Aged↗

DNA methylation and machine learning: challenges and perspective toward enhanced clinical diagnostics.

DNA methylation is an epigenetic modification that regulates gene expression by adding methyl groups to DNA, affecting cellular function and disease development. Machine learning, a subset of artificial intelligence, analyzes large datasets to identify patterns and make predictions. Over the past two decades, advances in bioinformatics technologies for arrays and sequencing have generated vast amounts of data, leading to the widespread adoption of machine learning methods for analyzing complex biological information for medical problems. This review explores recent advancements in DNA methylation studies that leverage emerging machine learning techniques for more precise, comprehensive, and rapid patient diagnostics based on DNA methylation markers. We present a general workflow for researchers, from clinical research questions to result interpretation and monitoring. Additionally, we showcase successful examples in diagnosing cancer, neurodevelopmental disorders, and multifactorial diseases. Some of these studies have led to the development of diagnostic platforms that have entered the global healthcare market, highlighting the promising future of this field.

Humans↗

Protein disorder prediction by condensed PSSM considering propensity for order or disorder.

BACKGROUND: More and more disordered regions have been discovered in protein sequences, and many of them are found to be functionally significant. Previous studies reveal that disordered regions of a protein can be predicted by its primary structure, the amino acid sequence. One observation that has been widely accepted is that ordered regions usually have compositional bias toward hydrophobic amino acids, and disordered regions are toward charged amino acids. Recent studies further show that employing evolutionary information such as position specific scoring matrices (PSSMs) improves the prediction accuracy of protein disorder. As more and more machine learning techniques have been introduced to protein disorder detection, extracting more useful features with biological insights attracts more attention. RESULTS: This paper first studies the effect of a condensed position specific scoring matrix with respect to physicochemical properties (PSSMP) on the prediction accuracy, where the PSSMP is derived by merging several amino acid columns of a PSSM belonging to a certain property into a single column. Next, we decompose each conventional physicochemical property of amino acids into two disjoint groups which have a propensity for order and disorder respectively, and show by experiments that some of the new properties perform better than their parent properties in predicting protein disorder. In order to get an effective and compact feature set on this problem, we propose a hybrid feature selection method that inherits the efficiency of uni-variant analysis and the effectiveness of the stepwise feature selection that explores combinations of multiple features. The experimental results show that the selected feature set improves the performance of a classifier built with Radial Basis Function Networks (RBFN) in comparison with the feature set constructed with PSSMs or PSSMPs that adopt simply the conventional physicochemical properties. CONCLUSION: Distinguishing disordered regions from ordered regions in protein sequences facilitates the exploration of protein structures and functions. Results based on independent testing data reveal that the proposed predicting model DisPSSMP performs the best among several of the existing packages doing similar tasks, without either under-predicting or over-predicting the disordered regions. Furthermore, the selected properties are demonstrated to be useful in finding discriminating patterns for order/disorder classification.

Amino Acid Sequence↗

Prediction-based fingerprints of protein-protein interactions.

The recognition of protein interaction sites is an important intermediate step toward identification of functionally relevant residues and understanding protein function, facilitating experimental efforts in that regard. Toward that goal, the authors propose a novel representation for the recognition of protein-protein interaction sites that integrates enhanced relative solvent accessibility (RSA) predictions with high resolution structural data. An observation that RSA predictions are biased toward the level of surface exposure consistent with protein complexes led the authors to investigate the difference between the predicted and actual (i.e., observed in an unbound structure) RSA of an amino acid residue as a fingerprint of interaction sites. The authors demonstrate that RSA prediction-based fingerprints of protein interactions significantly improve the discrimination between interacting and noninteracting sites, compared with evolutionary conservation, physicochemical characteristics, structure-derived and other features considered before. On the basis of these observations, the authors developed a new method for the prediction of protein-protein interaction sites, using machine learning approaches to combine the most informative features into the final predictor. For training and validation, the authors used several large sets of protein complexes and derived from them nonredundant representative chains, with interaction sites mapped from multiple complexes. Alternative machine learning techniques are used, including Support Vector Machines and Neural Networks, so as to evaluate the relative effects of the choice of a representation and a specific learning algorithm. The effects of induced fit and uncertainty of the negative (noninteracting) class assignment are also evaluated. Several representative methods from the literature are reimplemented to enable direct comparison of the results. Using rigorous validation protocols, the authors estimated that the new method yields the overall classification accuracy of about 74% and Matthews correlation coefficients of 0.42, as opposed to up to 70% classification accuracy and up to 0.3 Matthews correlation coefficient for methods that do not utilize RSA prediction-based fingerprints. The new method is available at http://sppider.cchmc.org.

Artificial Intelligence↗

Dietary Polyphenol Acteoside-Related Molecular Signatures in Clear Cell Renal Cell Carcinoma: Multi-Omics Profiling and Functional Validation of IMPDH1.

Clear cell renal cell carcinoma (ccRCC) is characterized by substantial metabolic and molecular heterogeneity, but the disease-relevant programs associated with acteoside, a dietary polyphenol, remain poorly understood. We integrated predicted acteoside targets with bulk, single-cell, and spatial transcriptomic data from ccRCC and combined molecular subtyping with cross-cohort machine-learning analysis. Acteoside-related signatures were preferentially enriched in malignant compartments and increased with tumor grade and stage. Consensus clustering identified two molecular subtypes with distinct biological and clinical features. C1 was associated with immune activation, metabolic activity, and more favorable survival, whereas C2 showed greater genomic instability, reduced renal epithelial differentiation, and poorer outcomes. We further benchmarked multiple machine-learning strategies and established a 10-gene prognostic model that retained predictive performance across independent cohorts, with IMPDH1 emerging as the strongest risk-associated feature. Functional experiments confirmed the biological relevance of IMPDH1: its knockdown suppressed ccRCC cell proliferation, DNA synthesis, colony formation, and migration, whereas overexpression produced the opposite effects. Together, these findings indicate that acteoside-related molecular signatures capture clinically relevant heterogeneity in ccRCC and provide a framework for linking dietary-polyphenol-related molecular space with tumor biology. The identification and functional validation of IMPDH1 further highlight its potential importance in ccRCC progression.

IMPDH1↗

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning↗

Serum Proteomic Profiling Implicates a Dysregulated Neurohormonal-Inflammatory Axis in Post-Fontan Sinus Tachycardia.

BACKGROUND: Postoperative sinus tachycardia is a poorly understood complication following the Fontan procedure. The molecular signaling cascades triggering acute tachycardia remain uncharacterized, limiting therapeutic innovation. Here, we present a retrospective study leveraging serum proteomics and machine learning to identify the molecular drivers of postoperative Fontan sinus tachycardia. METHODS: We integrated a clinically relevant ovine Fontan model with continuous telemetric heart rate monitoring and human patient data. Serum proteomics coupled with least absolute shrinkage and selection operator and Boruta machine learning algorithms were used to identify protein panels predictive of postoperative sinus tachycardia. Cross-species validation was performed by comparing proteomic signatures from sheep and pediatric patients undergoing Glenn or Fontan surgery. RESULTS: Ovine Fontan animals demonstrated significant heart rate elevation beginning on postoperative day 1, peaking at postoperative day 3 (159.4±11.7 bpm versus preoperative, 105.3±10.5 bpm; P=0.0002), before trending toward baseline by postoperative day 10. This pattern was mirrored in human patients with a more modest magnitude. Surgical controls did not exhibit tachycardia. The principal component most correlated with heart rate (principal component 1: r=0.78, P=2.2×10-4) was enriched for inflammatory and neural pathways. The Boruta algorithm identified an 11-protein panel with strong predictive power (area under the receiver operating characteristic curve, 0.963). Cross-species comparison demonstrated that angiotensinogen, angiotensin-converting enzyme, and pentraxin 3 were similarly dysregulated in both species postoperatively. CONCLUSIONS: This study provides molecular evidence implicating a dysregulated neurohormonal-inflammatory axis in acute postoperative Fontan sinus tachycardia and establishes a foundation for developing targeted diagnostics and therapeutics for this complication.

Animals↗

Comparison of genetic algorithms and other classification methods in the diagnosis of female urinary incontinence.

Galactica, a newly developed machine-learning system that utilizes a genetic algorithm for learning, was compared with discriminant analysis, logistic regression, k-means cluster analysis, a C4.5 decision-tree generator and a random bit climber hill-climbing algorithm. The methods were evaluated in the diagnosis of female urinary incontinence in terms of prediction accuracy of classifiers, on the basis of patient data. The best methods were discriminant analysis, logistic regression, C4.5 and Galactica. Practically no statistically significant differences existed between the prediction accuracy of these classification methods. We consider that machine-learning systems C4.5 and Galactica are preferable for automatic construction of medical decision aids, because they can cope with missing data values directly and can present a classifier in a comprehensible form. Galactica performed nearly as well as C4.5. The results are in agreement with the results of earlier research, indicating that genetic algorithms are a competitive method for constructing classifiers from medical data.

Algorithms↗

MULTIPREVENT: Integrated screening for smoking-related multimorbidity using low-dose chest computed tomography.

OBJECTIVES: Tobacco consumption, combined with individual genetic predispositions, contributes to an age-dependent risk not only for lung cancer but also for other non-communicable diseases (NCDs) such as cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), osteoporosis, and diabetes. The MULTIPREVENT project aims to validate whether low-dose computed tomography (LDCT) of the chest, combined with simple biomarkers, functional tests, and genomic profiling, can serve as an effective tool for comprehensive health assessment and risk prediction of multimorbidity in adults. STUDY DESIGN: The study is based on a prospective epidemiological design involving 3000 participants from the MOLTEST-BIS lung cancer screening cohort (2016-2018). These participants, aged 50-79 years (during MOLTEST-BIS) and with a smoking history of at least 30 pack-years, will undergo two follow-up assessments in 2025-2027 and 2030-2032. METHODS: Each follow-up includes LDCT, spirometry, standardized blood pressure measurement, anthropometric evaluation, biomarker assessment (lipid profile, lipoprotein(a), glycated haemoglobin), and health-related questionnaires. Genetic profiling will be performed using the Illumina Infinium Global Screening Arrays approach to identify inherited predispositions to major NCDs. All data, clinical, imaging (including radiomics), molecular, and genetic, will be integrated through machine learning algorithms to develop AI-based risk prediction models. RESULTS: The MULTIPREVENT study is expected to generate a wide range of scientific, clinical, and infrastructural results that will serve as a foundation for future public health initiatives in integrated prevention. CONCLUSIONS: By linking imaging and biochemical markers, genetic susceptibility, and clinical parameters within a longitudinal design, MULTIPREVENT will establish data-driven, AI-supported prevention strategies aimed at reducing morbidity and mortality among adults exposed to tobacco. The project will also serve as a model for population-based multimorbidity prevention programs.

Humans↗

CRNPRED: highly accurate prediction of one-dimensional protein structures by large-scale critical random networks.

BACKGROUND: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. RESULTS: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. CONCLUSION: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.

Algorithms↗

Identification of Biomarkers for Right Ventricular Dysfunction in Idiopathic Dilated Cardiomyopathy Via Urinary Proteomics and Machine Learning.

BACKGROUND: Right ventricular dysfunction (RVD) is a common complication of idiopathic dilated cardiomyopathy linked to poor outcomes. However, reliable noninvasive biomarkers for RVD remain lacking. This study aimed to identify urinary proteomic markers using mass spectrometry and machine learning. METHODS: In this prospective cohort, patients with idiopathic dilated cardiomyopathy were classified by cardiac magnetic resonance imaging into groups with RVD (RV ejection fraction <45%) and without RVD groups. Baseline urine samples were profiled by data-independent acquisition mass spectrometry. Differentially expressed proteins were identified and selected by least absolute shrinkage and selection operator regression to build a diagnostic model, developed in a training set, and validated in a test set. The primary end point was a composite of cardiovascular death, heart failure rehospitalization, left ventricular assist device implantation, or heart transplantation. RESULTS: The study enrolled 147 patients with idiopathic dilated cardiomyopathy (64 with RVD, 83 without), with a median follow-up of 19.3&#x2009;months. Of 3579 quantified urinary proteins, 46 were differentially expressed between groups. A 3-protein panel (RARRES1 [retinoic acid receptor responder protein 1], MVB12B [multivesicular body subunit 12B], GSK3A [glycogen synthase kinase 3 alpha]) was identified and showed excellent diagnostic accuracy (training area under the curve 0.946; validation area under the curve0.935), outperforming both NT-proBNP (N-terminal pro-brain natriuretic peptide) and tricuspid annular plane systolic excursion. The risk score derived from this panel effectively stratified patients, with the high-risk group exhibiting significantly worse outcomes than the low-risk group (hazard ratio, 3.24 [95% CI, 1.56-6.71], P=0.002). CONCLUSIONS: The urinary proteomic panel developed in this study demonstrates diagnostic and prognostic potential for identifying RVD in idiopathic dilated cardiomyopathy, providing a promising noninvasive tool for precise detection and clinical risk stratification.

Humans↗

Metagenomic polymorphic toxin effector and immunity profiling predicts microbiome development and disease-related dysbiosis.

Bacteria use antagonistic interbacterial weapons, such as polymorphic toxin secretion systems (TSS), to compete for niches in the human gut microbiome. We hypothesized that TSS influence gut microbiome development and disease-related dysbiosis. We developed a bioinformatic marker gene approach (PolyProf) to quantify TSS including ~200 effector and immunity genes and applied it to ~15,000 publicly available human metagenomes. PolyProf alpha and beta diversity readily distinguished 12 different human disease states and enabled the construction of highly accurate linear regression classifier machine learning models. Elastic net machine learning models integrating bacterial taxonomy with PolyProf had strong predictive value for 12 disease states, outperforming models utilizing taxonomy alone. During microbiome development in the first year of life, PolyProf alpha diversity increases, and beta diversity becomes increasingly like the maternal microbiome, influenced by vertical transfer, delivery mode, and breastfeeding. PolyProf is related to strain sharing among adults through social interactions. In summary, TSS genes strongly correlate with microbiome development and interpersonal strain sharing, suggesting roles for interbacterial antagonism. Since PolyProf distinguishes diverse adult disease statuses, these dynamics may contribute to non-genetic inheritance.IMPORTANCEPrevious research has demonstrated that bacteria compete within the gut microbiome using toxin secretion systems (TSS). How TSS contribute to human microbiome development and the microbiome alterations observed in human diseases is not known. This study develops a new bioinformatic tool for profiling TSS-related genes in metagenomic data. Application of this approach to large-scale human fecal metagenomic data demonstrates the dynamic association of TSS during microbiome development, including the exchange of strains among social contacts. TSS gene abundance patterns are highly predictive of 12 disease states. This study advances the field by enabling TSS profiling in metagenomes and by identifying disease and microbiome development biomarkers that provide hypotheses for future mechanistic studies and may be useful for disease diagnosis.

Dysbiosis↗

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗