PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Computational identification of residues that modulate voltage sensitivity of voltage-gated potassium channels.

BACKGROUND: Studies of the structure-function relationship in proteins for which no 3D structure is available are often based on inspection of multiple sequence alignments. Many functionally important residues of proteins can be identified because they are conserved during evolution. However, residues that vary can also be critically important if their variation is responsible for diversity of protein function and improved phenotypes. If too few sequences are studied, the support for hypotheses on the role of a given residue will be weak, but analysis of large multiple alignments is too complex for simple inspection. When a large body of sequence and functional data are available for a protein family, mature data mining tools, such as machine learning, can be applied to extract information more easily, sensitively and reliably. We have undertaken such an analysis of voltage-gated potassium channels, a transmembrane protein family whose members play indispensable roles in electrically excitable cells. RESULTS: We applied different learning algorithms, combined in various implementations, to obtain a model that predicts the half activation voltage of a voltage-gated potassium channel based on its amino acid sequence. The best result was obtained with a k-nearest neighbor classifier combined with a wrapper algorithm for feature selection, producing a mean absolute error of prediction of 7.0 mV. The predictor was validated by permutation test and evaluation of independent experimental data. Feature selection identified a number of residues that are predicted to be involved in the voltage sensitive conformation changes; these residues are good target candidates for mutagenesis analysis. CONCLUSION: Machine learning analysis can identify new testable hypotheses about the structure/function relationship in the voltage-gated potassium channel family. This approach should be applicable to any protein family if the number of training examples and the sequence diversity of the training set that are necessary for robust prediction are empirically validated. The predictor and datasets can be found at the VKCDB web site.

Algorithms↗

Discrimination of outer membrane proteins using machine learning algorithms.

Discriminating outer membrane proteins (OMPs) from other folding types of globular and membrane proteins is an important task both for identifying OMPs from genomic sequences and for the successful prediction of their secondary and tertiary structures. In this work, we have analyzed the performance of different methods, based on Bayes rules, logistic functions, neural networks, support vector machines, decision trees, etc. for discriminating OMPs. We found that most of the machine learning techniques discriminate OMPs with similar accuracy. The neural network-based method could discriminate the OMPs from other proteins [globular/transmembrane helical (TMH)] at the fivefold cross-validation accuracy of 91.0% in a dataset of 1,088 proteins. The accuracy of discriminating globular proteins is 88.8% and that of TMH proteins is 93.7%. Further, the neural network method is tested with globular proteins belonging to 30 different folding types and it could successfully exclude 95% of the considered proteins. The proteins with SAM domain such as knottins, rubredoxin, and thioredoxin folds are eliminated with 100% accuracy. These accuracy levels are comparable to or better than other methods in the literature. We suggest that this method could be effectively used to discriminate OMPs and for detecting OMPs in genomic sequences.

Algorithms↗

Oncogenic EME1 promotes tumor progression and immune modulation in human cancers with therapeutic targeting potential.

BACKGROUND: EME1, a critical DNA repair endonuclease, has emerged as a potential oncogene implicated in genome instability and cancer progression. However, its pan-cancer roles, prognostic significance, immune interactions, and therapeutic targeting remain underexplored. METHODS: We conducted a comprehensive pan-cancer analysis integrating multi-omics data from public databases, including TIMER2.0, GEPIA2, TISIDB, and cBioPortal, to evaluate EME1 expression, genetic alterations, and their association with clinical outcomes, immune infiltration, and molecular pathways. Virtual screening of 3180 FDA-approved drugs and molecular dynamics (MD) simulations were employed to identify and validate potential EME1 inhibitors. RESULTS: EME1 was significantly overexpressed in various human cancers and positively associated with advanced tumor grade and stage. High EME1 expression and mutations were linked to poor overall and disease-free survival. Immunogenomic profiling revealed strong positive correlations between EME1 and myeloid-derived suppressor cells (MDSCs), alongside a negative association with endothelial cell function, suggesting immunosuppressive roles. Machine learning models based on EME1-associated genes demonstrated high predictive accuracy for liver hepatocellular carcinoma (AUC > 0.90). Virtual screening identified eight promising drug candidates, including Everolimus and Dioscin, with strong binding affinities. MD simulations confirmed the stability of these interactions, particularly for Dioscin. CONCLUSION: This study reveals the multifaceted oncogenic roles of EME1 in tumor progression, immune evasion, and prognosis. It proposes EME1 as a promising biomarker and therapeutic target across multiple cancer types. The identified drug candidates warrant further in vitro and in vivo validation for potential repurposing in EME1-targeted cancer therapy.

EME1↗

Assessment of hepatotoxic liabilities by transcript profiling.

Male Wistar rats were treated with various model compounds or the appropriate vehicle controls in order to create a reference database for toxicogenomics assessment of novel compounds. Hepatotoxic compounds in the database were either known hepatotoxicants or showed hepatotoxicity during preclinical testing. Histopathology and clinical chemistry data were used to anchor the transcript profiles to an established endpoint (steatosis, cholestasis, direct acting, peroxisomal proliferation or nontoxic/control). These reference data were analyzed using a supervised learning method (support vector machines, SVM) to generate classification rules. This predictive model was subsequently used to assess compounds with regard to a potential hepatotoxic liability. A steatotic and a non-hepatotoxic 5HT(6) receptor antagonist compound from the same series were successfully discriminated by this toxicogenomics model. Additionally, an example is shown where a hepatotoxic liability was correctly recognized in the absence of pathological findings. In vitro experiments and a dog study confirmed the correctness of the toxicogenomics alert. Another interesting observation was that transcript profiles indicate toxicologically relevant changes at an earlier timepoint than routinely used methods. Together, these results support the useful application of toxicogenomics in raising alerts for adverse effects and generating mechanistic hypotheses that can be followed up by confirmatory experiments.

Animals↗

Diagnosing anorexia based on partial least squares, back propagation neural network, and support vector machines.

Support vector machine (SVM), as a novel type of learning machine, for the first time, was used to develop a predictive model for early diagnosis of anorexia. It was based on the concentration of six elements (Zn, Fe, Mg, Cu, Ca, and Mn) and the age extracted from 90 cases. Compared with the results obtained from two other classifiers, partial least squares (PLS) and back-propagation neural network (BPNN), the SVM method exhibited the best whole performance. The accuracies for the test set by PLS, BPNN, and SVM methods were 52%, 65%, and 87%, respectively. Moreover, the models we proposed could also provide some insight into what factors were related to anorexia.

Anorexia↗

Discovery and performance of DNA methylation panels for cancer detection and classification in blood.

Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.

DNA methylation↗

Recent advances in computational prediction of drug absorption and permeability in drug discovery.

Approximately 40%-60% of developing drugs failed during the clinical trials because of ADME/Tox deficiencies. Virtual screening should not be restricted to optimize binding affinity and improve selectivity; and the pharmacokinetic properties should also be included as important filters in virtual screening. Here, the current development in theoretical models to predict drug absorption-related properties, such as intestinal absorption, Caco-2 permeability, and blood-brain partitioning are reviewed. The important physicochemical properties used in the prediction of drug absorption, and the relevance of predictive models in the evaluation of passive drug absorption are discussed. Recent developments in the prediction of drug absorption, especially with the application of new machine learning methods and newly developed software are also discussed. Future directions for research are outlined.

Computer Simulation↗

Prediction of Saccharomyces cerevisiae protein functional class from functional domain composition.

MOTIVATION: A key goal of genomics is to assign function to genes, especially for orphan sequences. RESULTS: We compared the clustered functional domains in the SBASE database to each protein sequence using BLASTP. This representation for a protein is a vector, where each of the non-zero entries in the vector indicates a significant match between the sequence of interest and the SBASE domain. The machine learning methods nearest neighbour algorithm (NNA) and support vector machines are used for predicting protein functional classes from this information. We find that the best results are found using the SBASE-A database and the NNA, namely 72% accuracy for 79% coverage. We tested an assigning function based on searching for InterPro sequence motifs and by taking the most significant BLAST match within the dataset. We applied the functional domain composition method to predict the functional class of 2018 currently unclassified yeast open reading frames. AVAILABILITY: A program for the prediction method, that uses NNA called Functional Class Prediction based on Functional Domains (FCPFD) is available and can be obtained by contacting Y.D.Cai at y.cai@umist.ac.uk

Algorithms↗

Developing a machine learning-based prognosis and immunotherapeutic response signature in colorectal cancer: insights from ferroptosis, fatty acid dynamics, and the tumor microenvironment.

INSTRUCTION: Colorectal cancer (CRC) poses a challenge to public health and is characterized by a high incidence rate. This study explored the relationship between ferroptosis and fatty acid metabolism in the tumor microenvironment (TME) of patients with CRC to identify how these interactions impact the prognosis and effectiveness of immunotherapy, focusing on patient outcomes and the potential for predicting treatment response. METHODS: Using datasets from multiple cohorts, including The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO), we conducted an in-depth multi-omics study to uncover the relationship between ferroptosis regulators and fatty acid metabolism in CRC. Through unsupervised clustering, we discovered unique patterns that link ferroptosis and fatty acid metabolism, and further investigated them in the context of immune cell infiltration and pathway analysis. We developed the FeFAMscore, a prognostic model created using a combination of machine learning algorithms, and assessed its predictive power for patient outcomes and responsiveness to treatment. The FeFAMscore signature expression level was confirmed using RT-PCR, and ACAA2 progression in cancer was further verified. RESULTS: This study revealed significant correlations between ferroptosis regulators and fatty acid metabolism-related genes with respect to tumor progression. Three distinct patient clusters with varied prognoses and immune cell infiltration were identified. The FeFAMscore demonstrated superior prognostic accuracy over existing models, with a C-index of 0.689 in the training cohort and values ranging from 0.648 to 0.720 in four independent validation cohorts. It also responses to immunotherapy and chemotherapy, indicating a sensitive response of special therapies (e.g., anti-PD-1, anti-CTLA4, osimertinib) in high FeFAMscore patients. CONCLUSION: Ferroptosis regulators and fatty acid metabolism-related genes not only enhance immune activation, but also contribute to immune escape. Thus, the FeFAMscore, a novel prognostic tool, is promising for predicting both the prognosis and efficacy of immunotherapeutic strategies in patients with CRC.

Ferroptosis↗

Local irritation/corrosion testing strategies: extending a decision support system by applying self-learning classifiers.

Procedures have been established and tested for the extension of a decision support system (DSS) for the prediction of the local irritation/corrosion potential of chemicals by using self-learning classifiers. The different approaches (decision trees, distances examinations in a multidimensional space, k-nearest-neighbour method) have been implemented, tested and evaluated independently. A combination of all of the established extension approaches was also developed and tested. Self-learning classifiers are constructed "automatically" by a computer, i.e. they are not derived by a human expert, and thus they can be constructed with minimal effort. The classifiers presented here extend the existing DSS in a manner that increased significantly the predictive power of the extended system. However, automatically calculated results of self-learning classifiers are produced by a machine, and a machine is incapable of explaining the toxicological relevance of the results obtained. Thus, these results must be accepted, despite an inability to prove their reliability. Only the mathematical correctness of the method and the prediction rates for suitable test cases can lend some credibility to predictions produced by a computer calculating on a self-learning basis. This may not be adequate for regulatory hazard assessment purposes.

Algorithms↗

Protein secondary structure prediction.

The past year has seen a consolidation of protein secondary structure prediction methods. The advantages of prediction from an aligned family of proteins have been highlighted by several accurate predictions made 'blind', before any X-ray or NMR structure was known for the family. New techniques that apply machine learning and discriminant analysis show promise as alternatives to neural networks.

Animals↗

Classification of prostatic carcinoma with artificial neural networks using comparative genomic hybridization and quantitative stereological data.

Staging of prostate cancer is a mainstay of treatment decisions and prognostication. In the present study, 50 pT2N0 and 28 pT3N0 prostatic adenocarcinomas were characterized by Gleason grading, comparative genomic hybridization (CGH), and histological texture analysis based on principles of stereology and stochastic geometry. The cases were classified by learning vector quantization and support vector machines. The quality of classification was tested by cross-validation. Correct prediction of stage from primary tumor data was possible with an accuracy of 74-80% from different data sets. The accuracy of prediction was similar when the Gleason score was used as input variable, when stereological data were used, or when a combination of CGH data and stereological data was used. The results of classification by learning vector quantization were slightly better than those by support vector machines. A method is briefly sketched by which training of neural networks can be adapted to unequal sample sizes per class. Progression from pT2 to pT3 prostate cancer is correlated with complex changes of the epithelial cells in terms of volume fraction, of surface area, and of second-order stereological properties. Genetically, this progression is accompanied by a significant global increase in losses and gains of DNA, and specifically by increased numerical aberrations on chromosome arms 1q, 7p, and 8p.

Adenocarcinoma↗

Identification of MHC Ligands Through Allele-Guided Isolation Combined With Machine Learning for Improved MHC Assignment Using ARDisplay-I.

The isolation of major histocompatibility complex (MHC) ligands and subsequent analysis by mass spectrometry is considered the gold standard for defining targets for T cell-based immunotherapies. However, as many targets of high tumor specificity are only presented at low abundance on the cell surface of tumor cells, the efficient isolation of these peptides is crucial for their successful detection. Here, we demonstrate how optimizing the MHC ligand isolation strategy, based on both the presenting MHC alleles and the individual peptide level, enhances the identification of specific MHC ligands. This ideally acknowledges not only the hydrophobicity but also the post-translational modifications of the respective MHC ligands. To further improve the identification and characterization of MHC ligands, we developed an MHC class I ligand prediction algorithm (ARDisplay-I) that outperforms current state-of-the-art tools when benchmarked against competitors such as netMHCpan 4.1, MixMHCpred, or MHCflurry. Implementing these strategies can augment the development of T cell receptor-based therapies by improving the identification of novel immunotherapy targets and enriching the resources available in the computational immunology field through a superior MHC presentation prediction algorithm.

Ligands↗

Support vector machines for prediction of protein signal sequences and their cleavage sites.

Given a nascent protein sequence, how can one predict its signal peptide or "Zipcode" sequence? This is an important problem for scientists to use signal peptides as a vehicle to find new drugs or to reprogram cells for gene therapy (see, e.g. K.C. Chou, Current Protein and Peptide Science 2002;3:615-22). In this paper, support vector machines (SVMs), a new machine learning method, is applied to approach this problem. The overall rate of correct prediction for 1939 secretary proteins and 1440 nonsecretary proteins was over 91%. It has not escaped our attention that the new method may also serve as a useful tool for further investigating many unclear details regarding the molecular mechanism of the ZIP code protein-sorting system in cells.

Genetic Vectors↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗

Beyond data and technology: the need for new thinking to enable the era of precision prevention.

BACKGROUND: Global flagship initiatives increasingly advocate for proactive health maintenance to alleviate the growing burden on reactive, disease-focused healthcare systems. Precision prevention is conceived as the targeted modulation of causal pathways across the disease continuum, from latent risk and pre-disease states to clinical manifestation, surpassing conventional public health prevention strategies that prioritise managing population-level risk factors. Traditional discovery and implementation models, however, remain poorly aligned with the pace and breadth of scientific and technological advances. This review outlines key barriers to scaling precision prevention and argues for the integration of conceptual, methodological, and policy perspectives into a single implementation&#x2011;oriented framework. MAIN: Individualised risk stratification lies at the core of precision prevention. Genomics serves as a stable substrate for lifetime susceptibility assessment, while meaningful prediction in multifactorial chronic disease requires additional risk monitoring using dynamic intermediate molecular markers and high-resolution exposomic data. Machine learning and other artificial intelligence (AI) methods are increasingly helpful tools for integrating large, heterogeneous and temporally structured real-world data to generate personalised predictions of health trajectories. Trustworthy AI-enabled risk prediction or decision-support systems are expected to provide transparency about model logic, assumptions and performance. In discovery, existing diagnostic classifications and conventional case-control designs can obscure mechanistic heterogeneity. Shifting toward precision phenotyping and biologically grounded disease redefinition could reveal a new layer of molecular understanding. Evidence generation strategies that reflect the temporal change of disease, including high&#x2011;risk enrichment, surrogate endpoints, and adaptive, trajectory-based monitoring, are particularly important for common conditions with prolonged latency periods (e.g., cancer, cardiovascular disease). Features often dismissed as "noise", such as stochastic molecular variation and minimal exposures, may in fact encode meaningful individual-level signals and thus merit investigation. CONCLUSION: To shift healthcare from reactive treatment toward proactive health maintenance requires coordinated action from stakeholders to reshape the pillars of discovery, reform outcome assessments and modernise implementation strategies.

Humans↗

transFold: a web server for predicting the structure and residue contacts of transmembrane beta-barrels.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins makes them an important protein class. At the present time, very few non-homologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane (TM) proteins. The transFold web server uses pairwise inter-strand residue statistical potentials derived from globular (non-outer-membrane) proteins to predict the supersecondary structure of TMB. Unlike all previous approaches, transFold does not use machine learning methods such as hidden Markov models or neural networks; instead, transFold employs multi-tape S-attribute grammars to describe all potential conformations, and then applies dynamic programming to determine the global minimum energy supersecondary structure. The transFold web server not only predicts secondary structure and TMB topology, but is the only method which additionally predicts the side-chain orientation of transmembrane beta-strand residues, inter-strand residue contacts and TM beta-strand inclination with respect to the membrane. The program transFold currently outperforms all other methods for accuracy of beta-barrel structure prediction. Available at http://bioinformatics.bc.edu/clotelab/transFold.

Amino Acids↗

A hybrid neural network system for prediction and recognition of promoter regions in human genome.

This paper proposes a high specificity and sensitivity algorithm called PromPredictor for recognizing promoter regions in the human genome. PromPredictor extracts compositional features and CpG islands information from genomic sequence, feeding these features as input for a hybrid neural network system (HNN) and then applies the HNN for prediction. It combines a novel promoter recognition model, coding theory, feature selection and dimensionality reduction with machine learning algorithm. Evaluation on Human chromosome 22 was approximately 66% in sensitivity and approximately 48% in specificity. Comparison with two other systems revealed that our method had superior sensitivity and specificity in predicting promoter regions. PromPredictor is written in MATLAB and requires Matlab to run. PromPredictor is freely available at http://www.whtelecom.com/Prompredictor.htm.

Computational Biology↗