PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Multiclass cancer classification using gene expression profiling and probabilistic neural networks.

Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 or more. In addition, microarray data exhibit a high degree of noise. Most of the discussed methods do not adequately address the problem of dimensionality and noise. Furthermore, although machine learning and data mining methods are based on statistics, most such techniques do not address the biologist's requirement for sound mathematical confidence measures. Finally, most machine learning and data mining classification methods fail to incorporate misclassification costs, i.e. they are indifferent to the costs associated with false positive and false negative classifications. In this paper, we present a probabilistic neural network (PNN) model that addresses all these issues. The PNN model provides sound statistical confidences for its decisions, and it is able to model asymmetrical misclassification costs. Furthermore, we demonstrate the performance of the PNN for multiclass gene expression data sets. Here, we compare the performance of the PNN with two machine learning methods, a decision tree and a neural network. To assess and evaluate the performance of the classifiers, we use a lift-based scoring system that allows a fair comparison of different models. The PNN clearly outperformed the other models. The results demonstrate the successful application of the PNN model for multiclass cancer classification.

Artificial Intelligence↗

[Fish-based assessment methods for the ecological status of aquatic systems].

A short overview of fish-based assessment methods for aquatic systems is presented. Multimetric indices, as, e.g., the index of biotic integrity (IBI), firstly developed in USA and later adapted for European river basins and other countries, are shortly described. Non-multimetric indices are also discussed, e.g. the ichthyological index (II) and the index of ecological status of fish communities (ISECI), both proposed for monitoring Italian rivers. Moreover, statistical and predicting methods based on machine learning techniques are described. Finally, a new approach for developing standardised fish-based methods useful to assess ecological status of Italian rivers is proposed. Although rather complex, the use of bony fish in biomonitoring is promising and requires a multidisciplinary approach to be adopted.

Animals↗

Conserved HSFA1-dependent chromatin dynamics drive heat stress responses in plants.

Eukaryotic organisms remodel chromatin landscapes to regulate gene expression in response to environmental stress. In plants, heat stress (HS) induces widespread chromatin changes, yet the role of heat shock transcription factors (HSFs) in chromatin remodeling and their evolutionary conservation remains unclear. Using Marchantia polymorpha Mphsf mutants and Arabidopsis thaliana Athsfa1s mutants, we identify HSFA1 as a key regulator of HS-induced cis-regulatory element (CRE) accessibility, a mechanism conserved across land plants, mice, and humans. Gene regulatory network modeling reveals parallel transcription factor subnetworks, with MpWRKY10 and MpABI5B acting as indirect and negative HS regulators. We further showed that ABA modulates gene expression in an HSFA1-dependent manner without inducing chromatin remodeling. Finally, we develop a machine learning framework integrating chromatin accessibility and CRE information to predict gene expression across species, revealing stress-responsive regulatory logic at the transcriptional level. These findings provide insights into how TFs coordinate chromatin architecture to drive stress adaptation.

Heat-Shock Response↗

Rapid assessment of clinical severity for salmonellosis cases via protein family domain analysis and machine learning.

Salmonella is a common pathogen, infecting more than a million people yearly. Rapid assessment of clinical case severity is essential for improving patient outcomes and optimizing healthcare resources. Advancements in genome sequencing technologies have enabled the analysis of bacterial genomes from many clinical cases, opening up new opportunities for precise and timely diagnosis. This study proposes a genome-based framework for identifying critical Salmonella cases before the onset of critical symptoms and facilitating early medical intervention. By leveraging protein family (Pfam) domains as the representation for genomic data, the complex genetic profiles of Salmonella cases are simplified into interpretable features. The severity levels of cases were investigated through rigorous data analysis, resulting in a set of 70 Pfam domains that could be potentially used as biomarkers. Machine Learning was employed to assess the predictive power of the curated Pfam biomarkers, achieving high accuracy (~93%) in sorting cases into critical, moderate, and mild categories. The results demonstrate the efficacy of the proposed approach. This framework highlights the potential of using bacterial genomic data in clinical decision-making, opening the window for timely personalized interventions for Salmonella infection management.

Domains of unknown function (DUFs)↗

A neural-network based method for prediction of gamma-turns in proteins from multiple sequence alignment.

In the present study, an attempt has been made to develop a method for predicting gamma-turns in proteins. First, we have implemented the commonly used statistical and machine-learning techniques in the field of protein structure prediction, for the prediction of gamma-turns. All the methods have been trained and tested on a set of 320 nonhomologous protein chains by a fivefold cross-validation technique. It has been observed that the performance of all methods is very poor, having a Matthew's Correlation Coefficient (MCC) </= 0.06. Second, predicted secondary structure obtained from PSIPRED is used in gamma-turn prediction. It has been found that machine-learning methods outperform statistical methods and achieve an MCC of 0.11 when secondary structure information is used. The performance of gamma-turn prediction is further improved when multiple sequence alignment is used as the input instead of a single sequence. Based on this study, we have developed a method, GammaPred, for gamma-turn prediction (MCC = 0.17). The GammaPred is a neural-network-based method, which predicts gamma-turns in two steps. In the first step, a sequence-to-structure network is used to predict the gamma-turns from multiple alignment of protein sequence. In the second step, it uses a structure-to-structure network in which input consists of predicted gamma-turns obtained from the first step and predicted secondary structure obtained from PSIPRED.

Databases, Protein↗

Analyzing tumor gene expression profiles.

A brief introduction to high throughput technologies for measuring and analyzing gene expression is given. Various supervised and unsupervised data mining methods for analyzing the produced high-dimensional data are discussed. The main emphasis is on supervised machine learning methods for classification and prediction of tumor gene expression profiles. Furthermore, methods to rank the genes according to their importance for the classification are explored. The approaches are illustrated by exploratory studies using two examples of retrospective clinical data from routine tests; diagnostic prediction of small round blue cell tumors (SRBCT) of childhood and determining the estrogen receptor (ER) status of sporadic breast cancer. The classification performance is gauged using blind tests. These studies demonstrate the feasibility of machine learning-based molecular cancer classification.

Adult↗

Splicing-site recognition of rice (Oryza sativa L.) DNA sequences by support vector machines.

MOTIVATION: It was found that high accuracy splicing-site recognition of rice (Oryza sativa L.) DNA sequence is especially difficult. We described a new method for the splicing-site recognition of rice DNA sequences. METHOD: Based on the intron in eukaryotic organisms conforming to the principle of GT-AG, we used support vector machines (SVM) to predict the splicing sites. By machine learning, we built a model and used it to test the effect of the test data set of true and pseudo splicing sites. RESULTS: The prediction accuracy we obtained was 87.53% at the true 5' end splicing site and 87.37% at the true 3' end splicing sites. The results suggested that the SVM approach could achieve higher accuracy than the previous approaches.

Algorithms↗

Multi-criteria decision making and its application to in silico discovery of vaccine candidates for Toxoplasma gondii.

Vaccine discovery against eukaryotic parasites is not trivial and few exist. Reverse vaccinology is an in silico vaccine discovery approach, designed to identify vaccine candidates from the thousands of protein sequences encoded by a target genome. Previously, we produced the Vacceed bioinformatics pipeline for identification of parasite membrane and excreted/secreted proteins that were likely be exposed to the hosts immune system. More recently, we improved upon machine learning as the final decision-making process to identify parasite proteins that induce a protective response in an animal model. Subsequently, we combined Vacceed with metrics on B and T cell epitope types to produce a new in silico discovery workflow. In this study we extend this in silico workflow to the developability of proteins as vaccines by the incorporation of metrics on the physicochemical properties of proteins. To demonstrate this process, every Toxoplasma gondii protein was ranked in its capacity to provide exposure to the immune system (Vacceed exposure score), presence of epitopes and solubility characteristics by several multicriteria decision making (MCDM) tools (such as TOPSIS, VIKOR and MABAC). A consensus rank was subsequently generated from the results of these tools using a variety of aggregate ranking methods. Levels of uncertainty in the aggregate protein rankings was assessed by conformal interval prediction in association with a machine learning model. Several of the top ranked proteins identified by this approach were novel, uncharacterized membrane transporters or proteins associated with RNA metabolism. In conclusion, MCDM automated the decision making using well known algorithms while conformal prediction intervals varied significantly across the 8000+ proteins of T. gondii. Highly ranked proteins (e.g. the top 100) typically generated low prediction intervals, providing high levels of confidence in their ranks.

Toxoplasma↗

Support vector machines for predicting protein structural class.

BACKGROUND: We apply a new machine learning method, the so-called Support Vector Machine method, to predict the protein structural class. Support Vector Machine method is performed based on the database derived from SCOP, in which protein domains are classified based on known structures and the evolutionary relationships and the principles that govern their 3-D structure. RESULTS: High rates of both self-consistency and jackknife tests are obtained. The good results indicate that the structural class of a protein is considerably correlated with its amino acid composition. CONCLUSIONS: It is expected that the Support Vector Machine method and the elegant component-coupled method, also named as the covariant discrimination algorithm, if complemented with each other, can provide a powerful computational tool for predicting the structural classes of proteins.

Algorithms↗

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq↗

Predicting genetic regulatory response using classification.

MOTIVATION: Studying gene regulatory mechanisms in simple model organisms through analysis of high-throughput genomic data has emerged as a central problem in computational biology. Most approaches in the literature have focused either on finding a few strong regulatory patterns or on learning descriptive models from training data. However, these approaches are not yet adequate for making accurate predictions about which genes will be up- or down-regulated in new or held-out experiments. By introducing a predictive methodology for this problem, we can use powerful tools from machine learning and assess the statistical significance of our predictions. RESULTS: We present a novel classification-based method for learning to predict gene regulatory response. Our approach is motivated by the hypothesis that in simple organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular experiment based on (1) the presence of binding site subsequences ('motifs') in the gene's regulatory region and (2) the expression levels of regulators such as transcription factors in the experiment ('parents'). Thus, our learning task integrates two qualitatively different data sources: genome-wide cDNA microarray data across multiple perturbation and mutant experiments along with motif profile data from regulatory sequences. We convert the regression task of predicting real-valued gene expression measurements to a classification task of predicting +1 and -1 labels, corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. The learning algorithm employed is boosting with a margin-based generalization of decision trees, alternating decision trees. This large-margin classifier is sufficiently flexible to allow complex logical functions, yet sufficiently simple to give insight into the combinatorial mechanisms of gene regulation. We observe encouraging prediction accuracy on experiments based on the Gasch S.cerevisiae dataset, and we show that we can accurately predict up- and down-regulation on held-out experiments. We also show how to extract significant regulators, motifs and motif-regulator pairs from the learned models for various stress responses. Our method thus provides predictive hypotheses, suggests biological experiments, and provides interpretable insight into the structure of genetic regulatory networks. AVAILABILITY: The MLJava package is available upon request to the authors. Supplementary: Additional results are available from http://www.cs.columbia.edu/compbio/geneclass

Binding Sites↗

New approaches in molecular structure prediction.

In the past years, much effort has been put on the development of new methodologies and algorithms for the prediction of protein secondary and tertiary structures from (sequence) data; this is reviewed in detail. New approaches for these predictions such as neural network methods, genetic algorithms, machine learning, and graph theoretical methods are discussed. Secondary structure prediction algorithms were improved mostly by considering families of related proteins; however, for the reliable tertiary structure modeling of proteins, knowledge-based techniques are still preferred. Methods and examples with more or less successful results are described. Also, programs and parameterizations for energy minimisations, molecular dynamics, and electrostatic interactions have been improved, especially with respect to their former limits of applicability. Other topics discussed in this review include the use of traditional and on-line databases, the docking problem and surface properties of biomolecules, packing of protein cores, de novo design and protein engineering, prediction of membrane protein structures, the verification and reliability of model structures, and progress made with currently available software and computer hardware. In summary, the prediction of the structure, function, and other properties of a protein is still possible only within limits, but these limits continue to be moved.

Chemical Phenomena↗

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n&#x2009;=&#x2009;83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans↗

Application of machine learning in SNP discovery.

BACKGROUND: Single nucleotide polymorphisms (SNP) constitute more than 90% of the genetic variation, and hence can account for most trait differences among individuals in a given species. Polymorphism detection software PolyBayes and PolyPhred give high false positive SNP predictions even with stringent parameter values. We developed a machine learning (ML) method to augment PolyBayes to improve its prediction accuracy. ML methods have also been successfully applied to other bioinformatics problems in predicting genes, promoters, transcription factor binding sites and protein structures. RESULTS: The ML program C4.5 was applied to a set of features in order to build a SNP classifier from training data based on human expert decisions (True/False). The training data were 27,275 candidate SNP generated by sequencing 1973 STS (sequence tag sites) (12 Mb) in both directions from 6 diverse homozygous soybean cultivars and PolyBayes analysis. Test data of 18,390 candidate SNP were generated similarly from 1359 additional STS (8 Mb). SNP from both sets were classified by experts. After training the ML classifier, it agreed with the experts on 97.3% of test data compared with 7.8% agreement between PolyBayes and experts. The PolyBayes positive predictive values (PPV) (i.e., fraction of candidate SNP being real) were 7.8% for all predictions and 16.7% for those with 100% posterior probability of being real. Using ML improved the PPV to 84.8%, a 5- to 10-fold increase. While both ML and PolyBayes produced a similar number of true positives, the ML program generated only 249 false positives as compared to 16,955 for PolyBayes. The complexity of the soybean genome may have contributed to high false SNP predictions by PolyBayes and hence results may differ for other genomes. CONCLUSION: A machine learning (ML) method was developed as a supplementary feature to the polymorphism detection software for improving prediction accuracies. The results from this study indicate that a trained ML classifier can significantly reduce human intervention and in this case achieved a 5-10 fold enhanced productivity. The optimized feature set and ML framework can also be applied to all polymorphism discovery software. ML support software is written in Perl and can be easily integrated into an existing SNP discovery pipeline.

Algorithms↗

Automatic discovery of cross-family sequence features associated with protein function.

BACKGROUND: Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterized protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. RESULTS: We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. CONCLUSION: We have developed a novel and useful approach for knowledge discovery in annotated sequence data. The technique is able to identify functionally important sequence features and does not require expert knowledge. By viewing protein function from a sequence perspective, the approach is also suitable for discovering unexpected links between biological processes, such as the recently discovered role of ubiquitination in transcription.

Algorithms↗

Diagnosing breast cancer based on support vector machines.

The Support Vector Machine (SVM) classification algorithm, recently developed from the machine learning community, was used to diagnose breast cancer. At the same time, the SVM was compared to several machine learning techniques currently used in this field. The classification task involves predicting the state of diseases, using data obtained from the UCI machine learning repository. SVM outperformed k-means cluster and two artificial neural networks on the whole. It can be concluded that nine samples could be mislabeled from the comparison of several machine learning techniques.

Algorithms↗

Data-driven consideration of genetic disorders for global genomic newborn screening programs.

PURPOSE: Over 30 international studies are exploring newborn sequencing (NBSeq) to expand the range of genetic disorders included in newborn screening. Substantial variability in gene selection across programs exists, highlighting the need for a systematic approach to prioritize genes. METHODS: We assembled a data set comprising 25 characteristics about each of the 4390 genes included in 27 NBSeq programs. We used regression analysis to identify several predictors of inclusion and developed a machine learning model to rank genes for public health consideration. RESULTS: Among 27 NBSeq programs, the number of genes analyzed ranged from 134 to 4299, with only 74 (1.7%) genes included by over 80% of programs. The most significant associations with gene inclusion across programs were presence on the US Recommended Uniform Screening Panel (inclusion increase of 74.7%, CI: 71.0%-78.4%), robust evidence on the natural history (29.5%, CI: 24.6%-34.4%), and treatment efficacy (17.0%, CI: 12.3%-21.7%) of the associated genetic disease. A boosted trees machine learning model using 13 predictors achieved high accuracy in predicting gene inclusion across programs (area under the curve = 0.915, R2 = 84%). CONCLUSION: The machine learning model developed here provides a ranked list of genes that can adapt to emerging evidence and regional needs, enabling more consistent and informed gene selection in NBSeq initiatives.

Humans↗

Using feature generation and feature selection for accurate prediction of translation initiation sites.

Correct prediction of the translation initiation site (TIS) is an important issue in genomic research. We show that feature generation together with correlation based feature selection can be used with a variety of machine learning algorithms to give highly accurate translation initiation site prediction. Only very few features are needed and the results achieve comparable accuracy to the best existing approaches. Our approach has the advantage that it does not require one to devise a special prediction method; rather standard machine learning classifiers are shown to give very good performance on the selected features. The raw and generated features which we have found to be important are the following: positions -3 and -1 in the sequence; upstream k-grams for k=3, 4, and 5; stop-codon frequency; downstream in-frame 3-gram; and the distance of ATG to the beginning of the sequence. The best result, with an overall accuracy of 90%, is obtained by selecting only seven features from this set. The same features retrained with the use of a scanning model achieves an overall accuracy of 94% on this dataset.

Codon, Initiator↗