PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Prediction of functional class of the SARS coronavirus proteins by a statistical learning method.

The complete genome of severe acute respiratory syndrome coronavirus (SARS-CoV) reveals the existence of putative proteins unique to SARS-CoV. Identification of their function facilitates a mechanistic understanding of SARS infection and drug development for its treatment. The sequence of the majority of these putative proteins has no significant similarity to those of known proteins, which complicates the task of using sequence analysis tools to probe their function. Support vector machines (SVM), useful for predicting the functional class of distantly related proteins, is employed to ascribe a possible functional class to SARS-CoV proteins. Testing results indicate that SVM is able to predict the functional class of 73% of the known SARS-CoV proteins with available sequences and 67% of 18 other novel viral proteins. A combination of the sequence comparison method BLAST and SVMProt can further improve the prediction accuracy of SMVProt such that the functional class of two additional SARS-CoV proteins is correctly predicted. Our study suggests that the SARS-CoV genome possibly contains a putative voltage-gated ion channel, structural proteins, a carbon-oxygen lyase, oxidoreductases acting on the CH-OH group of donors, and an ATP-binding cassette transporter. A web version of our software, SVMProt, is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi .

Adenosine Triphosphate↗

How to interpret an anonymous bacterial genome: machine learning approach to gene identification.

In this report we address the problem of accurate statistical modeling of DNA sequences, either coding or noncoding, for a bacterial species whose genome (or a large portion) was sequenced but not yet characterized experimentally. Availability of these models is critical for successful solution of the genome annotation task by statistical methods of gene finding. We present the method, GeneMark-Genesis, which learns the parameters of Markov models of protein-coding and noncoding regions from anonymous bacterial genomic sequence. These models are subsequently used in the GeneMark and GeneMark.hmm gene-finding programs. Although there is basically one model of a noncoding region for a given genome, several models of protein-coding region are automatically obtained by GeneMark-Genesis. The diversity of protein-coding models reflects the diversity of oligonucleotide compositions, particularly the diversity of codon usage strategies observed in genes from one and the same genome. In the simplest and the most important case, there are just two gene models-typical and atypical ones. We show that the atypical model allows one to predict genes that escape identification by the typical model. Many genes predicted by the atypical model appear to be horizontally transferred genes. The early versions of GeneMark-Genesis were used for annotating the genomes of Methanoccocus jannaschii and Helicobacter pylori. We report the results of accuracy testing of the full-scale version of GeneMark-Genesis on 10 completely sequenced bacterial genomes. Interestingly, the GeneMark.hmm program that employed the typical and atypical models defined by GeneMark-Genesis was able to predict 683 new atypical genes with 176 of them confirmed by similarity search.

Algorithms↗

Two-sample comparison based on prediction error, with applications to candidate gene association studies.

To take advantage of the increasingly available high-density SNP maps across the genome, various tests that compare multilocus genotypes or estimated haplotypes between cases and controls have been developed for candidate gene association studies. Here we view this two-sample testing problem from the perspective of supervised machine learning and propose a new association test. The approach adopts the flexible and easy-to-understand classification tree model as the learning machine, and uses the estimated prediction error of the resulting prediction rule as the test statistic. This procedure not only provides an association test but also generates a prediction rule that can be useful in understanding the mechanisms underlying complex disease. Under the set-up of a haplotype-based transmission/disequilibrium test (TDT) type of analysis, we find through simulation studies that the proposed procedure has the correct type I error rates and is robust to population stratification. The power of the proposed procedure is sensitive to the chosen prediction error estimator. Among commonly used prediction error estimators, the .632+ estimator results in a test that has the best overall performance. We also find that the test using the .632+ estimator is more powerful than the standard single-point TDT analysis, the Pearson's goodness-of-fit test based on estimated haplotype frequencies, and two haplotype-based global tests implemented in the genetic analysis package FBAT. To illustrate the application of the proposed method in population-based association studies, we use the procedure to study the association between non-Hodgkin lymphoma and the IL10 gene.

Adult↗

Prediction of cytochrome P450 3A4, 2D6, and 2C9 inhibitors and substrates by using support vector machines.

Statistical learning methods have been used in developing filters for predicting inhibitors of two P450 isoenzymes, CYP3A4 and CYP2D6. This work explores the use of different statistical learning methods for predicting inhibitors of these enzymes and an additional P450 enzyme, CYP2C9, and the substrates of the three P450 isoenzymes. Two consensus support vector machine (CSVM) methods, "positive majority" (PM-CSVM) and "positive probability" (PP-CSVM), were used in this work. These methods were first tested for the prediction of inhibitors of CYP3A4 and CYP2D6 by using a significantly higher number of inhibitors and noninhibitors than that used in earlier studies. They were then applied to the prediction of inhibitors of CYP2C9 and substrates of the three enzymes. Both methods predict inhibitors of CYP3A4 and CYP2D6 at a similar level of accuracy as those of earlier studies. For classification of inhibitors of CYP2C9, the best CSVM method gives an accuracy of 88.9% for inhibitors and 96.3% for noninhibitors. The accuracies for classification of substrates and nonsubstrates of CYP3A4, CYP2D6, and CYP2C9 are 98.2 and 90.9%, 96.6 and 94.4%, and 85.7 and 98.8%, respectively. Both CSVM methods are potentially useful as filters for predicting inhibitors and substrates of P450 isoenzymes. These methods generally give better accuracies than single SVM classification systems, and the performance of the PP-CSVM method is slightly better than that of the PM-CSVM method.

Algorithms↗

Modeling obesity using abductive networks.

This paper investigates the use of abductive-network machine learning for modeling and predicting outcome parameters in terms of input parameters in medical survey data. Here we consider modeling obesity as represented by the waist-to-hip ratio (WHR) risk factor to investigate the influence of various parameters. The same approach would be useful in predicting values of clinical parameters that are difficult or expensive to measure from others that are more readily available. The AIM abductive network machine learning tool was used to model the WHR from 13 other health parameters. Survey data were collected for a randomly selected sample of 1100 persons aged 20 yr and over attending nine primary health care centers at Al-Khobar, Saudi Arabia. Models were synthesized by training on a randomly selected set of 800 cases, using both continuous and categorical representations of the parameters, and evaluated by predicting the WHR value for the remaining 300 cases. Models for WHR as a continuous variable predict the actual values within an error of 7.5% at the 90% confidence limits. Categorical models predict the correct logical value of WHR with an error in only 2 of the 300 evaluation cases. Analytical relationships derived from simple categorical models explain global observations on the total survey population to an accuracy as high as 99%. Simple continuous models represented as analytical functions highlight global relationships and trends. Results confirm the strong correlation between WHR and diastolic blood pressure, cholesterol level, and family history of obesity. Compared to other statistical and neural network approaches, AIM abductive networks provide faster and more automated model synthesis. A review is given of other areas where the proposed modeling approach can be useful in clinical practice.

Adult↗

Protein beta-turn prediction using nearest-neighbor method.

MOTIVATION: With the emerging success of protein secondary structure prediction through the applications of various statistical and machine learning techniques, similar techniques have been applied to protein beta-turn prediction. In this study, we perform protein beta-turn prediction using a k-nearest neighbor method, which is combined with a filter that uses predicted protein secondary structure information. Traditional beta-turn prediction from k-nearest neighbor method is modified to account for the unbalanced ratio of the natural occurrence of beta-turns and non-beta-turns. RESULTS: Our prediction scheme is tested on a set of 426 non-homologous protein sequences. The prediction scheme consists of two stages: k-nearest neighbor method stage and filtering stage. Variations of the k-nearest neighbor method were used to take property of beta-turns into consideration. Our filtering method uses beta-turn/non-beta-turn estimates from the k-nearest neighbor method stage and predicted protein secondary structure information from PSI-PRED in order to get new beta-turn/non-beta-turn estimate. Our result is compared with the previously best known beta-turn prediction method on the dataset of 426 non-homologous protein sequences and is shown to give slightly superior performance at significantly lower computational complexity. AVAILABILITY: Contact the author for information on the source code of the programs used.

Algorithms↗

Bio-medical entity extraction using support vector machines.

OBJECTIVE: Support vector machines (SVMs) have achieved state-of-the-art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named entity task to special domains with large-scale terminologies such as those in medicine and related disciplines. METHODS AND MATERIALS: The foundation for the model is a sample of text annotated by a domain expert according to an ontology of concepts, properties and relations. The model then learns to annotate unseen terms in new texts and contexts. The results can be used for a variety of intelligent language processing applications. We illustrate SVMs capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. RESULTS: Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance. CONCLUSION: Our experiments indicate a relationship between feature window size and the amount of training data and that a combination of surface words, orthographic features and head noun features achieve the best performance among the feature sets tested.

Algorithms↗

Optimal amnesic probabilistic automata or how to learn and classify proteins in linear time and space.

Statistical modeling of sequences is a central paradigm of machine learning that finds multiple uses in computational molecular biology and many other domains. The probabilistic automata typically built in these contexts are subtended by uniform, fixed-memory Markov models. In practice, such automata tend to be unnecessarily bulky and computationally imposing both during their synthesis and use. Recently, D. Ron, Y. Singer, and N. Tishby built much more compact, tree-shaped variants of probabilistic automata under the assumption of an underlying Markov process of variable memory length. These variants, called Probabilistic Suffix Trees (PSTs) were subsequently adapted by G. Bejerano and G. Yona and applied successfully to learning and prediction of protein families. The process of learning the automaton from a given training set S of sequences requires theta(Ln2) worst-case time, where n is the total length of the sequences in S and L is the length of a longest substring of S to be considered for a candidate state in the automaton. Once the automaton is built, predicting the likelihood of a query sequence of m characters may cost time theta(m2) in the worst case. The main contribution of this paper is to introduce automata equivalent to PSTs but having the following properties: Learning the automaton, for any L, takes O (n) time. Prediction of a string of m symbols by the automaton takes O (m) time. Along the way, the paper presents an evolving learning scheme and addresses notions of empirical probability and related efficient computation, which is a by-product possibly of more general interest.

Algorithms↗

Artificial Intelligence Technologies in Nursing Clinical Decision-Making: An Umbrella Review.

AIM: To describe contemporary peer-reviewed literature on artificial intelligence in nurses' clinical decision-making. METHODS: An umbrella review of literature reviews. DATA SOURCES: Four major databases were searched for reviews published between 2019 and 2024. RESULTS: Sixteen literature reviews reported on 965 nursing artificial intelligence primary studies. The studies focused on technology development and emerging performance evaluations, whilst real-world testing or implementation in nursing clinical settings was rare. Rigorous comparative analyses were lacking. While artificial intelligence demonstrates promise in decision-making, challenges such as a lack of controlled studies, algorithmic bias, limited reproducibility and insufficient clinical trials hinder its practical impact. Ethical concerns, transparency and patient data privacy issues pose barriers to AI integration in nursing practice. Ethical and legal guidelines for patient privacy are needed and should be taught along with AI literacy training for nurses. CONCLUSIONS: Artificial intelligence has the potential to enhance clinical nursing decision-making, although evidence is limited by too few examples of nurse participation during development. Underutilisation in administrative nursing functions hinders implementation. Nurses should assume a central role in the design and development of AI applications to ensure that these technologies address the realities of nursing practice. With such improvements, artificial intelligence can transform nursing practice, improve nurses' clinical decision-making and ultimately enhance consumer healthcare outcomes. PATIENT OR PUBLIC INVOLVEMENT: No Patient or Public Involvement. REPORTING METHOD: While there is no reporting checklist for umbrella reviews, the PRISMA guide for systematic reviews was followed.

Artificial Intelligence↗

Automated discovery of structural signatures of protein fold and function.

There are constraints on a protein sequence/structure for it to adopt a particular fold. These constraints could be either a local signature involving particular sequences or arrangements of secondary structure or a global signature involving features along the entire chain. To search systematically for protein fold signatures, we have explored the use of Inductive Logic Programming (ILP). ILP is a machine learning technique which derives rules from observation and encoded principles. The derived rules are readily interpreted in terms of concepts used by experts. For 20 populated folds in SCOP, 59 rules were found automatically. The accuracy of these rules, which is defined as the number of true positive plus true negative over the total number of examples, is 74% (cross-validated value). Further analysis was carried out for 23 signatures covering 30% or more positive examples of a particular fold. The work showed that signatures of protein folds exist, about half of rules discovered automatically coincide with the level of fold in the SCOP classification. Other signatures correspond to homologous family and may be the consequence of a functional requirement. Examination of the rules shows that many correspond to established principles published in specific literature. However, in general, the list of signatures is not part of standard biological databases of protein patterns. We find that the length of the loops makes an important contribution to the signatures, suggesting that this is an important determinant of the identity of protein folds. With the expansion in the number of determined protein structures, stimulated by structural genomics initiatives, there will be an increased need for automated methods to extract principles of protein folding from coordinates.

Algorithms↗

GenSo-FDSS: a neural-fuzzy decision support system for pediatric ALL cancer subtype identification using gene expression data.

OBJECTIVE: Acute lymphoblastic leukemia (ALL) is the most common malignancy of childhood, representing nearly one third of all pediatric cancers. Currently, the treatment of pediatric ALL is centered on tailoring the intensity of the therapy applied to a patient's risk of relapse, which is linked to the type of leukemia the patient has. Hence, accurate and correct diagnosis of the various leukemia subtypes becomes an important first step in the treatment process. Recently, gene expression profiling using DNA microarrays has been shown to be a viable and accurate diagnostic tool to identify the known prognostically important ALL subtypes. Thus, there is currently a huge interest in developing autonomous classification systems for cancer diagnosis using gene expression data. This is to achieve an unbiased analysis of the data and also partly to handle the large amount of genetic information extracted from the DNA microarrays. METHODOLOGY: Generally, existing medical decision support systems (DSS) for cancer classification and diagnosis are based on traditional statistical methods such as Bayesian decision theory and machine learning models such as neural networks (NN) and support vector machine (SVM). Though high accuracies have been reported for these systems, they fall short on certain critical areas. These included (a) being able to present the extracted knowledge and explain the computed solutions to the users; (b) having a logical deduction process that is similar and intuitive to the human reasoning process; and (c) flexible enough to incorporate new knowledge without running the risk of eroding old but valid information. On the other hand, a neural fuzzy system, which is synthesized to emulate the human ability to learn and reason in the presence of imprecise and incomplete information, has the ability to overcome the above-mentioned shortcomings. However, existing neural fuzzy systems have their own limitations when used in the design and implementation of DSS. Hence, this paper proposed the use of a novel neural fuzzy system: the generic self-organising fuzzy neural network (GenSoFNN) with truth-value restriction (TVR) fuzzy inference, as a fuzzy DSS (denoted as GenSo-FDSS) for the classification of ALL subtypes using gene expression data. RESULTS AND CONCLUSION: The performance of the GenSo-FDSS system is encouraging when benchmarked against those of NN, SVM and the K-nearest neighbor (K-NN) classifier. On average, a classification rate of above 90% has been achieved using the GenSo-FDSS system.

Algorithms↗

Predicting the toxicity of complex mixtures using artificial neural networks.

Industrial and municipal wastewaters constitute major sources of contamination of the aquatic compartment and represent a threat to aquatic life. Artificial neural networks based on three different learning paradigms were studied as a means of predicting acute toxicity to trout (5 days exposure to wastewaters) using input data from two simple microbiotests requiring only 5 or 15 min of incubation. These microbiotests were 1) the chemoluminescent peroxidase (Cl-Per) assay, which can detect radical scavengers and enzyme-inhibiting substances, and 2) the luminescent bacteria toxicity test (Microtox), in which reduction of light emission by bacteria during exposure is taken as a measure of toxicity. The responses obtained with the trout bioassay, the Cl-Per and the Microtox test were analyzed through statistical correlation (Pearson product-moment correlation), unsupervised learning by a self-organizing network, and assisted learning by the backpropagation and the Boltzmann machine (probabilistic) paradigms. No significant correlation (p < 0.05) was found between the responses obtained with either the Cl-Per assay (p = 0.121) or the Microtox (p = 0.061) microbiotest and those resulting from the trout bioassay. The self-organizing network was able to identify by itself a maximum of five classes that were more or less relevant for predicting toxicity to fish: class 1 contained 2 samples that were toxic to fish, class 2 contained 2/3 samples that were toxic, class 3 showed 6/8 samples that were non toxic, class 4 contained 5/6 samples that were non-toxic and class 5 comprised one sample that was toxic. Supervised learning with backpropagation analysis yielded two kinds of networks that hold potential. The first one was able to predict the actual toxic wastewater concentration with an overall performance of 65% when fed fresh data, while the second one, which was designed to differentiate between toxic and non-toxic effluents, exhibited a much better performance (90%). However, the probabilistic network also proved to be a very good predictive model for toxicity to fish, with an overall performance of 90%. Although more data are needed, the network based on the backpropagation paradigm seems to be a better predictor or classifier of trout toxicity when used with the Cl-Per and the Microtox microbiotests.

Algorithms↗

The future of pediatric vesicoureteral reflux management.

BACKGROUND AND OBJECTIVE: Vesicoureteral reflux (VUR) is a common condition in pediatric urology, yet important uncertainties persist regarding risk stratification, imaging strategies, and prevention of long-term renal damage. Emerging technologies may help address these challenges. This review provides a forward-looking overview of recent advances in artificial intelligence (AI) and immunomodulation that may influence future management of pediatric VUR. METHODS: A forward-looking literature review was performed using the PubMed database (January 2000-March 2025), focusing on studies addressing AI, immunomodulation, or vaccination in the context of VUR and urinary tract infections. Criteria of inclusion were the relevance to pediatric VUR, the novelty of the proposed concept, the potential clinical implications and, for the AI literature, the existence of a clinical evaluation of the algorithm on a dataset from patients. KEY FINDINGS AND LIMITATIONS: AI-based models show promising performance in supporting clinical decision-making, including prediction of the need for voiding cystourethrography, automated grading of VUR, estimation of recurrent urinary tract infection risk and prediction of chemoprophylaxis. These tools may facilitate more individualized diagnostic and therapeutic strategies, although current evidence is largely retrospective and requires prospective validation. Immunization and immunomodulatory approaches aim to reduce infection burden and modulate inflammatory pathways associated with renal scarring. While early experimental and adult clinical data are encouraging, pediatric-specific evidence remains limited, and clinical applicability in children with VUR is not yet established. CONCLUSION: Artificial intelligence and immunologically targeted strategies represent complementary, emerging approaches that may contribute to more personalized management of pediatric VUR. At present, both should be regarded as exploratory tools whose clinical impact will depend on further validation and appropriately designed pediatric studies.

Humans↗

Application of latent semantic analysis to protein remote homology detection.

MOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise.

Algorithms↗

Urinary nucleosides as potential tumor markers evaluated by learning vector quantization.

Modified nucleosides were recently presented as potential tumor markers for breast cancer. The patterns of the levels of urinary nucleosides are different for tumor bearing individuals and for healthy individuals. Thus, a powerful pattern recognition method is needed. Although backpropagation (BP) neural networks are becoming increasingly common in medical literature for pattern recognition, it has been shown that often-superior methods exist like learning vector quantization (LVQ) and support vector machines (SVM). The aim of this feasibility study is to get an indication of the performance of urinary nucleoside levels evaluated by LVQ in contrast to the evaluation the popular BP and SVM networks. Urine samples were collected from female breast cancer patients and from healthy females. Twelve different ribonucleosides were isolated and quantified by a high performance liquid chromatography (HPLC) procedure. LVQ, SVM and BP networks were trained and the performance was evaluated by the classification of the test sets into the categories "cancer" and "healthy". All methods showed a good classification with a sensitivity ranging from 58.8 to 70.6% at a specificity of 88.4-94.2% for the test patterns. Although the classification performance of all methods is comparable, the LVQ implementations are superior in terms of more qualitative features: the results of LVQ networks are more reproducible, as the initialization is deterministic. The LVQ networks can be trained by unbalanced sizes of the different classes. LVQ networks are fast during training, need only few parameters adjusted for training and can be retrained by patterns of "local individuals". As at least some of these features play an important role in an implementation into a medical decision support system, it is recommended to use LVQ for an extended study.

Adult↗

Machine learning can improve prediction of severity in acute pancreatitis using admission values of APACHE II score and C-reactive protein.

BACKGROUND: Acute pancreatitis (AP) has a variable course. Accurate early prediction of severity is essential to direct clinical care. Current assessment tools are inaccurate, and unable to adapt to new parameters. None of the current systems uses C-reactive protein (CRP). Modern machine-learning tools can address these issues. METHODS: 370 patients admitted with AP in a 5-year period were retrospectively assessed; after exclusions, 265 patients were studied. First recorded values for physical examination and blood tests, aetiology, severity and complications were recorded. A kernel logistic regression model was used to remove redundant features, and identify the relationships between relevant features and outcome. Bootstrapping was used to make the best use of data and obtain confidence estimates on the parameters of the model. RESULTS: A model containing 8 variables (age, CRP, respiratory rate, pO2 on air, arterial pH, serum creatinine, white cell count and GCS) predicted a severe attack with an area under the receiver-operating characteristic curve (AUC) of 0.82 (SD 0.01). The optimum cut-off value for predicting severity gave sensitivity and specificity of 0.87 and 0.71 respectively. The predictions were significantly better (p = 0.0036) than admission APACHE II scores in the same patients (AUC 0.74) and better than historical admission APACHE II data (AUC 0.68-0.75). CONCLUSIONS: This system for the first time combines admission values of selected components of APACHE II and CRP for prediction of severe AP. The score is simple to use, and is more accurate than admission APACHE II alone. It is adaptable and would allow incorporation of new predictive factors.

APACHE↗

Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolites.

MOTIVATION: Our knowledge of metabolism is far from complete, and the gaps in our knowledge are being revealed by metabolomic detection of small-molecules not previously known to exist in cells. An important challenge is to determine the reactions in which these compounds participate, which can lead to the identification of gene products responsible for novel metabolic pathways. To address this challenge, we investigate how machine learning can be used to predict potential substrates and products of oxidoreductase-catalyzed reactions. RESULTS: We examined 1956 oxidation/reduction reactions in the KEGG database. The vast majority of these reactions (1626) can be divided into 12 subclasses, each of which is marked by a particular type of functional group transformation. For a given transformation, the local structures of reaction centers in substrates and products can be characterized by patterns. These patterns are not unique to reactants but are widely distributed among KEGG metabolites. To distinguish reactants from non-reactants, we trained classifiers (linear-kernel Support Vector Machines) using negative and positive examples. The input to a classifier is a set of atomic features that can be determined from the 2D chemical structure of a compound. Depending on the subclass of reaction, the accuracy of prediction for positives (negatives) is 64 to 93% (44 to 92%) when asking if a compound is a substrate and 71 to 98% (50 to 92%) when asking if a compound is a product. Sensitivity analysis reveals that this performance is robust to variations of the training data. Our results suggest that metabolic connectivity can be predicted with reasonable accuracy from the presence or absence of local structural motifs in compounds and their readily calculated atomic features. AVAILABILITY: Classifiers reported here can be used freely for noncommercial purposes via a Java program available upon request.

Algorithms↗

Prediction of RNA-binding proteins from primary sequence by a support vector machine approach.

Elucidation of the interaction of proteins with different molecules is of significance in the understanding of cellular processes. Computational methods have been developed for the prediction of protein-protein interactions. But insufficient attention has been paid to the prediction of protein-RNA interactions, which play central roles in regulating gene expression and certain RNA-mediated enzymatic processes. This work explored the use of a machine learning method, support vector machines (SVM), for the prediction of RNA-binding proteins directly from their primary sequence. Based on the knowledge of known RNA-binding and non-RNA-binding proteins, an SVM system was trained to recognize RNA-binding proteins. A total of 4011 RNA-binding and 9781 non-RNA-binding proteins was used to train and test the SVM classification system, and an independent set of 447 RNA-binding and 4881 non-RNA-binding proteins was used to evaluate the classification accuracy. Testing results using this independent evaluation set show a prediction accuracy of 94.1%, 79.3%, and 94.1% for rRNA-, mRNA-, and tRNA-binding proteins, and 98.7%, 96.5%, and 99.9% for non-rRNA-, non-mRNA-, and non-tRNA-binding proteins, respectively. The SVM classification system was further tested on a small class of snRNA-binding proteins with only 60 available sequences. The prediction accuracy is 40.0% and 99.9% for snRNA-binding and non-snRNA-binding proteins, indicating a need for a sufficient number of proteins to train SVM. The SVM classification systems trained in this work were added to our Web-based protein functional classification software SVMProt, at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi. Our study suggests the potential of SVM as a useful tool for facilitating the prediction of protein-RNA interactions.

Algorithms↗