PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Homology induction: the use of machine learning to improve sequence similarity searches.

BACKGROUND: The inference of homology between proteins is a key problem in molecular biology The current best approaches only identify approximately 50% of homologies (with a false positive rate set at 1/1000). RESULTS: We present Homology Induction (HI), a new approach to inferring homology. HI uses machine learning to bootstrap from standard sequence similarity search methods. First a standard method is run, then HI learns rules which are true for sequences of high similarity to the target (assumed homologues) and not true for general sequences, these rules are then used to discriminate sequences in the twilight zone. To learn the rules HI describes the sequences in a novel way based on a bioinformatic knowledge base, and the machine learning method of inductive logic programming. To evaluate HI we used the PDB40D benchmark which lists sequences of known homology but low sequence similarity. We compared the HI methodology with PSI-BLAST alone and found HI performed significantly better. In addition, Receiver Operating Characteristic (ROC) curve analysis showed that these improvements were robust for all reasonable error costs. The predictive homology rules learnt by HI by can be interpreted biologically to provide insight into conserved features of homologous protein families. CONCLUSIONS: HI is a new technique for the detection of remote protein homology--a central bioinformatic problem. HI with PSI-BLAST is shown to outperform PSI-BLAST for all error costs. It is expect that similar improvements would be obtained using HI with any sequence similarity method.

Algorithms↗

Predicting the effect of missense mutations on protein function: analysis with Bayesian networks.

BACKGROUND: A number of methods that use both protein structural and evolutionary information are available to predict the functional consequences of missense mutations. However, many of these methods break down if either one of the two types of data are missing. Furthermore, there is a lack of rigorous assessment of how important the different factors are to prediction. RESULTS: Here we use Bayesian networks to predict whether or not a missense mutation will affect the function of the protein. Bayesian networks provide a concise representation for inferring models from data, and are known to generalise well to new data. More importantly, they can handle the noisy, incomplete and uncertain nature of biological data. Our Bayesian network achieved comparable performance with previous machine learning methods. The predictive performance of learned model structures was no better than a naïve Bayes classifier. However, analysis of the posterior distribution of model structures allows biologically meaningful interpretation of relationships between the input variables. CONCLUSION: The ability of the Bayesian network to make predictions when only structural or evolutionary data was observed allowed us to conclude that structural information is a significantly better predictor of the functional consequences of a missense mutation than evolutionary information, for the dataset used. Analysis of the posterior distribution of model structures revealed that the top three strongest connections with the class node all involved structural nodes. With this in mind, we derived a simplified Bayesian network that used just these three structural descriptors, with comparable performance to that of an all node network.

Algorithms↗

Machine learning in sedimentation modelling.

The paper presents machine learning (ML) models that predict sedimentation in the harbour basin of the Port of Rotterdam. The important factors affecting the sedimentation process such as waves, wind, tides, surge, river discharge, etc. are studied, the corresponding time series data is analysed, missing values are estimated and the most important variables behind the process are chosen as the inputs. Two ML methods are used: MLP ANN and M5 model tree. The latter is a collection of piece-wise linear regression models, each being an expert for a particular region of the input space. The models are trained on the data collected during 1992-1998 and tested by the data of 1999-2000. The predictive accuracy of the models is found to be adequate for the potential use in the operational decision making.

Algorithms↗

Effect of molecular descriptor feature selection in support vector machine classification of pharmacokinetic and toxicological properties of chemical agents.

Statistical-learning methods have been developed for facilitating the prediction of pharmacokinetic and toxicological properties of chemical agents. These methods employ a variety of molecular descriptors to characterize structural and physicochemical properties of molecules. Some of these descriptors are specifically designed for the study of a particular type of properties or agents, and their use for other properties or agents might generate noise and affect the prediction accuracy of a statistical learning system. This work examines to what extent the reduction of this noise can improve the prediction accuracy of a statistical learning system. A feature selection method, recursive feature elimination (RFE), is used to automatically select molecular descriptors for support vector machines (SVM) prediction of P-glycoprotein substrates (P-gp), human intestinal absorption of molecules (HIA), and agents that cause torsades de pointes (TdP), a rare but serious side effect. RFE significantly reduces the number of descriptors for each of these properties thereby increasing the computational speed for their classification. The SVM prediction accuracies of P-gp and HIA are substantially increased and that of TdP remains unchanged by RFE. These prediction accuracies are comparable to those of earlier studies derived from a selective set of descriptors. Our study suggests that molecular feature selection is useful for improving the speed and, in some cases, the accuracy of statistical learning methods for the prediction of pharmacokinetic and toxicological properties of chemical agents.

Algorithms↗

[Application of support vector machines to classification of blood cells].

The support vector machine (SVM) is a new learning technique based on the statistical learning theory. It was originally developed for two-class classification. In this paper, the SVM approach is extended to multi-class classification problems, a hierarchical SVM is applied to classify blood cells in different maturation stages from bone marrow. Based on stepwise decomposition, a hierarchical clustering method is presented to construct the architecture of the hierarchical (tree-like) SVM, then the optimal control parameters of SVM are determined by some criterion for each discriminant step. To verify the performances of classifiers, the SVM method is compared with three classical classifiers using 3-fold cross validation. The preliminary results indicate that the proposed method avoids the curse of dimensionality and has greater generalization. Thus, the method can improve the classification correctness for blood cells from bone marrow.

Algorithms↗

RNA secondary structure prediction from sequence alignments using a network of k-nearest neighbor classifiers.

We present a machine learning method (a hierarchical network of k-nearest neighbor classifiers) that uses an RNA sequence alignment in order to predict a consensus RNA secondary structure. The input to the network is the mutual information, the fraction of complementary nucleotides, and a novel consensus RNAfold secondary structure prediction of a pair of alignment columns and its nearest neighbors. Given this input, the network computes a prediction as to whether a particular pair of alignment columns corresponds to a base pair. By using a comprehensive test set of 49 RFAM alignments, the program KNetFold achieves an average Matthews correlation coefficient of 0.81. This is a significant improvement compared with the secondary structure prediction methods PFOLD and RNAalifold. By using the example of archaeal RNase P, we show that the program can also predict pseudoknot interactions.

Algorithms↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗

Analysis of differentially-regulated genes within a regulatory network by GPS genome navigation.

MOTIVATION: A critical challenge of the post-genomic era is to understand how genes are differentially regulated even when they belong to a given network. Because the fundamental mechanism controlling gene expression operates at the level of transcription initiation, computational techniques have been developed that identify cis regulatory features and map such features into expression patterns to classify genes into distinct networks. However, these methods are not focused on distinguishing between differentially regulated genes within a given network. Here we describe an unsupervised machine learning method, termed GPS for gene promoter scan, that discriminates among co-regulated promoters by simultaneously considering both cis-acting regulatory features and gene expression. GPS is particularly useful for knowledge discovery in environments with reduced datasets and high levels of uncertainty. RESULTS: Application of this method to the enteric bacteria Escherichia coli and Salmonella enterica uncovered novel members, as well as regulatory interactions in the regulon controlled by the PhoP protein that were not discovered using previous approaches. The predictions made by GPS were experimentally validated to establish that the PhoP protein uses multiple mechanisms to control gene transcription, and is a central element in a highly connected network. AVAILABILITY: The scripts and programs used in this work are accessible from the gps-tools.wustl.edu website. Data and predictions are available by request.

Algorithms↗

Prediction of the axillary lymph node status in mammary cancer on the basis of clinicopathological data and flow cytometry.

Axillary lymph node status is a major prognostic factor in mammary carcinoma. It is clinically desirable to predict the axillary lymph node status from data from the mammary cancer specimen. In the study, the axillary lymph node status, routine histological parameters and flow-cytometric data were retrospectively obtained from 1139 specimens of invasive mammary cancer. The ten variables: age, tumour type, tumour grade, tumour size, skin infiltration, lymphangiosis carcinomatosa, pT4 category, percentage of tumour cells in G2/M- and S-phases of the cell cycle, and ploidy index were considered as predictor variables, and the single variable lymph node metastasis pN (0 for pN0, or 1 for pN1 or pN2) was used as an output variable. A stepwise logistic regression analysis, with the axillary lymph node as a dependent variable, was used for feature selection. Only lymphangiosis carcinomatosa and tumour size proved to be significant as independent predictor variables; the other variables were non-contributory. Three paradigms with supervised learning rules (multilayer perceptron, learning vector quantisation and support vector machines) were used for the purpose of prediction. If any of these paradigms was used with the information from all ten input variables, 73% of cases could be correctly predicted, with specificity ranging from 82 to 84% and sensitivity ranging from 60 to 63%. If only the two significant input variables were used, lymphangiosis carcinomatosa and tumour diameter, the prediction accuracy was no worse. Nearly identical results were obtained by two different techniques of cross-validation (leave-one-out against ten-fold cross validation). It was concluded that: artificial neural networks can be used for risk stratification on the basis of routine data in individual cases of mammary cancer; and lymphangiosis carcinomatosa and tumour size are independent predictors of axillary lymph node metastasis in mammary cancer.

Algorithms↗

A gene mapping expert system.

Expert systems are now commonly developed to solve practical problems. Nevertheless, genetics has just begun to benefit from this new technology, since genetic expert systems are extremely rare and often purely experimental. A prototype for risk calculation in pedigrees was developed at the University of Utah, using a commercial frames/rules developmental shell (Intelligence Compiler), which runs on an IBM PC. When small data sets were used, the implementation functioned well, but it could not handle larger data sets. Performance became a major issue, with two possible solutions. The first possibility would have been to port the system to a more powerful machine, and the second would have been to use several different shells or languages, each efficiently representing a specific type of knowledge. Neither of these solutions was applicable in this case. From this experience, we learned that performance, portability, and modifiability were three major requirements for genetic expert systems. To achieve these goals, we implemented the gene mapping expert system GMES: (GMES is unrelated to the gene mapping system, GMS in Lisp combined with a frame/object shell (FROBS). We were able to efficiently represent, control, and optimize a gene mapping experiment, achieving portability by building GMES on top of a C-based version of Common Lisp. Lisp combined with the FROBS expert system shell permitted a declarative representation of each of the components of the experiment, resulting in a transplant specification of the problem within a maintainable system.

Algorithms↗

A spatio-temporal Bayesian network classifier for understanding visual field deterioration.

OBJECTIVE: Progressive loss of the field of vision is characteristic of a number of eye diseases such as glaucoma which is a leading cause of irreversible blindness in the world. Recently, there has been an explosion in the amount of data being stored on patients who suffer from visual deterioration including field test data, retinal image data and patient demographic data. However, there has been relatively little work in modelling the spatial and temporal relationships common to such data. In this paper we introduce a novel method for classifying visual field (VF) data that explicitly models these spatial and temporal relationships. METHODOLOGY: We carry out an analysis of our proposed spatio-temporal Bayesian classifier and compare it to a number of classifiers from the machine learning and statistical communities. These are all tested on two datasets of VF and clinical data. We investigate the receiver operating characteristics curves, the resulting network structures and also make use of existing anatomical knowledge of the eye in order to validate the discovered models. RESULTS: Results are very encouraging showing that our classifiers are comparable to existing statistical models whilst also facilitating the understanding of underlying spatial and temporal relationships within VF data. The results reveal the potential of using such models for knowledge discovery within ophthalmic databases, such as networks reflecting the 'nasal step', an early indicator of the onset of glaucoma. CONCLUSION: The results outlined in this paper pave the way for a substantial program of study involving many other spatial and temporal datasets, including retinal image and clinical data.

Algorithms↗

Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.

MOTIVATION: Human single nucleotide polymorphisms (SNPs) are the most frequent type of genetic variation in human population. One of the most important goals of SNP projects is to understand which human genotype variations are related to Mendelian and complex diseases. Great interest is focused on non-synonymous coding SNPs (nsSNPs) that are responsible of protein single point mutation. nsSNPs can be neutral or disease associated. It is known that the mutation of only one residue in a protein sequence can be related to a number of pathological conditions of dramatic social impact such as Alzheimer's, Parkinson's and Creutzfeldt-Jakob's diseases. The quality and completeness of presently available SNPs databases allows the application of machine learning techniques to predict the insurgence of human diseases due to single point protein mutation starting from the protein sequence. RESULTS: In this paper, we develop a method based on support vector machines (SVMs) that starting from the protein sequence information can predict whether a new phenotype derived from a nsSNP can be related to a genetic disease in humans. Using a dataset of 21 185 single point mutations, 61% of which are disease-related, out of 3587 proteins, we show that our predictor can reach more than 74% accuracy in the specific task of predicting whether a single point mutation can be disease related or not. Our method, although based on less information, outperforms other web-available predictors implementing different approaches. AVAILABILITY: A beta version of the web tool is available at http://gpcr.biocomp.unibo.it/cgi/predictors/PhD-SNP/PhD-SNP.cgi

Algorithms↗

Using symbolic knowledge in the UMLS to disambiguate words in small datasets with a naïve Bayes classifier.

Current approaches to word sense disambiguation use and combine various machine-learning techniques. Most refer to characteristics of the ambiguous word and surrounding words and are based on hundreds of examples. Unfortunately, developing large training sets is time-consuming. We investigate the use of symbolic knowledge to augment machine-learning techniques for small datasets. UMLS semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. A naïve Bayes classifier was trained for 15 words with 100 examples for each. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in eight experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. In a follow-up evaluation, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators.

Abstracting and Indexing↗

A neural-network-based method for predicting protein stability changes upon single point mutations.

MOTIVATION: One important requirement for protein design is to be able to predict changes of protein stability upon mutation. Different methods addressing this task have been described and their performance tested considering global linear correlation between predicted and experimental data. Neither is direct statistical evaluation of their prediction performance available, nor is a direct comparison among different approaches possible. Recently, a significant database of thermodynamic data on protein stability changes upon single point mutation has been generated (ProTherm). This allows the application of machine learning techniques to predicting free energy stability changes upon mutation starting from the protein sequence. RESULTS: In this paper, we present a neural-network-based method to predict if a given mutation increases or decreases the protein thermodynamic stability with respect to the native structure. Using a dataset consisting of 1615 mutations, our predictor correctly classifies >80% of the mutations in the database. On the same task and using the same data, our predictor performs better than other methods available on the Web. Moreover, when our system is coupled with energy-based methods, the joint prediction accuracy increases up to 90%, suggesting that it can be used to increase also the performance of pre-existing methods, and generally to improve protein design strategies. AVAILABILITY: The server is under construction and will be available at http://www.biocomp.unibo.it

Algorithms↗

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans↗

Feature selection in MLPs and SVMs based on maximum output information.

This paper presents feature selection algorithms for multilayer perceptrons (MLPs) and multiclass support vector machines (SVMs), using mutual information between class labels and classifier outputs, as an objective function. This objective function involves inexpensive computation of information measures only on discrete variables; provides immunity to prior class probabilities; and brackets the probability of error of the classifier. The maximum output information (MOI) algorithms employ this function for feature subset selection by greedy elimination and directed search. The output of the MOI algorithms is a feature subset of user-defined size and an associated trained classifier (MLP/SVM). These algorithms compare favorably with a number of other methods in terms of performance on various artificial and real-world data sets.

Algorithms↗

Prediction of enzyme classification from protein sequence without the use of sequence similarity.

We describe a novel approach for predicting the function of a protein from its amino-acid sequence. Given features that can be computed from the amino-acid sequence in a straightforward fashion (such as pI, molecular weight, and amino-acid composition), the technique allows us to answer questions such as: Is the protein an enzyme? If so, in which Enzyme Commission (EC) class does it belong? Our approach uses machine learning (ML) techniques to induce classifiers that predict the EC class of an enzyme from features extracted from its primary sequence. We report on a variety of experiments in which we explored the use of three different ML techniques in conjunction with training datasets derived from PDB and from Swiss-Prot. We also explored the use of several different feature sets. Our method is able to predict the first EC number of an enzyme with 74% accuracy (thereby assigning the enzyme to one of six broad categories of enzyme function), and to predict the second EC number of an enzyme with 68% accuracy (thereby assigning the enzyme to one of 57 subcategories of enzyme function). This technique could be a valuable complement to sequence-similarity searches and to pathway-analysis methods.

Algorithms↗

Knowledge-based analysis of microarray gene expression data by using support vector machines.

We introduce a method of functionally classifying genes by using gene expression data from DNA microarray hybridization experiments. The method is based on the theory of support vector machines (SVMs). SVMs are considered a supervised computer learning method because they exploit prior knowledge of gene function to identify unknown genes of similar function from expression data. SVMs avoid several problems associated with unsupervised clustering methods, such as hierarchical clustering and self-organizing maps. SVMs have many mathematical features that make them attractive for gene expression analysis, including their flexibility in choosing a similarity function, sparseness of solution when dealing with large data sets, the ability to handle large feature spaces, and the ability to identify outliers. We test several SVMs that use different similarity metrics, as well as some other supervised learning methods, and find that the SVMs best identify sets of genes with a common function using expression data. Finally, we use SVMs to predict functional roles for uncharacterized yeast ORFs based on their expression data.

Algorithms↗