PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Fedflow: cloud orchestration for federated learning with the FeatureCloud platform.

MOTIVATION: Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. RESULTS: We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. AVAILABILITY: Fedflow is open-source and available at https://github.com/W-L/fedflow.

Journal Article↗

Learning Boolean queries for article quality filtering.

Prior research has shown that Support Vector Machine models have the ability to identify high quality content-specific articles in the domain of internal medicine. These models, though powerful, cannot be used in Boolean search engines nor can the content of the models be verified via human inspection. In this paper, we use decision trees combined with several feature selection methods to generate Boolean query filters for the same domain and task. The resulting trees are generated automatically and exhibit high performance. The trees are understandable, manageable, and able to be validated by humans. The subsequent Boolean queries are sensible and can be readily used as filters by Boolean search engines.

Algorithms↗

Improving the performance of dictionary-based approaches in protein name recognition.

Dictionary-based protein name recognition is often a first step in extracting information from biomedical documents because it can provide ID information on recognized terms. However, dictionary-based approaches present two fundamental difficulties: (1) false recognition mainly caused by short names; (2) low recall due to spelling variations. In this paper, we tackle the former problem using machine learning to filter out false positives and present two alternative methods for alleviating the latter problem of spelling variations. The first is achieved by using approximate string searching, and the second by expanding the dictionary with a probabilistic variant generator, which we propose in this paper. Experimental results using the GENIA corpus revealed that filtering using a naive Bayes classifier greatly improved precision with only a slight loss of recall, resulting in 10.8% improvement in F-measure, and dictionary expansion with the variant generator gave further 1.6% improvement and achieved an F-measure of 66.6%.

Abstracting and Indexing↗

Weighted quality estimates in machine learning.

MOTIVATION: Machine learning methods such as neural networks, support vector machines, and other classification and regression methods rely on iterative optimization of the model quality in the space of the parameters of the method. Model quality measures (accuracies, correlations, etc.) are frequently overly optimistic because the training sets are dominated by particular families and subfamilies. To overcome the bias, the dataset is usually reduced by filtering out closely related objects. However, such filtering uses fixed similarity thresholds and ignores a part of the training information. RESULTS: We suggested a novel approach to calculate prediction model quality based on assigning to each data point inverse density weights derived from the postulated distance metric. We demonstrated that our new weighted measures estimate the model generalization better and are consistent with the machine learning theory. The Vapnik-Chervonenkis theorem was reformulated and applied to derive the space-uniform error estimates. Two examples were used to illustrate the advantages of the inverse density weighting. First, we demonstrated on a set with a built-in bias that the unweighted cross-validation procedure leads to an overly optimistic quality estimate, while the density-weighted quality estimates are more realistic. Second, an analytical equation for weighted quality estimates was used to derive an SVM model for signal peptide prediction using a full set of known signal peptides, instead of the usual filtered subset.

Algorithms↗

Profile-based string kernels for remote homology detection and motif extraction.

We introduce novel profile-based string kernels for use with support vector machines (SVMs) for the problems of protein classification and remote homology detection. These kernels use probabilistic profiles, such as those produced by the PSI-BLAST algorithm, to define position-dependent mutation neighborhoods along protein sequences for inexact matching of k-length subsequences ("k-mers") in the data. By use of an efficient data structure, the kernels are fast to compute once the profiles have been obtained. For example, the time needed to run PSI-BLAST in order to build the pro- files is significantly longer than both the kernel computation time and the SVM training time. We present remote homology detection experiments based on the SCOP database where we show that profile-based string kernels used with SVM classifiers strongly outperform all recently presented supervised SVM methods. We also show how we can use the learned SVM classifier to extract "discriminative sequence motifs" -- short regions of the original profile that contribute almost all the weight of the SVM classification score -- and show that these discriminative motifs correspond to meaningful structural features in the protein data. The use of PSI-BLAST profiles can be seen as a semi-supervised learning technique, since PSI-BLAST leverages unlabeled data from a large sequence database to build more informative profiles. Recently presented "cluster kernels" give general semi-supervised methods for improving SVM protein classification performance. We show that our profile kernel results are comparable to cluster kernels while providing much better scalability to large datasets.

Algorithms↗

A machine learning approach for automated recognition of movement patterns using basic, kinetic and kinematic gait data.

This paper investigated application of a machine learning approach (Support vector machine, SVM) for the automatic recognition of gait changes due to ageing using three types of gait measures: basic temporal/spatial, kinetic and kinematic. The gaits of 12 young and 12 elderly participants were recorded and analysed using a synchronized PEAK motion analysis system and a force platform during normal walking. Altogether, 24 gait features describing the three types of gait characteristics were extracted for developing gait recognition models and later testing of generalization performance. Test results indicated an overall accuracy of 91.7% by the SVM in its capacity to distinguish the two gait patterns. The classification ability of the SVM was found to be unaffected across six kernel functions (linear, polynomial, radial basis, exponential radial basis, multi-layer perceptron and spline). Gait recognition rate improved when features were selected from different gait data type. A feature selection algorithm demonstrated that as little as three gait features, one selected from each data type, could effectively distinguish the age groups with 100% accuracy. These results demonstrate considerable potential in applying SVMs in gait classification for many applications.

Aged↗

Feature subset selection for splice site prediction.

MOTIVATION: The large amount of available annotated Arabidopsis thaliana sequences allows the induction of splice site prediction models with supervised learning algorithms (see Haussler (1998) for a review and references). These algorithms need information sources or features from which the models can be computed. For splice site prediction, the features we consider in this study are the presence or absence of certain nucleotides in close proximity to the splice site. Since it is not known how many and which nucleotides are relevant for splice site prediction, the set of features is chosen large enough such that the probability that all relevant information sources are in the set is very high. Using only those features that are relevant for constructing a splice site prediction system might improve the system and might also provide us with useful biological knowledge. Using fewer features will of course also improve the prediction speed of the system. RESULTS: A wrapper-based feature subset selection algorithm using a support vector machine or a naive Bayes prediction method was evaluated against the traditional method for selecting features relevant for splice site prediction. Our results show that this wrapper approach selects features that improve the performance against the use of all features and against the use of the features selected by the traditional method. AVAILABILITY: The data and additional interactive graphs on the selected feature subsets are available at http://www.psb.rug.ac.be/gps

Arabidopsis↗

Handling missing values in support vector machine classifiers.

This paper discusses the task of learning a classifier from observed data containing missing values amongst the inputs which are missing completely at random. A non-parametric perspective is adopted by defining a modified risk taking into account the uncertainty of the predicted outputs when missing values are involved. It is shown that this approach generalizes the approach of mean imputation in the linear case and the resulting kernel machine reduces to the standard Support Vector Machine (SVM) when no input values are missing. Furthermore, the method is extended to the multivariate case of fitting additive models using componentwise kernel machines, and an efficient implementation is based on the Least Squares Support Vector Machine (LS-SVM) classifier formulation.

Algorithms↗

Differentially expressed genes in gastric tumors identified by cDNA array.

Using cDNA fragments from the FAPESP/lICR Cancer Genome Project, we constructed a cDNA array having 4512 elements and determined gene expression in six normal and six tumor gastric tissues. Using t-statistics, we identified 80 cDNAs whose expression in normal and tumor samples differed more than 3.5 sample standard deviations. Using Self-Organizing Map, the expression profile of these cDNAs allowed perfect separation of malignant and non-malignant samples. Using the supervised learning procedure Support Vector Machine, we identified trios of cDNAs that could be used to classify samples as normal or tumor, based on single-array analysis. Finally, we identified genes with altered linear correlation when their expression in normal and tumor samples were compared. Further investigation concerning the function of these genes could contribute to the understanding of gastric carcinogenesis and may prove useful in molecular diagnostics.

Algorithms↗

Drug discovery using support vector machines. The case studies of drug-likeness, agrochemical-likeness, and enzyme inhibition predictions.

Support Vector Machines (SVM) is a powerful classification and regression tool that is becoming increasingly popular in various machine learning applications. We tested the ability of SVM, in comparison with well-known neural network techniques, to predict drug-likeness and agrochemical-likeness for large compound collections. For both kinds of data, SVM outperforms various neural networks using the same set of descriptors. We also used SVM for estimating the activity of Carbonic Anhydrase II (CA II) enzyme inhibitors and found that the prediction quality of our SVM model is better than that reported earlier for conventional QSAR. Model characteristics and data set features were studied in detail.

Agrochemicals↗

ABC stenosis morphology classification and outcome of coronary angioplasty: reassessment with computing techniques.

BACKGROUND: The American College of Cardiology/American Heart Association (ACC/AHA) stenosis morphology classification (MC) stratifies coronary lesions for probability of success and complications after coronary angioplasty (PTCA). Modern computing techniques were used to evaluate the individual predictive value of MC in random PTCA cases. METHODS AND RESULTS: MC was attributed to the target lesions by consensus of 2 observers. The predictive value regarding procedural success (PS) and major adverse cardiac events (MACE) of MC was analyzed by conventional logistic regression analyses and by inductive machine learning models. The study was adequately powered for the methods applied with 325 target lesions of 250 cases. Overall, PS decreased and MACE increased from type A to type C lesions. Regression analysis identified no single factor as predictive. Logistic regression showed an error rate of 42%. Machine learning techniques achieved an individual predictive error of only 10%, which could be further reduced to 2% by addition of parameters. For PS, MC parameters showed a high ranking for building the model. For MACE, variables of the medical history showed more impact. CONCLUSIONS: MC per se cannot individually predict PS or MACE. However, when all MC parameters are integrated together with additional lesion-specific and history variables, a high individual predictive value can be achieved. This technique may be clinically helpful for risk stratification in the catheterization laboratory and improvement of classification systems in interventional cardiology.

Algorithms↗

Distributed machine learning: scaling up with coarse-grained parallelism.

Machine learning methods are becoming accepted as additions to the biologists data-analysis tool kit. However, scaling these techniques up to large data sets, such as those in biological and medical domains, is problematic in terms of both the required computational search effort and required memory (and the detrimental effects of excessive swapping). Our approach to tackling the problem of scaling up to large datasets is to take advantage of the ubiquitous workstation networks that are generally available in scientific and engineering environments. This paper introduces the notion of the invariant-partitioning property--that for certain evaluation criteria it is possible to partition a data set across multiple processors such that any rule that is satisfactory over the entire data set will also be satisfactory on at least one subset. In addition, by taking advantage of cooperation through interprocess communication, it is possible to build distributed learning algorithms such that only rules that are satisfactory over the entire data set will be learned. We describe a distributed learning system, CorPRL, that takes advantage of the invariant-partitioning property to learn from very large data sets, and present results demonstrating CorPRL's effectiveness in analyzing data from two databases.

Database Management Systems↗

The Helmholtz machine.

Discovering the structure inherent in a set of patterns is a fundamental aim of statistical inference or learning. One fruitful approach is to build a parameterized stochastic generative model, independent draws from which are likely to produce the patterns. For all but the simplest generative models, each pattern can be generated in exponentially many ways. It is thus intractable to adjust the parameters to maximize the probability of the observed patterns. We describe a way of finessing this combinatorial explosion by maximizing an easily computed lower bound on the probability of the observations. Our method can be viewed as a form of hierarchical self-supervised learning that may relate to the function of bottom-up and top-down cortical processing pathways.

Algorithms↗

Feature mining and predictive model construction from severe trauma patient's data.

In management of severe trauma patients, trauma surgeons need to decide which patients are eligible for damage control. Such decision may be supported by utilizing models that predict the patient's outcome. The study described in this paper investigates the possibility to construct patient outcome prediction models from retrospective patient's data at the end of initial damage control surgery by using feature mining and machine learning techniques. As the data used comprises rather excessive number of features, special attention was paid to the problem of selecting only the most relevant features. We show that a small subset of features may carry enough information to construct reasonably accurate prognostic models. Furthermore, the techniques used in our study identified two factors, namely the pH value when admitted to ICU and the worst partial active thromboplastin time, to be of highest importance for prediction. This finding is pathophysiologically reasonable and represents two of three major problems with severe trauma patients, metabolic acidosis, hypothermia, and coagulopathy.

Algorithms↗

Learning Petri net models of non-linear gene interactions.

Understanding how an individual's genetic make-up influences their risk of disease is a problem of paramount importance. Although machine-learning techniques are able to uncover the relationships between genotype and disease, the problem of automatically building the best biochemical model or "explanation" of the relationship has received less attention. In this paper, I describe a method based on random hill climbing that automatically builds Petri net models of non-linear (or multi-factorial) disease-causing gene-gene interactions. Petri nets are a suitable formalism for this problem, because they are used to model concurrent, dynamic processes analogous to biochemical reaction networks. I show that this method is routinely able to identify perfect Petri net models for three disease-causing gene-gene interactions recently reported in the literature.

Algorithms↗

On the relationship between deterministic and probabilistic directed Graphical models: from Bayesian networks to recursive neural networks.

Machine learning methods that can handle variable-size structured data such as sequences and graphs include Bayesian networks (BNs) and Recursive Neural Networks (RNNs). In both classes of models, the data is modeled using a set of observed and hidden variables associated with the nodes of a directed acyclic graph. In BNs, the conditional relationships between parent and child variables are probabilistic, whereas in RNNs they are deterministic and parameterized by neural networks. Here, we study the formal relationship between both classes of models and show that when the source nodes variables are observed, RNNs can be viewed as limits, both in distribution and probability, of BNs with local conditional distributions that have vanishing covariance matrices and converge to delta functions. Conditions for uniform convergence are also given together with an analysis of the behavior and exactness of Belief Propagation (BP) in 'deterministic' BNs. Implications for the design of mixed architectures and the corresponding inference algorithms are briefly discussed.

Bayes Theorem↗

Predicting risk of coronary artery disease from DNA microarray-based genotyping using neural networks and other statistical analysis tool.

This paper presents a novel approach for complex disease prediction that we have developed, exemplified by a study on risk of coronary artery disease (CAD). This multi-disciplinary approach straddles fields of microarray technology and genetics, neural networks (NN), data mining and machine learning, as well as traditional statistical analysis techniques, namely principal components analysis (PCA) and factor analysis (FA). A description of the biological background of the study is given, followed by a detailed description of how the problem has been modeled for analyses by neural networks and FA. A committee learning approach for NN has been used to improve generalization rates. We show that our NN approach is able to yield promising prediction results despite using only the most fundamental network structures. More interestingly, through the statistical analysis process, genes of similar biological functions have been clustered. In addition, a gene marker involved in breaking down lipids has been found to be the most correlated to CAD.

Algorithms↗

Gaussian processes for classification: mean-field algorithms.

We derive a mean-field algorithm for binary classification with gaussian processes that is based on the TAP approach originally proposed in statistical physics of disordered systems. The theory also yields an approximate leave-one-out estimator for the generalization error, which is computed with no extra computational cost. We show that from the TAP approach, it is possible to derive both a simpler "naive" mean-field theory and support vector machines (SVMs) as limiting cases. For both mean-field algorithms and support vector machines, simulation results for three small benchmark data sets are presented. They show that one may get state-of-the-art performance by using the leave-one-out estimator for model selection and the built-in leave-one-out estimators are extremely precise when compared to the exact leave-one-out estimate. The second result is taken as strong support for the internal consistency of the mean-field approach.

Algorithms↗