PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Prediction of ultrasound-mediated disruption of cell membranes using machine learning techniques and statistical analysis of acoustic spectra.

Although biological effects of ultrasound must be avoided for safe diagnostic applications, ultrasound's ability to disrupt cell membranes has attracted interest as a method to facilitate drug and gene delivery. This paper seeks to develop "prediction rules" for predicting the degree of cell membrane disruption based on specified ultrasound parameters and measured acoustic signals. Three techniques for generating prediction rules (regression analysis, classification trees and discriminant analysis) are applied to data obtained from a sequence of experiments on bovine red blood cells. For each experiment, the data consist of four ultrasound parameters, acoustic measurements at 400 frequencies, and a measure of cell membrane disruption. To avoid over-training, various combinations of the 404 predictor variables are used when applying the rule generation methods. The results indicate that the variable combination consisting of ultrasound exposure time and acoustic signals measured at the driving frequency and its higher harmonics yields the best rule for all three rule generation methods. The methods used for deriving the prediction rules are broadly applicable, and could be used to develop prediciton rules in other scenarios involving different cell types or tissues. These rules and the methods used to derive them could be used for real-time feedback about ultrasound's biological effects.

Animals↗

Dynamic probability estimator for machine learning.

An efficient algorithm for dynamic estimation of probabilities without division on unlimited number of input data is presented. The method estimates probabilities of the sampled data from the raw sample count, while keeping the total count value constant. Accuracy of the estimate depends on the counter size, rather than on the total number of data points. Estimator follows variations of the incoming data probability within a fixed window size, without explicit implementation of the windowing technique. Total design area is very small and all probabilities are estimated concurrently. Dynamic probability estimator was implemented using a programmable gate array from Xilinx. The performance of this implementation is evaluated in terms of the area efficiency and execution time. This method is suitable for the highly integrated design of artificial neural networks where a large number of dynamic probability estimators can work concurrently.

Artificial Intelligence↗

Evaluating robustness of gait event detection based on machine learning and natural sensors.

A real-time system for deriving timing control for functional electrical stimulation for foot-drop correction, using peripheral nerve activity as a sensor input, was tested for reliability to investigate the potential for clinical use. The system, which was previously reported on, was tested on a hemiplegic subject instrumented with a recording cuff electrode on the Sural nerve, and a stimulation cuff electrode on the Peroneal cuff. Implanted devices enabled recording and stimulation through telelinks. An input domain was derived from the recorded electroneurogram and fed to a detection algorithm based on an adaptive logic network for controlling the stimulation timing. The reliability was tested by letting the subject wear different foot wear and walk on different surfaces than when the training data was recorded. The detection system was also evaluated several months after training. The detection system proved able to successfully detect when walking with different footwear on varying surfaces up to 374 days after training, and thereby showed great potential for being clinically useful.

Action Potentials↗

Theodor Bücher Lecture. Metabolomics, modelling and machine learning in systems biology - towards an understanding of the languages of cells. Delivered on 3 July 2005 at the 30th FEBS Congress and the 9th IUBMB conference in Budapest.

The newly emerging field of systems biology involves a judicious interplay between high-throughput 'wet' experimentation, computational modelling and technology development, coupled to the world of ideas and theory. This interplay involves iterative cycles, such that systems biology is not at all confined to hypothesis-dependent studies, with intelligent, principled, hypothesis-generating studies being of high importance and consequently very far from aimless fishing expeditions. I seek to illustrate each of these facets. Novel technology development in metabolomics can increase substantially the dynamic range and number of metabolites that one can detect, and these can be exploited as disease markers and in the consequent and principled generation of hypotheses that are consistent with the data and achieve this in a value-free manner. Much of classical biochemistry and signalling pathway analysis has concentrated on the analyses of changes in the concentrations of intermediates, with 'local' equations - such as that of Michaelis and Menten v=(Vmax x S)/(S+K m) - that describe individual steps being based solely on the instantaneous values of these concentrations. Recent work using single cells (that are not subject to the intellectually unsupportable averaging of the variable displayed by heterogeneous cells possessing nonlinear kinetics) has led to the recognition that some protein signalling pathways may encode their signals not (just) as concentrations (AM or amplitude-modulated in a radio analogy) but via changes in the dynamics of those concentrations (the signals are FM or frequency-modulated). This contributes in principle to a straightforward solution of the crosstalk problem, leads to a profound reassessment of how to understand the downstream effects of dynamic changes in the concentrations of elements in these pathways, and stresses the role of signal processing (and not merely the intermediates) in biological signalling. It is this signal processing that lies at the heart of understanding the languages of cells. The resolution of many of the modern and postgenomic problems of biochemistry requires the development of a myriad of new technologies (and maybe a new culture), and thus regular input from the physical sciences, engineering, mathematics and computer science. One solution, that we are adopting in the Manchester Interdisciplinary Biocentre (http://www.mib.ac.uk/) and the Manchester Centre for Integrative Systems Biology (http://www.mcisb.org/), is thus to colocate individuals with the necessary combinations of skills. Novel disciplines that require such an integrative approach continue to emerge. These include fields such as chemical genomics, synthetic biology, distributed computational environments for biological data and modelling, single cell diagnostics/bionanotechnology, and computational linguistics/text mining.

Artificial Intelligence↗

Rapid and quantitative detection of the microbial spoilage of meat by fourier transform infrared spectroscopy and machine learning.

Fourier transform infrared (FT-IR) spectroscopy is a rapid, noninvasive technique with considerable potential for application in the food and related industries. We show here that this technique can be used directly on the surface of food to produce biochemically interpretable "fingerprints." Spoilage in meat is the result of decomposition and the formation of metabolites caused by the growth and enzymatic activity of microorganisms. FT-IR was exploited to measure biochemical changes within the meat substrate, enhancing and accelerating the detection of microbial spoilage. Chicken breasts were purchased from a national retailer, comminuted for 10 s, and left to spoil at room temperature for 24 h. Every hour, FT-IR measurements were taken directly from the meat surface using attenuated total reflectance, and the total viable counts were obtained by classical plating methods. Quantitative interpretation of FT-IR spectra was possible using partial least-squares regression and allowed accurate estimates of bacterial loads to be calculated directly from the meat surface in 60 s. Genetic programming was used to derive rules showing that at levels of 10(7) bacteria.g(-1) the main biochemical indicator of spoilage was the onset of proteolysis. Thus, using FT-IR we were able to acquire a metabolic snapshot and quantify, noninvasively, the microbial loads of food samples accurately and rapidly in 60 s, directly from the sample surface. We believe this approach will aid in the Hazard Analysis Critical Control Point process for the assessment of the microbiological safety of food at the production, processing, manufacturing, packaging, and storage levels.

Artificial Intelligence↗

Development of neural mechanisms for machine learning.

The goal of this work is to develop a humanoid robot's perceptual mechanisms through the use of learning aids. We describe methods to enable learning on a humanoid robot using learning aids such as books, drawing materials, boards, educational videos or other children toys. Visual properties of objects are learned and inserted into a recognition scheme, which is then applied to acquire new object representations - we propose learning through developmental stages. Inspired in infant development, we will also boost the robot's perceptual capabilities by having a human caregiver performing educational and play activities with the robot (such as drawing, painting or playing with a toy train on a railway). We describe original algorithms to extract meaningful percepts from such learning experiments. Experimental evaluation of the algorithms corroborates the theoretical framework.

Algorithms↗

Application of machine learning and visualization of heterogeneous datasets to uncover relationships between translation and developmental stage expression of C. elegans mRNAs.

The relationships between genes in neighboring clusters in a self-organizing map (SOM) and properties attributed to them are sometimes difficult to discern, especially when heterogeneous datasets are used. We report a novel approach to identify correlations between heterogeneous datasets. One dataset, derived from microarray analysis of polysomal distribution, contained changes in the translational efficiency of Caenorhabditis elegans mRNAs resulting from loss of specific eIF4E isoform. The other dataset contained expression patterns of mRNAs across all developmental stages. Two algorithms were applied to these datasets: a classical scatter plot and an SOM. The outputs were linked using a two-dimensional color scale. This revealed that an mRNA's eIF4E-dependent translational efficiency is strongly dependent on its expression during development. This correlation was not detectable with a traditional one-dimensional color scale.

Algorithms↗

A machine learning method for extracting symbolic knowledge from recurrent neural networks.

Neural networks do not readily provide an explanation of the knowledge stored in their weights as part of their information processing. Until recently, neural networks were considered to be black boxes, with the knowledge stored in their weights not readily accessible. Since then, research has resulted in a number of algorithms for extracting knowledge in symbolic form from trained neural networks. This article addresses the extraction of knowledge in symbolic form from recurrent neural networks trained to behave like deterministic finite-state automata (DFAs). To date, methods used to extract knowledge from such networks have relied on the hypothesis that networks' states tend to cluster and that clusters of network states correspond to DFA states. The computational complexity of such a cluster analysis has led to heuristics that either limit the number of clusters that may form during training or limit the exploration of the space of hidden recurrent state neurons. These limitations, while necessary, may lead to decreased fidelity, in which the extracted knowledge may not model the true behavior of a trained network, perhaps not even for the training set. The method proposed here uses a polynomial time, symbolic learning algorithm to infer DFAs solely from the observation of a trained network's input-output behavior. Thus, this method has the potential to increase the fidelity of the extracted knowledge.

Algorithms↗

Machine learning paradigms for pattern recognition and image understanding.

In this paper some issues are considered related to the encoding of spatial information and associated perceptual learning algorithms which, it is claimed, are necessary for robust pattern and object recognition in multi-object (natural) scenes. The types of learning requirements within a 'recognition-by-parts' paradigm are contrasted with findings from alternative models.

Form Perception↗

A machine learning strategy to identify candidate binding sites in human protein-coding sequence.

BACKGROUND: The splicing of RNA transcripts is thought to be partly promoted and regulated by sequences embedded within exons. Known sequences include binding sites for SR proteins, which are thought to mediate interactions between splicing factors bound to the 5' and 3' splice sites. It would be useful to identify further candidate sequences, however identifying them computationally is hard since exon sequences are also constrained by their functional role in coding for proteins. RESULTS: This strategy identified a collection of motifs including several previously reported splice enhancer elements. Although only trained on coding exons, the model discriminates both coding and non-coding exons from intragenic sequence. CONCLUSION: We have trained a computational model able to detect signals in coding exons which seem to be orthogonal to the sequences' primary function of coding for proteins. We believe that many of the motifs detected here represent binding sites for both previously unrecognized proteins which influence RNA splicing as well as other regulatory elements.

Algorithms↗

Automating the assignment of diagnosis codes to patient encounters using example-based and machine learning techniques.

OBJECTIVE: Human classification of diagnoses is a labor intensive process that consumes significant resources. Most medical practices use specially trained medical coders to categorize diagnoses for billing and research purposes. METHODS: We have developed an automated coding system designed to assign codes to clinical diagnoses. The system uses the notion of certainty to recommend subsequent processing. Codes with the highest certainty are generated by matching the diagnostic text to frequent examples in a database of 22 million manually coded entries. These code assignments are not subject to subsequent manual review. Codes at a lower certainty level are assigned by matching to previously infrequently coded examples. The least certain codes are generated by a naïve Bayes classifier. The latter two types of codes are subsequently manually reviewed. MEASUREMENTS: Standard information retrieval accuracy measurements of precision, recall and f-measure were used. Micro- and macro-averaged results were computed. RESULTS At least 48% of all EMR problem list entries at the Mayo Clinic can be automatically classified with macro-averaged 98.0% precision, 98.3% recall and an f-score of 98.2%. An additional 34% of the entries are classified with macro-averaged 90.1% precision, 95.6% recall and 93.1% f-score. The remaining 18% of the entries are classified with macro-averaged 58.5%. CONCLUSION: Over two thirds of all diagnoses are coded automatically with high accuracy. The system has been successfully implemented at the Mayo Clinic, which resulted in a reduction of staff engaged in manual coding from thirty-four coders to seven verifiers.

Abstracting and Indexing↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

Machine learning for the quality of life in inflammatory bowel disease.

Presence of a chronic disease influences patients' lives and reinforces demands to accept and then cope with the illness. In the case of inflammatory bowel disease, quality of life greatly differs through phases of remissions and relapses. Could the quality of life questionnaire tell the difference? In this study we are disclosing possibilities of assessing patients' perspectives by analysing analogue scale statements regarding concerns and worries related to ulcerative colitis. Some two hundred Swedish patients, 3/4 in remission and 1/4 in relapse, filled out a booklet containing 36 statements. To characterise the disease activity, we have used multivariate discrimination. To structure and describe in details paths distinguishing the remission from relapse, we have used an artificial intelligence procedure. Applications of the CART (Classification And Regression Trees) algorithm resulted in a set of classifiers which are, based on the similar subsets of significant variables, i.e. statements. Best reached classification accuracy did not exceed 80% in any case. Other classifiers namely, K-nearest-neighbour (KNN), Learning Vector Quantization (LVQ) and Back Propagation Neural Network (BPNN) confirmed that outcome. An expectation that the disease activity should clearly speak throughout the questionnaire held for a certain number of the observations such as pain and suffering, loss of bowel control, dying early, feeling alone, ability to have children, being treated as different and concerns regarding the medication. To highlight the difference of incorrect 20%, K-means clustering was performed. The results settled a basis for a hypothesis that the studied quality of life instrument captures more than the disease activity.

Algorithms↗

The effect of sample size and disease prevalence on supervised machine learning of narrative data.

This paper examines the independent effects of outcome prevalence and training sample sizes on inductive learning performance. We trained 3 inductive learning algorithms (MC4, IB, and Naïve-Bayes) on 60 simulated datasets of parsed radiology text reports labeled with 6 disease states. Data sets were constructed to define positive outcome states at 4 prevalence rates (1, 5, 10, 25, and 50%) in training set sizes of 200 and 2,000 cases. We found that the effect of outcome prevalence is significant when outcome classes drop below 10% of cases. The effect appeared independent of sample size, induction algorithm used, or class label. Work is needed to identify methods of improving classifier performance when output classes are rare.

Algorithms↗

On selecting features from splice junctions: an analysis using information theoretic and machine learning approaches.

The computational recognition of precise splice junctions is a challenge faced in the analysis of newly sequenced genomes. This is challenging due to the fact that the distribution of sequence patterns in these regions is not always distinct. Our objective is to understand the sequence signatures at the splice junctions, not simply to create an artificial recognition system. We use a combination of a neural network based calliper randomization approach and an information theoretic based feature selection approach for this purpose. This has been done in an effort to understand regions that harbor information content and to extract features relevant for the prediction of splice junctions. The analysis using the neural network based calliper randomization approach revealed regions important in the internal representation of the network model. The calliper approach captured both correlated as well as independently important features. The feature selection approach captures features that are independently informative. The two different methods can capture features with different properties. Comparative analysis of the results using both the methods help to infer about the kind of information present in the region.

Alternative Splicing↗

The use of machine learning program LERS-LB 2.5 in knowledge acquisition for expert system development in nursing.

LERS-LB (Learning from Examples using Rough Sets Lower Boundaries) is a computer program based on rough set theory for knowledge acquisition, which extracts patterns from real-world data in generating production rules for expert system development. From LERS-LB evaluation of an SPSS-X data file containing data for recovery room patients, it was concluded that both statistical data files and existing databases can be converted to decision-table format needed by LERS-LB, but it is less desirable to work with statistical files than a well-developed database. It was also concluded that choosing a well-developed database and checking it thoroughly for accuracy and completeness should be done before running LERS-LB, or other learning programs, to avoid problems with data errors. Using rough set theory and a technique called 'dropping conditions' LERS-LB offers, at least in theory, a possible method for identifying which data items are critical to nursing practice. Further research and continued LERS-LB program enhancements still may help with identifying critical data items versus redundant data for nursing practice. LERS-LB, and other learning programs, offer techniques which will help reduce the knowledge acquisition bottleneck in nursing expert system development. It is doubtful, however, that learning programs will eliminate the need for involving domain experts in evaluating rules and expert systems for clinical decision support.

Artificial Intelligence↗