PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Rapid identification of closely related muscle foods by vibrational spectroscopy and machine learning.

Muscle foods are an integral part of the human diet and during the last few decades consumption of poultry products in particular has increased significantly. It is important for consumers, retailers and food regulatory bodies that these products are of a consistently high quality, authentic, and have not been subjected to adulteration by any lower-grade material either by accident or for economic gain. A variety of methods have been developed for the identification and authentication of muscle foods. However, none of these are rapid or non-invasive, all are time-consuming and difficulties have been encountered in discriminating between the commercially important avian species. Whilst previous attempts have been made to discriminate between muscle foods using infrared spectroscopy, these have had limited success, in particular regarding the closely related poultry species, chicken and turkey. Moreover, this study includes novel data since no attempts have been made to discriminate between both the species and the distinct muscle groups within these species, and this is the first application of Raman spectroscopy to the study of muscle foods. Samples of pre-packed meat and poultry were acquired and FT-IR and Raman measurements taken directly from the meat surface. Qualitative interpretation of FT-IR and Raman spectra at the species and muscle group levels were possible using discriminant function analysis. Genetic algorithms were used to elucidate meaningful interpretation of FT-IR results in (bio)chemical terms and we show that specific wavenumbers, and therefore chemical species, were discriminatory for each type (species and muscle) of poultry sample. We believe that this approach would aid food regulatory bodies in the rapid identification of meat and poultry products and shows particular potential for rapid assessment of food adulteration.

Algorithms↗

Structure-activity relationships derived by machine learning: the use of atoms and their bond connectivities to predict mutagenicity by inductive logic programming.

We present a general approach to forming structure-activity relationships (SARs). This approach is based on representing chemical structure by atoms and their bond connectivities in combination with the inductive logic programming (ILP) algorithm PROGOL. Existing SAR methods describe chemical structure by using attributes which are general properties of an object. It is not possible to map chemical structure directly to attribute-based descriptions, as such descriptions have no internal organization. A more natural and general way to describe chemical structure is to use a relational description, where the internal construction of the description maps that of the object described. Our atom and bond connectivities representation is a relational description. ILP algorithms can form SARs with relational descriptions. We have tested the relational approach by investigating the SARs of 230 aromatic and heteroaromatic nitro compounds. These compounds had been split previously into two subsets, 188 compounds that were amenable to regression and 42 that were not. For the 188 compounds, a SAR was found that was as accurate as the best statistical or neural network-generated SARs. The PROGOL SAR has the advantages that it did not need the use of any indicator variables handcrafted by an expert, and the generated rules were easily comprehensible. For the 42 compounds, PROGOL formed a SAR that was significantly (P < 0.025) more accurate than linear regression, quadratic regression, and back-propagation. This SAR is based on an automatically generated structural alert for mutagenicity.

Algorithms↗

Machine learning methods applied on dental fear and behavior management problems in children.

The etiologies of dental fear and dental behavior management problems in children were investigated in a database of information on 2,257 Swedish children 4-6 and 9-11 years old. The analyses were performed using computerized inductive techniques within the field of artificial intelligence. The database held information regarding dental fear levels and behavior management problems, which were defined as outcomes, i.e. dependent variables. The attributes, i.e. independent variables, included data on dental health and dental treatments, information about parental dental fear, general anxiety, socioeconomic variables, etc. The data contained both numerical and discrete variables. The analyses were performed using an inductive analysis program (XpertRule Analyser, Attar Software Ltd, Lancashire, UK) that presents the results in a hierarchic diagram called a knowledge tree. The importance of the different attributes is represented by their position in this diagram. The results show that inductive methods are well suited for analyzing multifactorial and complex relationships in large data sets, and are thus a useful complement to multivariate statistical techniques. The knowledge trees for the two outcomes, dental fear and behavior management problems, were very different from each other, suggesting that the two phenomena are not equivalent. Dental fear was found to be more related to non-dental variables, whereas dental behavior management problems seemed connected to dental variables.

Artificial Intelligence↗

Application of metabolomics to plant genotype discrimination using statistics and machine learning.

MOTIVATION: Metabolomics is a post genomic technology which seeks to provide a comprehensive profile of all the metabolites present in a biological sample. This complements the mRNA profiles provided by microarrays, and the protein profiles provided by proteomics. To test the power of metabolome analysis we selected the problem of discrimating between related genotypes of Arabidopsis. Specifically, the problem tackled was to discrimate between two background genotypes (Col0 and C24) and, more significantly, the offspring produced by the crossbreeding of these two lines, the progeny (whose genotypes would differ only in their maternally inherited mitichondia and chloroplasts). OVERVIEW: A gas chromotography--mass spectrometry (GCMS) profiling protocol was used to identify 433 metabolites in the samples. The metabolomic profiles were compared using descriptive statistics which indicated that key primary metabolites vary more than other metabolites. We then applied neural networks to discriminate between the genotypes. This showed clearly that the two background lines can be discrimated between each other and their progeny, and indicated that the two progeny lines can also be discriminated. We applied Euclidean hierarchical and Principal Component Analysis (PCA) to help understand the basis of genotype discrimination. PCA indicated that malic acid and citrate are the two most important metabolites for discriminating between the background lines, and glucose and fructose are two most important metabolites for discriminating between the crosses. These results are consistant with genotype differences in mitochondia and chloroplasts.

Algorithms↗

An ENSEMBLE machine learning approach for the prediction of all-alpha membrane proteins.

MOTIVATION: All-alpha membrane proteins constitute a functionally relevant subset of the whole proteome. Their content ranges from about 10 to 30% of the cell proteins, based on sequence comparison and specific predictive methods. Due to the paucity of membrane proteins solved with atomic resolution, the training/testing sets of predictive methods for protein topography and topology routinely include very few well-solved structures mixed with a hundred proteins known with low resolution. Moreover, available predictors fail in predicting recently crystallised membrane proteins (Chen et al., 2002). Presently the number of well-solved membrane proteins comprises some 59 chains of low sequence homology. It is therefore possible to train/test predictors only with the set of proteins known with atomic resolution and evaluate more thoroughly the performance of different methods. RESULTS: We implement a cascade-neural network (NN), two different hidden Markov models (HMM), and their ensemble (ENSEMBLE) as a new method. We train and test in cross validation the three methods and ENSEMBLE on the 59 well resolved membrane proteins. ENSEMBLE scores with a per-protein accuracy of 90% for topography and 71% for topology, outperforming the best single method of 7 and 5 percentage points, respectively. When tested on a low resolution set of 151 proteins, with no homology with the 59 proteins, the per-protein accuracy of ENSEMBLE is 76% for topography and 68% for topology. Our results also indicate that the performance of ENSEMBLE is higher than that of the best predictors presently available on the Web.

Algorithms↗

How to interpret an anonymous bacterial genome: machine learning approach to gene identification.

In this report we address the problem of accurate statistical modeling of DNA sequences, either coding or noncoding, for a bacterial species whose genome (or a large portion) was sequenced but not yet characterized experimentally. Availability of these models is critical for successful solution of the genome annotation task by statistical methods of gene finding. We present the method, GeneMark-Genesis, which learns the parameters of Markov models of protein-coding and noncoding regions from anonymous bacterial genomic sequence. These models are subsequently used in the GeneMark and GeneMark.hmm gene-finding programs. Although there is basically one model of a noncoding region for a given genome, several models of protein-coding region are automatically obtained by GeneMark-Genesis. The diversity of protein-coding models reflects the diversity of oligonucleotide compositions, particularly the diversity of codon usage strategies observed in genes from one and the same genome. In the simplest and the most important case, there are just two gene models-typical and atypical ones. We show that the atypical model allows one to predict genes that escape identification by the typical model. Many genes predicted by the atypical model appear to be horizontally transferred genes. The early versions of GeneMark-Genesis were used for annotating the genomes of Methanoccocus jannaschii and Helicobacter pylori. We report the results of accuracy testing of the full-scale version of GeneMark-Genesis on 10 completely sequenced bacterial genomes. Interestingly, the GeneMark.hmm program that employed the typical and atypical models defined by GeneMark-Genesis was able to predict 683 new atypical genes with 176 of them confirmed by similarity search.

Algorithms↗

Machine-learning techniques for macromolecular crystallization data.

Systematizing belief systems regarding macromolecular crystallization has two major advantages: automation and clarification. In this paper, methodologies are presented for systematizing and representing knowledge about the chemical and physical properties of additives used in crystallization experiments. A novel autonomous discovery program is introduced as a method to prune rule-based models produced from crystallization data augmented with such knowledge. Computational experiments indicate that such a system can retain and present informative rules pertaining to protein crystallization that warrant further confirmation via experimental techniques.

Algorithms↗

Estimation of the stapes-bone thickness in the stapedotomy surgical procedure using a machine-learning technique.

Stapedotomy is a surgical procedure aimed at the treatment of hearing impairment due to otosclerosis. The treatment consists of drilling a hole through the stapes bone in the inner ear in order to insert a prosthesis. Safety precautions require knowledge of the nonmeasurable stapes thickness. The technical goal herein has been the design of high-level controls for an intelligent mechatronics drilling tool in order to enable the estimation of stapes thickness from measurable drilling data. The goal has been met by learning a map between drilling features, hence no model of the physical system has been necessary. Learning has been achieved as explained in this paper by a scheme, namely the d-sigma Fuzzy Lattice Neurocomputing (d sigma-FLN) scheme for classification, within the framework of fuzzy lattices. The successful application of the d sigma-FLN scheme is demonstrated in estimating the thickness of a stapes bone "on-line" using drilling data obtained experimentally in the laboratory.

Deafness↗

Gait event detection for FES using accelerometers and supervised machine learning.

Rule based detectors were used with a single cluster of accelerometers attached to the shank for the real time detection of the main phases of normal gait during walking. The gait phase detectors were synthesized from two rule induction algorithms, Rough Sets (RS) and Adaptive Logic Networks (ALNs), and compared with to a previously reported stance/swing detector based on a hand crafted, rule based algorithm. Data was sampled at 100 Hz and the detection errors determined at each sample for 50 steps. For three able bodied subjects, the sample by sample accuracy of stance/swing detection ranged within 94-97%, 87-94%, and 87-95% for the RS, ALN, and the handcrafted methods, respectively. A heuristically formulated postdetector filter improved the RS and ALN detectors' accuracy to 98%. RS and ALN also detected five gait phases to an overall accuracy of 82-89% and 86-91%, respectively. The postdetector filter localized the errors to the phase transitions, but did not change the detection accuracy. The average duration of the error at each transition was 40 ms and 23 ms for RS and ALN, respectively. When implemented on a microcontroller, the RS-based detector executed ten times faster and required one tenth of the memory than the ALN-based detector.

Acceleration↗

Prediction of ultrasound-mediated disruption of cell membranes using machine learning techniques and statistical analysis of acoustic spectra.

Although biological effects of ultrasound must be avoided for safe diagnostic applications, ultrasound's ability to disrupt cell membranes has attracted interest as a method to facilitate drug and gene delivery. This paper seeks to develop "prediction rules" for predicting the degree of cell membrane disruption based on specified ultrasound parameters and measured acoustic signals. Three techniques for generating prediction rules (regression analysis, classification trees and discriminant analysis) are applied to data obtained from a sequence of experiments on bovine red blood cells. For each experiment, the data consist of four ultrasound parameters, acoustic measurements at 400 frequencies, and a measure of cell membrane disruption. To avoid over-training, various combinations of the 404 predictor variables are used when applying the rule generation methods. The results indicate that the variable combination consisting of ultrasound exposure time and acoustic signals measured at the driving frequency and its higher harmonics yields the best rule for all three rule generation methods. The methods used for deriving the prediction rules are broadly applicable, and could be used to develop prediciton rules in other scenarios involving different cell types or tissues. These rules and the methods used to derive them could be used for real-time feedback about ultrasound's biological effects.

Animals↗

Dynamic probability estimator for machine learning.

An efficient algorithm for dynamic estimation of probabilities without division on unlimited number of input data is presented. The method estimates probabilities of the sampled data from the raw sample count, while keeping the total count value constant. Accuracy of the estimate depends on the counter size, rather than on the total number of data points. Estimator follows variations of the incoming data probability within a fixed window size, without explicit implementation of the windowing technique. Total design area is very small and all probabilities are estimated concurrently. Dynamic probability estimator was implemented using a programmable gate array from Xilinx. The performance of this implementation is evaluated in terms of the area efficiency and execution time. This method is suitable for the highly integrated design of artificial neural networks where a large number of dynamic probability estimators can work concurrently.

Artificial Intelligence↗

Evaluating robustness of gait event detection based on machine learning and natural sensors.

A real-time system for deriving timing control for functional electrical stimulation for foot-drop correction, using peripheral nerve activity as a sensor input, was tested for reliability to investigate the potential for clinical use. The system, which was previously reported on, was tested on a hemiplegic subject instrumented with a recording cuff electrode on the Sural nerve, and a stimulation cuff electrode on the Peroneal cuff. Implanted devices enabled recording and stimulation through telelinks. An input domain was derived from the recorded electroneurogram and fed to a detection algorithm based on an adaptive logic network for controlling the stimulation timing. The reliability was tested by letting the subject wear different foot wear and walk on different surfaces than when the training data was recorded. The detection system was also evaluated several months after training. The detection system proved able to successfully detect when walking with different footwear on varying surfaces up to 374 days after training, and thereby showed great potential for being clinically useful.

Action Potentials↗