PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

A machine learning approach to the analysis of time-frequency maps, and its application to neural dynamics.

The statistical analysis of experimentally recorded brain activity patterns may require comparisons between large sets of complex signals in order to find meaningful similarities and differences between signals with large variability. High-level representations such as time-frequency maps convey a wealth of useful information, but they involve a large number of parameters that make statistical investigations of many signals difficult at present. In this paper, we describe a method that performs drastic reduction in the complexity of time-frequency representations through a modelling of the maps by elementary functions. The method is validated on artificial signals and subsequently applied to electrophysiological brain signals (local field potential) recorded from the olfactory bulb of rats while they are trained to recognize odours. From hundreds of experimental recordings, reproducible time-frequency events are detected, and relevant features are extracted, which allow further information processing, such as automatic classification.

Algorithms↗

On the optimization of classes for the assignment of unidentified reading frames in functional genomics programmes: the need for machine learning.

At present, the assignment of function to novel genes uncovered by the systematic genome-sequencing programmes is a problem. Many studies anticipate that this can be achieved by analysing patterns of gene expression via the transcriptome, proteome and metabolome. Thus, functional genomics is, in part, an exercise in pattern classification. Because many genes have known functional classes, the problem of predicting their functional class is a supervised learning problem. However, most pattern classification methods that have been applied to the problem have been unsupervised clustering methods. Consequently, the best classification tools have not always been used. Furthermore, the present functional classes are suboptimal and new unsupervised clustering methods are needed to improve them. Better-structured functional classes will facilitate the prediction of biochemically testable functions.

Animals↗

Relating clinical and neurophysiological assessment of spasticity by machine learning.

Spasticity following spinal cord injury (SCI) is most often assessed clinically using a five-point Ashworth score (AS). A more objective assessment of altered motor control may be achieved by using a comprehensive protocol based on a surface electromyographic (sEMG) activity recorded from thigh and leg muscles. However, the relationship between the clinical and neurophysiological assessments is still unknown. In this paper we employ three different classification methods to investigate this relationship. The experimental results indicate that, if the appropriate set of sEMG features is used, the neurophysiological assessment is related to clinical findings and can be used to predict the AS. A comprehensive sEMG assessment may be proven useful as an objective method of evaluating the effectiveness of various interventions and for follow-up of SCI patients.

Artificial Intelligence↗

Predicting protein-ligand binding affinities using novel geometrical descriptors and machine-learning methods.

Inspired by the concept of knowledge-based scoring functions, a new quantitative structure-activity relationship (QSAR) approach is introduced for scoring protein-ligand interactions. This approach considers that the strength of ligand binding is correlated with the nature of specific ligand/binding site atom pairs in a distance-dependent manner. In this technique, atom pair occurrence and distance-dependent atom pair features are used to generate an interaction score. Scoring and pattern recognition results obtained using Kernel PLS (partial least squares) modeling and a genetic algorithm-based feature selection method are discussed.

Algorithms↗

Application of machine learning to improve the results of high-throughput docking against the HIV-1 protease.

We have previously reported that the application of a Laplacian-modified naive Bayesian (NB) classifier may be used to improve the ranking of known inhibitors from a random database of compounds after High-Throughput Docking (HTD). The method relies upon the frequency of substructural features among the active and inactive compounds from 2D fingerprint information of the compounds. Here we present an investigation of the role of extended connectivity fingerprints in training the NB classifier against HTD studies on the HIV-1 protease using three docking programs: Glide, FlexX, and GOLD. The results show that the performance of the NB classifier is due to the presence of a large number of features common to the set of known active compounds rather than a single structural or substructural scaffold. We demonstrate that the Laplacian-modified naive Bayesian classifier trained with data from high-throughput docking is superior at identifying active compounds from a target database in comparison to conventional two-dimensional substructure search methods alone.

Algorithms↗

New methods for ligand-based virtual screening: use of data fusion and machine learning to enhance the effectiveness of similarity searching.

Similarity searching using a single bioactive reference structure is a well-established technique for accessing chemical structure databases. This paper describes two extensions of the basic approach. First, we discuss the use of group fusion to combine the results of similarity searches when multiple reference structures are available. We demonstrate that this technique is notably more effective than conventional similarity searching in scaffold-hopping searches for structurally diverse sets of active molecules; conversely, the technique will do little to improve the search performance if the actives are structurally homogeneous. Second, we make the assumption that the nearest neighbors resulting from a similarity search, using a single bioactive reference structure, are also active and use this assumption to implement approximate forms of group fusion, substructural analysis, and binary kernel discrimination. This approach, called turbo similarity searching, is notably more effective than conventional similarity searching.

Artificial Intelligence↗

Modeling of human cytochrome p450-mediated drug metabolism using unsupervised machine learning approach.

We developed a computational algorithm for evaluating the possibility of cytochrome P450-mediated metabolic transformations that xenobiotics molecules undergo in the human body. First, we compiled a database of known human cytochrome P-450 substrates, products, and nonsubstrates for 38 enzyme-specific groups (total of 2200 compounds). Second, we determined the cytochrome-mediated metabolic reactions most typical for each group and examined the substrates and products of these reactions. To assess the probability of P450 transformations of novel compounds, we built a nonlinear quantitative structure-metabolism relationships (QSMR) model based on Kohonen self-organizing maps (SOM). This neural network QSMR model incorporated a predefined set of physicochemical descriptors encoding the key molecular properties that define the metabolic fate of individual molecules. Isozyme-specific groups of substrate molecules were visualized, thus facilitating prediction of tissue-specific metabolism. The developed algorithm can be used in early stages of drug discovery as an efficient tool for the assessment of human metabolism and toxicity of novel compounds in designing discovery libraries and in lead optimization.

Algorithms↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Rapid identification of closely related muscle foods by vibrational spectroscopy and machine learning.

Muscle foods are an integral part of the human diet and during the last few decades consumption of poultry products in particular has increased significantly. It is important for consumers, retailers and food regulatory bodies that these products are of a consistently high quality, authentic, and have not been subjected to adulteration by any lower-grade material either by accident or for economic gain. A variety of methods have been developed for the identification and authentication of muscle foods. However, none of these are rapid or non-invasive, all are time-consuming and difficulties have been encountered in discriminating between the commercially important avian species. Whilst previous attempts have been made to discriminate between muscle foods using infrared spectroscopy, these have had limited success, in particular regarding the closely related poultry species, chicken and turkey. Moreover, this study includes novel data since no attempts have been made to discriminate between both the species and the distinct muscle groups within these species, and this is the first application of Raman spectroscopy to the study of muscle foods. Samples of pre-packed meat and poultry were acquired and FT-IR and Raman measurements taken directly from the meat surface. Qualitative interpretation of FT-IR and Raman spectra at the species and muscle group levels were possible using discriminant function analysis. Genetic algorithms were used to elucidate meaningful interpretation of FT-IR results in (bio)chemical terms and we show that specific wavenumbers, and therefore chemical species, were discriminatory for each type (species and muscle) of poultry sample. We believe that this approach would aid food regulatory bodies in the rapid identification of meat and poultry products and shows particular potential for rapid assessment of food adulteration.

Algorithms↗

Structure-activity relationships derived by machine learning: the use of atoms and their bond connectivities to predict mutagenicity by inductive logic programming.

We present a general approach to forming structure-activity relationships (SARs). This approach is based on representing chemical structure by atoms and their bond connectivities in combination with the inductive logic programming (ILP) algorithm PROGOL. Existing SAR methods describe chemical structure by using attributes which are general properties of an object. It is not possible to map chemical structure directly to attribute-based descriptions, as such descriptions have no internal organization. A more natural and general way to describe chemical structure is to use a relational description, where the internal construction of the description maps that of the object described. Our atom and bond connectivities representation is a relational description. ILP algorithms can form SARs with relational descriptions. We have tested the relational approach by investigating the SARs of 230 aromatic and heteroaromatic nitro compounds. These compounds had been split previously into two subsets, 188 compounds that were amenable to regression and 42 that were not. For the 188 compounds, a SAR was found that was as accurate as the best statistical or neural network-generated SARs. The PROGOL SAR has the advantages that it did not need the use of any indicator variables handcrafted by an expert, and the generated rules were easily comprehensible. For the 42 compounds, PROGOL formed a SAR that was significantly (P < 0.025) more accurate than linear regression, quadratic regression, and back-propagation. This SAR is based on an automatically generated structural alert for mutagenicity.

Algorithms↗

Machine learning methods applied on dental fear and behavior management problems in children.

The etiologies of dental fear and dental behavior management problems in children were investigated in a database of information on 2,257 Swedish children 4-6 and 9-11 years old. The analyses were performed using computerized inductive techniques within the field of artificial intelligence. The database held information regarding dental fear levels and behavior management problems, which were defined as outcomes, i.e. dependent variables. The attributes, i.e. independent variables, included data on dental health and dental treatments, information about parental dental fear, general anxiety, socioeconomic variables, etc. The data contained both numerical and discrete variables. The analyses were performed using an inductive analysis program (XpertRule Analyser, Attar Software Ltd, Lancashire, UK) that presents the results in a hierarchic diagram called a knowledge tree. The importance of the different attributes is represented by their position in this diagram. The results show that inductive methods are well suited for analyzing multifactorial and complex relationships in large data sets, and are thus a useful complement to multivariate statistical techniques. The knowledge trees for the two outcomes, dental fear and behavior management problems, were very different from each other, suggesting that the two phenomena are not equivalent. Dental fear was found to be more related to non-dental variables, whereas dental behavior management problems seemed connected to dental variables.

Artificial Intelligence↗

Application of metabolomics to plant genotype discrimination using statistics and machine learning.

MOTIVATION: Metabolomics is a post genomic technology which seeks to provide a comprehensive profile of all the metabolites present in a biological sample. This complements the mRNA profiles provided by microarrays, and the protein profiles provided by proteomics. To test the power of metabolome analysis we selected the problem of discrimating between related genotypes of Arabidopsis. Specifically, the problem tackled was to discrimate between two background genotypes (Col0 and C24) and, more significantly, the offspring produced by the crossbreeding of these two lines, the progeny (whose genotypes would differ only in their maternally inherited mitichondia and chloroplasts). OVERVIEW: A gas chromotography--mass spectrometry (GCMS) profiling protocol was used to identify 433 metabolites in the samples. The metabolomic profiles were compared using descriptive statistics which indicated that key primary metabolites vary more than other metabolites. We then applied neural networks to discriminate between the genotypes. This showed clearly that the two background lines can be discrimated between each other and their progeny, and indicated that the two progeny lines can also be discriminated. We applied Euclidean hierarchical and Principal Component Analysis (PCA) to help understand the basis of genotype discrimination. PCA indicated that malic acid and citrate are the two most important metabolites for discriminating between the background lines, and glucose and fructose are two most important metabolites for discriminating between the crosses. These results are consistant with genotype differences in mitochondia and chloroplasts.

Algorithms↗

An ENSEMBLE machine learning approach for the prediction of all-alpha membrane proteins.

MOTIVATION: All-alpha membrane proteins constitute a functionally relevant subset of the whole proteome. Their content ranges from about 10 to 30% of the cell proteins, based on sequence comparison and specific predictive methods. Due to the paucity of membrane proteins solved with atomic resolution, the training/testing sets of predictive methods for protein topography and topology routinely include very few well-solved structures mixed with a hundred proteins known with low resolution. Moreover, available predictors fail in predicting recently crystallised membrane proteins (Chen et al., 2002). Presently the number of well-solved membrane proteins comprises some 59 chains of low sequence homology. It is therefore possible to train/test predictors only with the set of proteins known with atomic resolution and evaluate more thoroughly the performance of different methods. RESULTS: We implement a cascade-neural network (NN), two different hidden Markov models (HMM), and their ensemble (ENSEMBLE) as a new method. We train and test in cross validation the three methods and ENSEMBLE on the 59 well resolved membrane proteins. ENSEMBLE scores with a per-protein accuracy of 90% for topography and 71% for topology, outperforming the best single method of 7 and 5 percentage points, respectively. When tested on a low resolution set of 151 proteins, with no homology with the 59 proteins, the per-protein accuracy of ENSEMBLE is 76% for topography and 68% for topology. Our results also indicate that the performance of ENSEMBLE is higher than that of the best predictors presently available on the Web.

Algorithms↗