PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Intensity-based protein identification by machine learning from a library of tandem mass spectra.

Tandem mass spectrometry (MS/MS) has emerged as a cornerstone of proteomics owing in part to robust spectral interpretation algorithms. Widely used algorithms do not fully exploit the intensity patterns present in mass spectra. Here, we demonstrate that intensity pattern modeling improves peptide and protein identification from MS/MS spectra. We modeled fragment ion intensities using a machine-learning approach that estimates the likelihood of observed intensities given peptide and fragment attributes. From 1,000,000 spectra, we chose 27,000 with high-quality, nonredundant matches as training data. Using the same 27,000 spectra, intensity was similarly modeled with mismatched peptides. We used these two probabilistic models to compute the relative likelihood of an observed spectrum given that a candidate peptide is matched or mismatched. We used a 'decoy' proteome approach to estimate incorrect match frequency, and demonstrated that an intensity-based method reduces peptide identification error by 50-96% without any loss in sensitivity.

Algorithms↗

Machine learning can improve prediction of severity in acute pancreatitis using admission values of APACHE II score and C-reactive protein.

BACKGROUND: Acute pancreatitis (AP) has a variable course. Accurate early prediction of severity is essential to direct clinical care. Current assessment tools are inaccurate, and unable to adapt to new parameters. None of the current systems uses C-reactive protein (CRP). Modern machine-learning tools can address these issues. METHODS: 370 patients admitted with AP in a 5-year period were retrospectively assessed; after exclusions, 265 patients were studied. First recorded values for physical examination and blood tests, aetiology, severity and complications were recorded. A kernel logistic regression model was used to remove redundant features, and identify the relationships between relevant features and outcome. Bootstrapping was used to make the best use of data and obtain confidence estimates on the parameters of the model. RESULTS: A model containing 8 variables (age, CRP, respiratory rate, pO2 on air, arterial pH, serum creatinine, white cell count and GCS) predicted a severe attack with an area under the receiver-operating characteristic curve (AUC) of 0.82 (SD 0.01). The optimum cut-off value for predicting severity gave sensitivity and specificity of 0.87 and 0.71 respectively. The predictions were significantly better (p = 0.0036) than admission APACHE II scores in the same patients (AUC 0.74) and better than historical admission APACHE II data (AUC 0.68-0.75). CONCLUSIONS: This system for the first time combines admission values of selected components of APACHE II and CRP for prediction of severe AP. The score is simple to use, and is more accurate than admission APACHE II alone. It is adaptable and would allow incorporation of new predictive factors.

APACHE↗

In silico prediction of pregnane X receptor activators by machine learning approaches.

Pregnane X receptor (PXR) regulates drug metabolism and is involved in drug-drug interactions. Prediction of PXR activators is important for evaluating drug metabolism and toxicity. Computational pharmacophore and quantitative structure-activity relationship models have been developed for predicting PXR activators. Because of the structural diversity of PXR activators, more efforts are needed for exploring methods applicable to a broader spectrum of compounds. We explored three machine learning methods (MLMs) for predicting PXR activators, which were trained and tested by using significantly higher number of compounds, 128 PXR activators (98 human) and 77 PXR non-activators, than those of previous studies. The recursive feature-selection method was used to select molecular descriptors relevant to PXR activator prediction, which are consistent with conclusions from other computational and structural studies. In a 10-fold cross-validation test, our MLM systems correctly predicted 81.2 to 84.0% of PXR activators, 80.8 to 85.0% of hPXR activators, 61.2 to 70.3% of PXR nonactivators, and 67.7 to 73.6% of hPXR nonactivators. Our systems also correctly predicted 73.3 to 86.7% of 15 newly published hPXR activators. MLMs seem to be useful for predicting PXR activators and for providing clues to physicochemical features of PXR activation.

Artificial Intelligence↗

Machine learning and word sense disambiguation in the biomedical domain: design and evaluation issues.

BACKGROUND: Word sense disambiguation (WSD) is critical in the biomedical domain for improving the precision of natural language processing (NLP), text mining, and information retrieval systems because ambiguous words negatively impact accurate access to literature containing biomolecular entities, such as genes, proteins, cells, diseases, and other important entities. Automated techniques have been developed that address the WSD problem for a number of text processing situations, but the problem is still a challenging one. Supervised WSD machine learning (ML) methods have been applied in the biomedical domain and have shown promising results, but the results typically incorporate a number of confounding factors, and it is problematic to truly understand the effectiveness and generalizability of the methods because these factors interact with each other and affect the final results. Thus, there is a need to explicitly address the factors and to systematically quantify their effects on performance. RESULTS: Experiments were designed to measure the effect of "sample size" (i.e. size of the datasets), "sense distribution" (i.e. the distribution of the different meanings of the ambiguous word) and "degree of difficulty" (i.e. the measure of the distances between the meanings of the senses of an ambiguous word) on the performance of WSD classifiers. Support Vector Machine (SVM) classifiers were applied to an automatically generated data set containing four ambiguous biomedical abbreviations: BPD, BSA, PCA, and RSV, which were chosen because of varying degrees of differences in their respective senses. Results showed that: 1) increasing the sample size generally reduced the error rate, but this was limited mainly to well-separated senses (i.e. cases where the distances between the senses were large); in difficult cases an unusually large increase in sample size was needed to increase performance slightly, which was impractical, 2) the sense distribution did not have an effect on performance when the senses were separable, 3) when there was a majority sense of over 90%, the WSD classifier was not better than use of the simple majority sense, 4) error rates were proportional to the similarity of senses, and 5) there was no statistical difference between results when using a 5-fold or 10-fold cross-validation method. Other issues that impact performance are also enumerated. CONCLUSION: Several different independent aspects affect performance when using ML techniques for WSD. We found that combining them into one single result obscures understanding of the underlying methods. Although we studied only four abbreviations, we utilized a well-established statistical method that guarantees the results are likely to be generalizable for abbreviations with similar characteristics. The results of our experiments show that in order to understand the performance of these ML methods it is critical that papers report on the baseline performance, the distribution and sample size of the senses in the datasets, and the standard deviation or confidence intervals. In addition, papers should also characterize the difficulty of the WSD task, the WSD situations addressed and not addressed, as well as the ML methods and features used. This should lead to an improved understanding of the generalizablility and the limitations of the methodology.

Algorithms↗

Machine learning for development of an expert system to predict premature birth.

Normal pregnancy involves a term of 40 weeks gestation. Problems associated with low birthweight and prematurity continue to plague childbearing families and the healthcare system because 8-12% of all newborns in the United States deliver prior to 37 weeks gestation. The high cost of caring for premature babies increasingly treats all pregnant women as if they are 'high risk' for preterm birth. Artificial intelligence techniques used a machine learning program named LERS1 with large datasets (n = 18,890; 214 variables), statistical analysis, expert verification techniques, and a prototype expert system2 that yielded improved accuracy (53-90%) over existing manual techniques (17-38%) for predicting preterm birth.

Adult↗

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article↗

Building an allergens ontology and maintaining it using machine learning techniques.

Ontologies are widely used for formalizing and organizing the knowledge of a particular domain of interest. This facilitates knowledge sharing and re-use by both people and systems. Ontologies are becoming increasingly important in the biomedical domain since they enable knowledge sharing in a formal, homogeneous and unambiguous way. Knowledge in a rapidly growing field such as biomedicine is usually evolving and therefore an ontology maintenance process is required to keep ontological knowledge up-to-date. This work presents our methodology for building a formally defined ontology, maintaining it exploiting machine learning techniques and domain specific corpora, and evaluating it using a well-defined experimental setting. The application of this methodology in the allergen domain is then discussed in detail presenting the ontology built, the specific techniques used and the evaluation settings.

Allergens↗

Predicting the phosphorylation sites using hidden Markov models and machine learning methods.

Accurately predicting phosphorylation sites in proteins is an important issue in postgenomics, for which how to efficiently extract the most predictive features from amino acid sequences for modeling is still challenging. Although both the distributed encoding method and the bio-basis function method work well, they still have some limits in use. The distributed encoding method is unable to code the biological content in sequences efficiently, whereas the bio-basis function method is a nonparametric method, which is often computationally expensive. As hidden Markov models (HMMs) can be used to generate one model for one cluster of aligned protein sequences, the aim in this study is to use HMMs to extract features from amino acid sequences, where sequence clusters are determined using available biological knowledge. In this novel method, HMMs are first constructed using functional sequences only. Both functional and nonfunctional training sequences are then inputted into the trained HMMs to generate functional and nonfunctional feature vectors. From this, a machine learning algorithm is used to construct a classifier based on these feature vectors. It is found in this work that (1) this method provides much better prediction accuracy than the use of HMMs only for prediction, and (2) the support vector machines (SVMs) algorithm outperforms decision trees and neural network algorithms when they are constructed on the features extracted using the trained HMMs.

Algorithms↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

Evaluating the C-section rate of different physician practices: using machine learning to model standard practice.

The C-section rate of a population of 22,175 expectant mothers is 16.8%; yet the 17 physician groups that serve this population have vastly different group C-section rates, ranging from 13% to 23%. Our goal is to determine retrospectively if the variations in the observed rates can be attributed to variations in the intrinsic risk of the patient sub-populations (i.e. some groups contain more "high-risk C-section" patients), or differences in physician practice (i.e. some groups do more C-sections). We apply machine learning to this problem by training models to predict standard practice from retrospective data. We then use the models of standard practice to evaluate the C-section rate of each physician practice. Our results indicate that although there is variation in intrinsic risk among the groups, there also is much variation in physician practice.

Artificial Intelligence↗

Machine learning techniques in disease forecasting: a case study on rice blast prediction.

BACKGROUND: Diverse modeling approaches viz. neural networks and multiple regression have been followed to date for disease prediction in plant populations. However, due to their inability to predict value of unknown data points and longer training times, there is need for exploiting new prediction softwares for better understanding of plant-pathogen-environment relationships. Further, there is no online tool available which can help the plant researchers or farmers in timely application of control measures. This paper introduces a new prediction approach based on support vector machines for developing weather-based prediction models of plant diseases. RESULTS: Six significant weather variables were selected as predictor variables. Two series of models (cross-location and cross-year) were developed and validated using a five-fold cross validation procedure. For cross-year models, the conventional multiple regression (REG) approach achieved an average correlation coefficient (r) of 0.50, which increased to 0.60 and percent mean absolute error (%MAE) decreased from 65.42 to 52.24 when back-propagation neural network (BPNN) was used. With generalized regression neural network (GRNN), the r increased to 0.70 and %MAE also improved to 46.30, which further increased to r = 0.77 and %MAE = 36.66 when support vector machine (SVM) based method was used. Similarly, cross-location validation achieved r = 0.48, 0.56 and 0.66 using REG, BPNN and GRNN respectively, with their corresponding %MAE as 77.54, 66.11 and 58.26. The SVM-based method outperformed all the three approaches by further increasing r to 0.74 with improvement in %MAE to 44.12. Overall, this SVM-based prediction approach will open new vistas in the area of forecasting plant diseases of various crops. CONCLUSION: Our case study demonstrated that SVM is better than existing machine learning techniques and conventional REG approaches in forecasting plant diseases. In this direction, we have also developed a SVM-based web server for rice blast prediction, a first of its kind worldwide, which can help the plant science community and farmers in their decision making process. The server is freely available at http://www.imtech.res.in/raghava/rbpred/.

Agriculture↗

Applying machine learning methods for finding significant amino acid properties in proteins.

There are several possibilities for definition and derivation of sequence patterns associated with structural motifs, in particular on the secondary structure level which may be used to predict these structure elements. Sequence patterns consist of a number of consecutive positions along the polypeptide chain from which a certain quantity is specified. One of the important factors in deriving sequence patterns in terms of amino acid properties is how to find the most characteristic properties to specify a certain position and thus to avoid redundant physical information. We have applied machine learning methods to select the most significant amino acid properties describing a structurally determined sequence position. Results are given for the beginning of alpha-helices. These methods may link the gap between amino acid patterns and property patterns and thus are valuable to improve protein structure prediction.

Amino Acids↗

Machine learning techniques to diagnose breast cancer from image-processed nuclear features of fine needle aspirates.

An interactive computer system evaluates and diagnoses based on cytologic features derived directly from a digital scan of fine-needle aspirate (FNA) slides. A consecutive series of 569 patients provided the data to develop the system and an additional 54 consecutive, new patients provided samples to test the system. The projected prospective accuracy of the system estimated by tenfold cross validation was 97%. The actual accuracy on 54 new samples (36 benign, 1 atypia, and 17 malignant) was 100%. Digital image analysis coupled with machine learning techniques will improve diagnostic accuracy of breast fine needle aspirates.

Biopsy, Needle↗

Metabolomics, machine learning and modelling: towards an understanding of the language of cells.

In answering the question 'Systems Biology--will it work?' (which it self-evidently has already), it is appropriate to highlight advances in philosophy, in new technique development and in novel findings. In terms of philosophy, we see that systems biology involves an iterative interplay between linked activities--instance, between theory and experiment, between induction and deduction and between measurements of parameters and variables--with more emphasis than has perhaps been common now being focused on the first in each of these pairs. In technique development, we highlight closed loop machine learning and its use in the optimization of scientific instrumentation, and the ability to effect high-quality and quasi-continuous optical images of cells. This leads to many important and novel findings. In the first case, these may involve new biomarkers for disease, whereas in the second case, we have determined that many biological signals may be frequency-rather than amplitude-encoded. This leads to a very different view of how signalling 'works' (equations such as that of Michaelis and Menten which use only amplitudes, i.e. concentrations, are inadequate descriptors), lays emphasis on the signal processing network elements that lie 'downstream' of what are traditionally considered the signals, and allows one simply to understand how cross-talk may be avoided between pathways which nevertheless use common signalling elements. The language of cells is much richer than we had supposed, and we are now well placed to decode it.

Cell Physiological Phenomena↗

Symbolic, neural, and Bayesian machine learning models for predicting carcinogenicity of chemical compounds.

Experimental programs have been underway for several years to determine the environmental effects of chemical compounds, mixtures, and the like. Among these programs is the National Toxicology Program (NTP) on rodent carcinogenicity. Because these experiments are costly and time-consuming, the rate at which test articles (i.e., chemicals) can be tested is limited. The ability to predict the outcome of the analysis at various points in the process would facilitate informed decisions about the allocation of testing resources. To assist human experts in organizing an empirical testing regime, and to try to shed light on mechanisms of toxicity, we constructed toxicity models using various machine learning and data mining methods, both existing and those of our own devising. These models took the form of decision trees, rule sets, neural networks, rules extracted from trained neural networks, and Bayesian classifiers. As a training set, we used recent results from rodent carcinogenicity bioassays conducted by the NTP on 226 test articles. We performed 10-way cross-validation on each of our models to approximate their expected error rates on unseen data. The data set consists of physical-chemical parameters of test articles, alerting chemical substructures, salmonella mutagenicity assay results, subchronic histopathology data, and information on route, strain, and sex/species for 744 individual experiments. These results contribute to the ongoing process of evaluating and interpreting the data collected from chemical toxicity studies.

Animals↗

Evaluating machine learning approaches for aiding probe selection for gene-expression arrays.

MOTIVATION: Microarrays are a fast and cost-effective method of performing thousands of DNA hybridization experiments simultaneously. DNA probes are typically used to measure the expression level of specific genes. Because probes greatly vary in the quality of their hybridizations, choosing good probes is a difficult task. If one could accurately choose probes that are likely to hybridize well, then fewer probes would be needed to represent each gene in a gene-expression microarray, and, hence, more genes could be placed on an array of a given physical size. Our goal is to empirically evaluate how successfully three standard machine-learning algorithms-naïve Bayes, decision trees, and artificial neural networks-can be applied to the task of predicting good probes. Fortunately it is relatively easy to get training examples for such a learning task: place various probes on a gene chip, add a sample where the corresponding genes are highly expressed, and then record how well each probe measures the presence of its corresponding gene. With such training examples, it is possible that an accurate predictor of probe quality can be learned. RESULTS: Two of the learning algorithms we investigate-naïve Bayes and neural networks-learn to predict probe quality surprisingly well. For example, in the top ten predicted probes for a given gene not used for training, on average about five rank in the top 2.5% of that gene's hundreds of possible probes. Decision-tree induction and the simple approach of using predicted melting temperature to rank probes perform significantly worse than these two algorithms. The features we use to represent probes are very easily computed and the time taken to score each candidate probe after training is minor. Training the naïve Bayes algorithm takes very little time, and while it takes over 10 times as long to train a neural network, that time is still not very substantial (on the order of a few hours on a desktop workstation). We also report the information contained in the features we use to describe the probes. We find the fraction of cytosine in the probe to be the most informative feature. We also find, not surprisingly, that the nucleotides in the middle of the probes sequence are more informative than those at the ends of the sequence.

Algorithms↗

An integrated machine learning system to computationally screen protein databases for protein binding peptide ligands.

A fairly large set of protein interactions is mediated by families of peptide binding domains, such as Src homology 2 (SH2), SH3, PDZ, major histocompatibility complex, etc. To identify their ligands by experimental screening is not only labor-intensive but almost futile in screening low abundance species due to the suppression by high abundance species. An ideal way of studying protein-protein interactions is to use high throughput computational approaches to screen protein sequence databases to direct the validating experiments toward the most promising peptides. Predictors with only good cross-validation were not good enough to screen protein databases. In the current study we built integrated machine learning systems using three novel coding methods and screened the Swiss-Prot and GenBank protein databases for potential ligands of 10 SH3 and three PDZ domains. A large fraction of predictions has already been experimentally confirmed by other independent research groups, indicating a satisfying generalization capability for future applications in identifying protein interactions.

Amino Acid Motifs↗

Classifying spatial patterns of brain activity with machine learning methods: application to lie detection.

Patterns of brain activity during deception have recently been characterized with fMRI on the multi-subject average group level. The clinical value of fMRI in lie detection will be determined by the ability to detect deception in individual subjects, rather than group averages. High-dimensional non-linear pattern classification methods applied to functional magnetic resonance (fMRI) images were used to discriminate between the spatial patterns of brain activity associated with lie and truth. In 22 participants performing a forced-choice deception task, 99% of the true and false responses were discriminated correctly. Predictive accuracy, assessed by cross-validation in participants not included in training, was 88%. The results demonstrate the potential of non-linear machine learning techniques in lie detection and other possible clinical applications of fMRI in individual subjects, and indicate that accurate clinical tests could be based on measurements of brain function with fMRI.

Adult↗