PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

Cross-Platform Proteomics and Machine Learning Algorithms Nominate Plasma Biomarkers of Stroke Diagnosis.

BACKGROUND: Blood-based biomarkers for stroke subtyping could improve triage in emergency settings. We used cross-platform proteomics to identify plasma biomarkers differentiating major stroke diagnostic groups. METHODS: We conducted a case-control study using 2 biorepositories. Plasma was collected in the emergency department from adults with suspected stroke before therapeutic intervention. Differentially enriched proteins were identified across acute ischemic stroke, intracerebral hemorrhage, transient ischemic attack, and stroke mimics using SomaScan discovery proteomics (Grady). Differentially enriched proteins were nominated using pairwise and multigroup comparisons and adjusted for clinical covariates. Protein panels were created using least absolute shrinkage and selection operator logistic regression. Internal validation used repeated nested cross-validation (rCV) and targeted mass spectrometry (MS), while external validation used data-independent acquisition &#xa0;mass spectrometry in an independent cohort (Yale). RESULTS: We included 100 subjects (40 with acute ischemic stroke, 20 with intracerebral hemorrhage, 20 with transient ischemic attack, 20 with stroke mimics) in discovery and 80 subjects (20 per group) in external validation cohorts. SomaScan quantified 7307 proteins, of which 61 differentiated stroke subtypes. We identified 7 protein classifiers for acute ischemic stroke (rCV-area under the curve, 0.82 [95% CI, 0.78-0.86]), 6 for intracerebral hemorrhage (rCV-area under the curve, 0.70 [95% CI, 0.64-0.76]), 8 for transient ischemic attack (rCV-area under the curve, 0.78 [95% CI, 0.73-0.84]), and 7 for stroke mimics (rCV-area under the curve, 0.81 [95% CI, 0.77-0.86]). Targeted proteomics internally validated 11 proteins, and data-independent acquisition-mass spectrometry externally validated 32 proteins, including VTN (vitronectin), PLG (plasminogen), and S100A9 as top stroke mimics, transient ischemic attack, and intracerebral hemorrhage classifiers. CONCLUSIONS: This study highlights plasma proteomics as a valuable tool for discovering protein biomarkers of stroke diagnosis. These findings support further validation in larger, multicenter cohorts to facilitate biomarker-guided stroke diagnosis in acute care.

Humans↗

The cerebellum: a neuronal learning machine?

Comparison of two seemingly quite different behaviors yields a surprisingly consistent picture of the role of the cerebellum in motor learning. Behavioral and physiological data about classical conditioning of the eyelid response and motor learning in the vestibulo-ocular reflex suggests that (i) plasticity is distributed between the cerebellar cortex and the deep cerebellar nuclei; (ii) the cerebellar cortex plays a special role in learning the timing of movement; and (iii) the cerebellar cortex guides learning in the deep nuclei, which may allow learning to be transferred from the cortex to the deep nuclei. Because many of the similarities in the data from the two systems typify general features of cerebellar organization, the cerebellar mechanisms of learning in these two systems may represent principles that apply to many motor systems.

Animals↗

Morphological classification of brains via high-dimensional shape transformations and machine learning methods.

A high-dimensional shape transformation posed in a mass-preserving framework is used as a morphological signature of a brain image. Population differences with complex spatial patterns are then determined by applying a nonlinear support vector machine (SVM) pattern classification method to the morphological signatures. Significant reduction of the dimensionality of the morphological signatures is achieved via wavelet decomposition and feature reduction methods. Applying the method to MR images with simulated atrophy shows that the method can correctly detect subtle and spatially complex atrophy, even when the simulated atrophy represents only a 5% variation from the original image. Applying this method to actual MR images shows that brains can be correctly determined to be male or female with a successful classification rate of 97%, using the leave-one-out method. This proposed method also shows a high classification rate for old adults' age classification, even under difficult test scenarios. The main characteristic of the proposed methodology is that, by applying multivariate pattern classification methods, it can detect subtle and spatially complex patterns of morphological group differences which are often not detectable by voxel-based morphometric methods, because these methods analyze morphological measurements voxel-by-voxel and do not consider the entirety of the data simultaneously.

Aged↗

Prediction of beta-turns with learning machines.

The support vector machine approach was introduced to predict the beta-turns in proteins. The overall self-consistency rate by the re-substitution test for the training or learning dataset reached 100%. Both the training dataset and independent testing dataset were taken from Chou [J. Pept. Res. 49 (1997) 120]. The success prediction rates by the jackknife test for the beta-turn subset of 455 tetrapeptides and non-beta-turn subset of 3807 tetrapeptides in the training dataset were 58.1 and 98.4%, respectively. The success rates with the independent dataset test for the beta-turn subset of 110 tetrapeptides and non-beta-turn subset of 30,231 tetrapeptides were 69.1 and 97.3%, respectively. The results obtained from this study support the conclusion that the residue-coupled effect along a tetrapeptide is important for the formation of a beta-turn.

Artificial Intelligence↗

Support vector machine learning model for the prediction of sentinel node status in patients with cutaneous melanoma.

BACKGROUND: Currently, approximately 80% of melanoma patients undergoing sentinel node biopsy (SNB) have negative sentinel lymph nodes (SLNs), and no prediction system is reliable enough to be implemented in the clinical setting to reduce the number of SNB procedures. In this study, the predictive power of support vector machine (SVM)-based statistical analysis was tested. METHODS: The clinical records of 246 patients who underwent SNB at our institution were used for this analysis. The following clinicopathologic variables were considered: the patient's age and sex and the tumor's histological subtype, Breslow thickness, Clark level, ulceration, mitotic index, lymphocyte infiltration, regression, angiolymphatic invasion, microsatellitosis, and growth phase. The results of SVM-based prediction of SLN status were compared with those achieved with logistic regression. RESULTS: The SLN positivity rate was 22% (52 of 234). When the accuracy was > or = 80%, the negative predictive value, positive predictive value, specificity, and sensitivity were 98%, 54%, 94%, and 77% and 82%, 41%, 69%, and 93% by using SVM and logistic regression, respectively. Moreover, SVM and logistic regression were associated with a diagnostic error and an SNB percentage reduction of (1) 1% and 60% and (2) 15% and 73%, respectively. CONCLUSIONS: The results from this pilot study suggest that SVM-based prediction of SLN status might be evaluated as a prognostic method to avoid the SNB procedure in 60% of patients currently eligible, with a very low error rate. If validated in larger series, this strategy would lead to obvious advantages in terms of both patient quality of life and costs for the health care system.

Artificial Intelligence↗

Support vector machine learning from heterogeneous data: an empirical analysis using protein sequence and structure.

MOTIVATION: Drawing inferences from large, heterogeneous sets of biological data requires a theoretical framework that is capable of representing, e.g. DNA and protein sequences, protein structures, microarray expression data, various types of interaction networks, etc. Recently, a class of algorithms known as kernel methods has emerged as a powerful framework for combining diverse types of data. The support vector machine (SVM) algorithm is the most popular kernel method, due to its theoretical underpinnings and strong empirical performance on a wide variety of classification tasks. Furthermore, several recently described extensions allow the SVM to assign relative weights to various datasets, depending upon their utilities in performing a given classification task. RESULTS: In this work, we empirically investigate the performance of the SVM on the task of inferring gene functional annotations from a combination of protein sequence and structure data. Our results suggest that the SVM is quite robust to noise in the input datasets. Consequently, in the presence of only two types of data, an SVM trained from an unweighted combination of datasets performs as well or better than a more sophisticated algorithm that assigns weights to individual data types. Indeed, for this simple case, we can demonstrate empirically that no solution is significantly better than the naive, unweighted average of the two datasets. On the other hand, when multiple noisy datasets are included in the experiment, then the naive approach fares worse than the weighted approach. Our results suggest that for many applications, a naive unweighted sum of kernels may be sufficient. AVAILABILITY: http://noble.gs.washington.edu/proj/seqstruct

Algorithms↗

Protein backbone angle prediction with machine learning approaches.

MOTIVATION: Protein backbone torsion angle prediction provides useful local structural information that goes beyond conventional three-state (alpha, beta and coil) secondary structure predictions. Accurate prediction of protein backbone torsion angles will substantially improve modeling procedures for local structures of protein sequence segments, especially in modeling loop conformations that do not form regular structures as in alpha-helices or beta-strands. RESULTS: We have devised two novel automated methods in protein backbone conformational state prediction: one method is based on support vector machines (SVMs); the other method combines a standard feed-forward back-propagation artificial neural network (NN) with a local structure-based sequence profile database (LSBSP1). Extensive benchmark experiments demonstrate that both methods have improved the prediction accuracy rate over the previously published methods for conformation state prediction when using an alphabet of three or four states. AVAILABILITY: LSBSP1 and the NN algorithm have been implemented in PrISM.1, which is available from www.columbia.edu/~ay1/. SUPPLEMENTARY INFORMATION: Supplementary data for the SVM method can be downloaded from the Website www.cs.columbia.edu/compbio/backbone.

Algorithms↗

Computation of conformational entropy from protein sequences using the machine-learning method--application to the study of the relationship between structural conservation and local structural stability.

A complete protein sequence can usually determine a unique conformation; however, the situation is different for shorter subsequences--some of them are able to adopt unique conformations, independent of context; while others assume diverse conformations in different contexts. The conformations of subsequences are determined by the interplay between local and nonlocal interactions. A quantitative measure of such structural conservation or variability will be useful in the understanding of the sequence-structure relationship. In this report, we developed an approach using the support vector machine method to compute the conformational variability directly from sequences, which is referred to as the sequence structural entropy. As a practical application, we studied the relationship between sequence structural entropy and the hydrogen exchange for a set of well-studied proteins. We found that the slowest exchange cores usually comprise amino acids of the lowest sequence structural entropy. Our results indicate that structural conservation is closely related to the local structural stability. This relationship may have interesting implications in the protein folding processes, and may be useful in the study of the sequence-structure relationship.

Amino Acid Sequence↗

Building a protein name dictionary from full text: a machine learning term extraction approach.

BACKGROUND: The majority of information in the biological literature resides in full text articles, instead of abstracts. Yet, abstracts remain the focus of many publicly available literature data mining tools. Most literature mining tools rely on pre-existing lexicons of biological names, often extracted from curated gene or protein databases. This is a limitation, because such databases have low coverage of the many name variants which are used to refer to biological entities in the literature. RESULTS: We present an approach to recognize named entities in full text. The approach collects high frequency terms in an article, and uses support vector machines (SVM) to identify biological entity names. It is also computationally efficient and robust to noise commonly found in full text material. We use the method to create a protein name dictionary from a set of 80,528 full text articles. Only 8.3% of the names in this dictionary match SwissProt description lines. We assess the quality of the dictionary by studying its protein name recognition performance in full text. CONCLUSION: This dictionary term lookup method compares favourably to other published methods, supporting the significance of our direct extraction approach. The method is strong in recognizing name variants not found in SwissProt.

Abstracting and Indexing↗

Machine learning in soil classification.

In a number of engineering problems, e.g. in geotechnics, petroleum engineering, etc. intervals of measured series data (signals) are to be attributed a class maintaining the constraint of contiguity and standard classification methods could be inadequate. Classification in this case needs involvement of an expert who observes the magnitude and trends of the signals in addition to any a priori information that might be available. In this paper, an approach for automating this classification procedure is presented. Firstly, a segmentation algorithm is developed and applied to segment the measured signals. Secondly, the salient features of these segments are extracted using boundary energy method. Based on the measured data and extracted features to assign classes to the segments classifiers are built; they employ Decision Trees, ANN and Support Vector Machines. The methodology was tested in classifying sub-surface soil using measured data from Cone Penetration Testing and satisfactory results were obtained.

Algorithms↗

Understanding protein dispensability through machine-learning analysis of high-throughput data.

MOTIVATION: Protein dispensability is fundamental to the understanding of gene function and evolution. Recent advances in generating high-throughput data such as genomic sequence data, protein-protein interaction data, gene-expression data and growth-rate data of mutants allow us to investigate protein dispensability systematically at the genome scale. RESULTS: In our studies, protein dispensability is represented as a fitness score that is measured by the growth rate of gene-deletion mutants. By the analyses of high-throughput data in yeast Saccharomyces cerevisiae, we found that a protein's dispensability had significant correlations with its evolutionary rate and duplication rate, as well as its connectivity in protein-protein interaction network and gene-expression correlation network. Neural network and support vector machine were applied to predict protein dispensability through high-throughput data. Our studies shed some lights on global characteristics of protein dispensability and evolution. AVAILABILITY: The original datasets for protein dispensability analysis and prediction, together with related scripts, are available at http://digbio.missouri.edu/~ychen/ProDispen/ CONTACT: xudong@missouri.edu.

Artificial Intelligence↗

Machine learning approaches for phenotype-genotype mapping: predicting heterozygous mutations in the CYP21B gene from steroid profiles.

OBJECTIVE: Non-linear relations between multiple biochemical parameters are the basis for the diagnosis of many diseases. Traditional linear analytical methods are not reliable predictors. Novel nonlinear techniques are increasingly used to improve the diagnostic accuracy of automated data interpretation. This has been exemplified in particular for the classification and diagnostic prediction of cancers based on expression profiling data. Our objective was to predict the genotype from complex biochemical data by comparing the performance of experienced clinicians to traditional linear analysis, and to novel non-linear analytical methods. DESIGN AND METHODS: As a model, we used a well-defined set of interconnected data consisting of unstimulated serum levels of steroid intermediates assessed in 54 subjects heterozygous for a mutation of the 21-hydroxylase gene (CYP21B) and in 43 healthy controls. RESULTS: The genetic alteration was predicted from the pattern of steroid levels with an accuracy of 39% by clinicians and of 64% by linear analysis. In contrast, non-linear analysis, such as self-organizing artificial neural networks, support vector machines, and nearest neighbour classifiers, allowed for higher accuracy up to 83%. CONCLUSIONS: The successful application of these non-linear adaptive methods to capture specific biochemical problems may have generalized implications for biochemical testing in many areas. Nonlinear analytical techniques such as neural networks, support vector machines, and nearest neighbour classifiers may serve as an important adjunct to the decision process of a human investigator not 'trained' in a specific complex clinical or laboratory setting and may aid them to classify the problem more directly.

Adult↗

Recognizing names in biomedical texts: a machine learning approach.

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective and efficient literature mining and knowledge discovery that can help biologists to gather and make use of the knowledge encoded in text documents. In order to make organized and structured information available, automatically recognizing biomedical entity names becomes critical and is important for information retrieval, information extraction and automated knowledge acquisition. RESULTS: In this paper, we present a named entity recognition system in the biomedical domain, called PowerBioNE. In order to deal with the special phenomena of naming conventions in the biomedical domain, we propose various evidential features: (1) word formation pattern; (2) morphological pattern, such as prefix and suffix; (3) part-of-speech; (4) head noun trigger; (5) special verb trigger and (6) name alias feature. All the features are integrated effectively and efficiently through a hidden Markov model (HMM) and a HMM-based named entity recognizer. In addition, a k-Nearest Neighbor (k-NN) algorithm is proposed to resolve the data sparseness problem in our system. Finally, we present a pattern-based post-processing to automatically extract rules from the training data to deal with the cascaded entity name phenomenon. From our best knowledge, PowerBioNE is the first system which deals with the cascaded entity name phenomenon. Evaluation shows that our system achieves the F-measure of 66.6 and 62.2 on the 23 classes of GENIA V3.0 and V1.1, respectively. In particular, our system achieves the F-measure of 75.8 on the "protein" class of GENIA V3.0. For comparison, our system outperforms the best published result by 7.8 on GENIA V1.1, without help of any dictionaries. It also shows that our HMM and the k-NN algorithm outperform other models, such as back-off HMM, linear interpolated HMM, support vector machines, C4.5, C4.5 rules and RIPPER, by effectively capturing the local context dependency and resolving the data sparseness problem. Moreover, evaluation on GENIA V3.0 shows that the post-processing for the cascaded entity name phenomenon improves the F-measure by 3.9. Finally, error analysis shows that about half of the errors are caused by the strict annotation scheme and the annotation inconsistency in the GENIA corpus. This suggests that our system achieves an acceptable F-measure of 83.6 on the 23 classes of GENIA V3.0 and in particular 86.2 on the "protein" class, without help of any dictionaries. We think that a F-measure of 90 on the 23 classes of GENIA V3.0 and in particular 92 on the "protein" class, can be achieved through refining of the annotation scheme in the GENIA corpus, such as flexible annotation scheme and annotation consistency, and inclusion of a reasonable biomedical dictionary. AVAILABILITY: A demo system is available at http://textmining.i2r.a-star.edu.sg/NLS/demo.htm. Technology license is available upon the bilateral agreement.

Abstracting and Indexing↗

Monitoring of complex industrial bioprocesses for metabolite concentrations using modern spectroscopies and machine learning: application to gibberellic acid production.

Two rapid vibrational spectroscopic approaches (diffuse reflectance-absorbance Fourier transform infrared [FT-IR] and dispersive Raman spectroscopy), and one mass spectrometric method based on in vacuo Curie-point pyrolysis (PyMS), were investigated in this study. A diverse range of unprocessed, industrial fed-batch fermentation broths containing the fungus Gibberella fujikuroi producing the natural product gibberellic acid, were analyzed directly without a priori chromatographic separation. Partial least squares regression (PLSR) and artificial neural networks (ANNs) were applied to all of the information-rich spectra obtained by each of the methods to obtain quantitative information on the gibberellic acid titer. These estimates were of good precision, and the typical root-mean-square error for predictions of concentrations in an independent test set was <10% over a very wide titer range from 0 to 4925 ppm. However, although PLSR and ANNs are very powerful techniques they are often described as "black box" methods because the information they use to construct the calibration model is largely inaccessible. Therefore, a variety of novel evolutionary computation-based methods, including genetic algorithms and genetic programming, were used to produce models that allowed the determination of those input variables that contributed most to the models formed, and to observe that these models were predominantly based on the concentration of gibberellic acid itself. This is the first time that these three modern analytical spectroscopies, in combination with advanced chemometric data analysis, have been compared for their ability to analyze a real commercial bioprocess. The results demonstrate unequivocally that all methods provide very rapid and accurate estimates of the progress of industrial fermentations, and indicate that, of the three methods studied, Raman spectroscopy is the ideal bioprocess monitoring method because it can be adapted for on-line analysis.

Algorithms↗

Machine learning approaches to lung cancer prediction from mass spectra.

We addressed the problem of discriminating between 24 diseased and 17 healthy specimens on the basis of protein mass spectra. To prepare the data, we performed mass to charge ratio (m/z) normalization, baseline elimination, and conversion of absolute peak height measures to height ratios. After preprocessing, the major difficulty encountered was the extremely large number of variables (1676 m/z values) versus the number of examples (41). Dimensionality reduction was treated as an integral part of the classification process; variable selection was coupled with model construction in a single ten-fold cross-validation loop. We explored different experimental setups involving two peak height representations, two variable selection methods, and six induction algorithms, all on both the original 1676-mass data set and on a prescreened 124-mass data set. Highest predictive accuracies (1-2 off-sample misclassifications) were achieved by a multilayer perceptron and Naïve Bayes, with the latter displaying more consistent performance (hence greater reliability) over varying experimental conditions. We attempted to identify the most discriminant peaks (proteins) on the basis of scores assigned by the two variable selection methods and by neural network based sensitivity analysis. These three scoring schemes consistently ranked four peaks as the most relevant discriminators: 11683, 1403, 17350 and 66107.

Algorithms↗

A shape-based machine learning tool for drug design.

Building predictive models for iterative drug design in the absence of a known target protein structure is an important challenge. We present a novel technique, Compass, that removes a major obstacle to accurate prediction by automatically selecting conformations and alignments of molecules without the benefit of a characterized active site. The technique combines explicit representation of molecular shape with neural network learning methods to produce highly predictive models, even across chemically distinct classes of molecules. We apply the method to predicting human perception of musk odor and show how the resulting models can provide graphical guidance for chemical modifications.

Algorithms↗

Learning machines applied to potential forest distribution.

The clearing of forests to obtain land for pasture and agriculture and the replacement of autochthonous species by other faster-growing varieties of trees for timber have both led to the loss of vast areas of forest worldwide. At present, many developed countries are attempting to reverse these effects, establishing policies for the restoration of older woodland systems. Reforestation is a complex matter, planned and carried out by experts who need objective information regarding the type of forest that can be sustained in each area. This information is obtained by drawing up feasibility models constructed using statistical methods that make use of the information provided by morphological and environmental variables (height, gradient, rainfall, etc.) that partially condition the presence or absence of a specific kind of forestation in an area. The aim of this work is to construct a set of feasibility models for woodland located in the basin of the River Liébana (NW Spain), to serve as a support tool for the experts entrusted with carrying out the reforestation project. The techniques used are multilayer perceptron neural networks and support vector machines. Their results will be compared to the results obtained by traditional techniques (such as discriminant analysis and logistic regression) by measuring the degree of fit between each model and the existing distribution of woodlands. The interpretation and problems of the feasibility models are commented on in the Discussion section.

Artificial Intelligence↗