PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Challenges in real-life emotion annotation and machine learning based detection.

Since the early studies of human behavior, emotion has attracted the interest of researchers in many disciplines of Neurosciences and Psychology. More recently, it is a growing field of research in computer science and machine learning. We are exploring how the expression of emotion is perceived by listeners and how to represent and automatically detect a subject's emotional state in speech. In contrast with most previous studies, conducted on artificial data with archetypal emotions, this paper addresses some of the challenges faced when studying real-life non-basic emotions. We present a new annotation scheme allowing the annotation of emotion mixtures. Our studies of real-life spoken dialogs from two call center services reveal the presence of many blended emotions, dependent on the dialog context. Several classification methods (SVM, decision trees) are compared to identify relevant emotional states from prosodic, disfluency and lexical cues extracted from the real-life spoken human-human interactions.

Artificial Intelligence↗

The forecast of the postoperative survival time of patients suffered from non-small cell lung cancer based on PCA and extreme learning machine.

In this paper, a new effective model is proposed to forecast how long the postoperative patients suffered from non-small cell lung cancer will survive. The new effective model which is based on the extreme learning machine (ELM) and principal component analysis (PCA) can forecast successfully the postoperative patients' survival time. The new model obtains better prediction accuracy and faster convergence rate which the model using backpropagation (BP) algorithm and the Levenberg-Marquardt (LM) algorithm to forecast the postoperative patients' survival time can not achieve. Finally, simulation results are given to verify the efficiency and effectiveness of our proposed new model.

Algorithms↗

Gene selection from microarray data for cancer classification--a machine learning approach.

A DNA microarray can track the expression levels of thousands of genes simultaneously. Previous research has demonstrated that this technology can be useful in the classification of cancers. Cancer microarray data normally contains a small number of samples which have a large number of gene expression levels as features. To select relevant genes involved in different types of cancer remains a challenge. In order to extract useful gene information from cancer microarray data and reduce dimensionality, feature selection algorithms were systematically investigated in this study. Using a correlation-based feature selector combined with machine learning algorithms such as decision trees, naïve Bayes and support vector machines, we show that classification performance at least as good as published results can be obtained on acute leukemia and diffuse large B-cell lymphoma microarray data sets. We also demonstrate that a combined use of different classification and feature selection approaches makes it possible to select relevant genes with high confidence. This is also the first paper which discusses both computational and biological evidence for the involvement of zyxin in leukaemogenesis.

Algorithms↗

Image analysis and machine learning applied to breast cancer diagnosis and prognosis.

Fine needle aspiration (FNA) accuracy is limited by, among other factors, the subjective interpretation of the aspirate. We have increased breast FNA accuracy by coupling digital image analysis methods with machine learning techniques. Additionally, our mathematical approach captures nuclear features ("grade") that are prognostically more accurate than are estimates based on tumor size and lymph node status. An interactive computer system evaluates, diagnoses and determines prognosis based on nuclear features derived directly from a digital scan of FNA slides. A consecutive series of 569 patients provided the data for the diagnostic study. A 166-patient subset provided the data for the prognostic study. An additional 75 consecutive, new patients provided samples to test the diagnostic system. The projected prospective accuracy of the diagnostic system was estimated to be 97% by 10-fold cross-validation, and the actual accuracy on 75 new samples was 100%. The projected prospective accuracy of the prognostic system was estimated to be 86% by leave-one-out testing.

Biopsy, Needle↗

Optimization of rifamycin B fermentation in shake flasks via a machine-learning-based approach.

Rifamycin B is an important polyketide antibiotic used in the treatment of tuberculosis and leprosy. We present results on medium optimization for Rifamycin B production via a barbital insensitive mutant strain of Amycolatopsis mediterranei S699. Machine-learning approaches such as Genetic algorithm (GA), Neighborhood analysis (NA) and Decision Tree technique (DT) were explored for optimizing the medium composition. Genetic algorithm was applied as a global search algorithm while NA was used for a guided local search and to develop medium predictors. The fermentation medium for Rifamycin B consisted of nine components. A large number of distinct medium compositions are possible by variation of concentration of each component. This presents a large combinatorial search space. Optimization was achieved within five generations via GA as well as NA. These five generations consisted of 178 shake-flask experiments, which is a small fraction of the search space. We detected multiple optima in the form of 11 distinct medium combinations. These medium combinations provided over 600% improvement in Rifamycin B productivity. Genetic algorithm performed better in optimizing fermentation medium as compared to NA. The Decision Tree technique revealed the media-media interactions qualitatively in the form of sets of rules for medium composition that give high as well as low productivity.

Actinomycetales↗

Machine learning approaches for the prediction of signal peptides and other protein sorting signals.

Prediction of protein sorting signals from the sequence of amino acids has great importance in the field of proteomics today. Recently, the growth of protein databases, combined with machine learning approaches, such as neural networks and hidden Markov models, have made it possible to achieve a level of reliability where practical use in, for example automatic database annotation is feasible. In this review, we concentrate on the present status and future perspectives of SignalP, our neural network-based method for prediction of the most well-known sorting signal: the secretory signal peptide. We discuss the problems associated with the use of SignalP on genomic sequences, showing that signal peptide prediction will improve further if integrated with predictions of start codons and transmembrane helices. As a step towards this goal, a hidden Markov model version of SignalP has been developed, making it possible to discriminate between cleaved signal peptides and uncleaved signal anchors. Furthermore, we show how SignalP can be used to characterize putative signal peptides from an archaeon, Methanococcus jannaschii. Finally, we briefly review a few methods for predicting other protein sorting signals and discuss the future of protein sorting prediction in general.

Algorithms↗

A hybrid machine-learning approach for segmentation of protein localization data.

MOTIVATION: Subcellular protein localization data are critical to the quantitative understanding of cellular function and regulation. Such data are acquired via observation and quantitative analysis of fluorescently labeled proteins in living cells. Differentiation of labeled protein from cellular artifacts remains an obstacle to accurate quantification. We have developed a novel hybrid machine-learning-based method to differentiate signal from artifact in membrane protein localization data by deriving positional information via surface fitting and combining this with fluorescence-intensity-based data to generate input for a support vector machine. RESULTS: We have employed this classifier to analyze signaling protein localization in T-cell activation. Our classifier displayed increased performance over previously available techniques, exhibiting both flexibility and adaptability: training on heterogeneous data yielded a general classifier with good overall performance; training on more specific data yielded an extremely high-performance specific classifier. We also demonstrate accurate automated learning utilizing additional experimental data.

Animals↗

A machine learning evaluation of an artificial immune system.

ARTIS is an artificial immune system framework which contains several adaptive mechanisms. LISYS is a version of ARTIS specialized for the problem of network intrusion detection. The adaptive mechanisms of LISYS are characterized in terms of their machine-learning counterparts, and a series of experiments is described, each of which isolates a different mechanism of LISYS and studies its contribution to the system's overall performance. The experiments were conducted on a new data set, which is more recent and realistic than earlier data sets. The network intrusion detection problem is challenging because it requires one-class learning in an on-line setting with concept drift. The experiments confirm earlier experimental results with LISYS, and they study in detail how LISYS achieves success on the new data set.

Algorithms↗

Predicting hepatitis B virus-positive metastatic hepatocellular carcinomas using gene expression profiling and supervised machine learning.

Hepatocellular carcinoma (HCC) is one of the most common and aggressive human malignancies. Its high mortality rate is mainly a result of intra-hepatic metastases. We analyzed the expression profiles of HCC samples without or with intra-hepatic metastases. Using a supervised machine-learning algorithm, we generated for the first time a molecular signature that can classify metastatic HCC patients and identified genes that were relevant to metastasis and patient survival. We found that the gene expression signature of primary HCCs with accompanying metastasis was very similar to that of their corresponding metastases, implying that genes favoring metastasis progression were initiated in the primary tumors. Osteopontin, which was identified as a lead gene in the signature, was over-expressed in metastatic HCC; an osteopontin-specific antibody effectively blocked HCC cell invasion in vitro and inhibited pulmonary metastasis of HCC cells in nude mice. Thus, osteopontin acts as both a diagnostic marker and a potential therapeutic target for metastatic HCC.

Algorithms↗

Three machine learning techniques for automatic determination of rules to control locomotion.

Automatic prediction of gait events (e.g., heel contact, flat foot, initiation of the swing, etc.) and corresponding profiles of the activations of muscles is important for real-time control of locomotion. This paper presents three supervised machine learning (ML) techniques for prediction of the activation patterns of muscles and sensory data, based on the history of sensory data, for walking assisted by a functional electrical stimulation (FES). Those ML's are: 1) a multilayer perceptron with Levenberg-Marquardt modification of backpropagation learning algorithm; 2) an adaptive-network-based fuzzy inference system (ANFIS); and 3) a combination of an entropy minimization type of inductive learning (IL) technique and a radial basis function (RBF) type of artificial neural network with orthogonal least squares learning algorithm. Here we show the prediction of the activation of the knee flexor muscles and the knee joint angle for seven consecutive strides based on the history of the knee joint angle and the ground reaction forces. The data used for training and testing of ML's was obtained from a simulation of walking assisted with an FES system [39]. The ability of generating rules for an FES controller was selected as the most important criterion when comparing the ML's. Other criteria such as generalization of results, computational complexity, and learning rate were also considered. The minimal number of rules and the most explicit and comprehensible rules were obtained by ANFIS. The best generalization was obtained by the IL and RBF network.

Algorithms↗

Simple models for estimating dementia severity using machine learning.

Estimating dementia severity using the Clinical Dementia Rating (CDR) Scale is a two-stage process that currently is costly and impractical in community settings, and at best has an interrater reliability of 80%. Because staging of dementia severity is economically and clinically important, we used Machine Learning (ML) algorithms with an Electronic Medical Record (EMR) to identify simpler models for estimating total CDR scores. Compared to a gold standard, which required 34 attributes to derive total CDR scores, ML algorithms identified models with as few as seven attributes. The classification accuracy varied with the algorithm used with naïve Bayes giving the highest. (76%) The mildly demented severity class was the only one with significantly reduced accuracy (59%). If one groups the severity classes into normal, very mild-to-mildly demented, and moderate-to-severely demented, then classification accuracies are clinically acceptable (85%). These simple models can be used in community settings where it is currently not possible to estimate dementia severity due to time and cost constraints.

Algorithms↗

[Risk assessment in ovarian hyperstimulation syndrome (OHS) using the machine learning system (Decision Master) in 155 in-vitro fertilisations and embryo-transfer (IVF/ET) cycles with a long stimulation protocol].

In 155 selected IVF/ET cycles stimulated with the long protocol 25 cycles with severe OHS are included which turned up later on (purposely overrepresented). An inductive machine learning program is described both in informatics and medical essentials. It is tested whether there exists an algorithm for ruling out the above-mentioned complication in the follicular phase of the same cycle already. By cross validation 89% of the OHS could be predicted and proven by practical rules using hormone and ultrasound values to avoid similar events in ongoing or further cycles.

Adult↗

A program for machine learning of counting criteria: empirical induction of logic-based classification rules.

A program has been developed which derives classification rules from empirical observations and expresses these rules in a knowledge representation format called 'counting criteria'. Decision rules derived in this format are often more comprehensible than rules derived by existing machine learning programs such as AQ11. Use of the program is illustrated by the inference of discrimination criteria for certain types of bacteria based upon their biochemical characteristics. The program may be useful for the conceptual analysis of data and for the automatic generation of prototype knowledge bases for expert systems.

Artificial Intelligence↗

Prognostic significance of DNA damage response-related markers in esophageal squamous cell carcinoma using machine learning approaches.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) lacks reliable prognostic biomarkers. Homologous recombination deficiency (HRD) has been implicated in genomic instability across multiple cancers, but its prognostic significance in ESCC remains unexplored. This study aimed to evaluate HRD score as a prognostic biomarker and develop a machine learning-based predictive model for ESCC. METHODS: Transcriptomic and clinical data from 78 ESCC patients were obtained from The Cancer Genome Atlas (TCGA) and randomly split into training (70%) and test (30%) cohorts. Prognostic models were constructed using 112 machine learning algorithm combinations based on DNA damage response (DDR)-related genes. Gene set enrichment analysis (GSEA), somatic mutation profiling, and immune cell infiltration estimation via CIBERSORT were performed to characterize HRD-associated molecular features. RESULTS: High HRD scores were significantly associated with poorer overall survival (P<0.05). Among 112 algorithm combinations, the survival support vector machine (Survival-SVM) model demonstrated optimal performance [training concordance index (C-index): 0.741; test C-index: 0.708], identifying six hub genes: PARP1, MBD4, TELO2, NSMCE3, SMUG1, and BABAM1. A nomogram incorporating risk score (RS) and clinical variables achieved strong predictive accuracy for 1- to 3-year survival [area under the curve (AUC) >0.7]. High-HRD tumors exhibited distinct mutational patterns (TP53 and TTN) and enriched glutathione metabolism and cytochrome P450 pathways. Immune infiltration analysis revealed significant differences in plasma cell and neutrophil infiltration between risk groups (P<0.05), suggesting HRD-associated immune microenvironment remodeling. CONCLUSIONS: We developed a novel HRD-based prognostic model incorporating six DDR-related genes that demonstrates robust predictive performance in ESCC. HRD score is identified as an independent prognostic factor associated with genomic instability, immune microenvironment alterations, and clinical outcomes. These findings provide a theoretical basis for personalized treatment strategies, including potential applications of PARP inhibitors and immunotherapy in ESCC.

Esophageal squamous cell carcinoma (ESCC)↗

Discrimination of outer membrane proteins using machine learning algorithms.

Discriminating outer membrane proteins (OMPs) from other folding types of globular and membrane proteins is an important task both for identifying OMPs from genomic sequences and for the successful prediction of their secondary and tertiary structures. In this work, we have analyzed the performance of different methods, based on Bayes rules, logistic functions, neural networks, support vector machines, decision trees, etc. for discriminating OMPs. We found that most of the machine learning techniques discriminate OMPs with similar accuracy. The neural network-based method could discriminate the OMPs from other proteins [globular/transmembrane helical (TMH)] at the fivefold cross-validation accuracy of 91.0% in a dataset of 1,088 proteins. The accuracy of discriminating globular proteins is 88.8% and that of TMH proteins is 93.7%. Further, the neural network method is tested with globular proteins belonging to 30 different folding types and it could successfully exclude 95% of the considered proteins. The proteins with SAM domain such as knottins, rubredoxin, and thioredoxin folds are eliminated with 100% accuracy. These accuracy levels are comparable to or better than other methods in the literature. We suggest that this method could be effectively used to discriminate OMPs and for detecting OMPs in genomic sequences.

Algorithms↗

Automated annotation of keywords for proteins related to mycoplasmataceae using machine learning techniques.

MOTIVATION: With the increase in submission of sequences to public databases, the curators of these are not able to cope with the amount of information. The motivation of this work is to generate a system for automated annotation of data we are particularly interested in, namely proteins related to the Mycoplasmataceae family. Following previous works on automatic annotation using symbolic machine learning techniques, the present work proposes a method of automatic annotation of keywords (a part of the SWISS-PROT annotation procedure), and the validation, by an expert, of the annotation rules generated. The aim of this procedure is twofold: to complete the annotation of keywords of those proteins which is far from adequate, and to produce a prototype of the validation environment, which is aimed at an expert who does not have a deep knowledge of the structure of the current databases containing the necessary information s/he needs. RESULTS: As for the first objective, a rate of correct keywords annotation of 60% is reported in the literature. Our preliminary results show that with a slightly different method, applied this method to data related to Mycoplasmataceae only, we are able to increase that rate of correct annotation.

Abstracting and Indexing↗

Automatic resolution of ambiguous terms based on machine learning and conceptual relations in the UMLS.

UNLABELLED: Motivation. The UMLS has been used in natural language processing applications such as information retrieval and information extraction systems. The mapping of free-text to UMLS concepts is important for these applications. To improve the mapping, we need a method to disambiguate terms that possess multiple UMLS concepts. In the general English domain, machine-learning techniques have been applied to sense-tagged corpora, in which senses (or concepts) of ambiguous terms have been annotated (mostly manually). Sense disambiguation classifiers are then derived to determine senses (or concepts) of those ambiguous terms automatically. However, manual annotation of a corpus is an expensive task. We propose an automatic method that constructs sense-tagged corpora for ambiguous terms in the UMLS using MEDLINE abstracts. METHODS: For a term W that represents multiple UMLS concepts, a collection of MEDLINE abstracts that contain W is extracted. For each abstract in the collection, occurrences of concepts that have relations with W as defined in the UMLS are automatically identified. A sense-tagged corpus, in which senses of W are annotated, is then derived based on those identified concepts. The method was evaluated on a set of 35 frequently occurring ambiguous biomedical abbreviations using a gold standard set that was automatically derived. The quality of the derived sense-tagged corpus was measured using precision and recall. RESULTS: The derived sense-tagged corpus had an overall precision of 92.9% and an overall recall of 47.4%. After removing rare senses and ignoring abbreviations with closely related senses, the overall precision was 96.8% and the overall recall was 50.6%. CONCLUSIONS: UMLS conceptual relations and MEDLINE abstracts can be used to automatically acquire knowledge needed for resolving ambiguity when mapping free-text to UMLS concepts.

Abbreviations as Topic↗