PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Computerized breast cancer diagnosis and prognosis from fine-needle aspirates.

OBJECTIVE: To use digital image analysis and machine learning to (1) improve breast mass diagnosis based on fine-needle aspirates and (2) improve breast cancer prognostic estimations. DESIGN: An interactive computer system evaluates, diagnoses, and determines prognosis based on cytologic features derived from a digital scan of fine-needle aspirate slides. SETTING: The University of Wisconsin (Madison) Departments of Computer Science and Surgery and the University of Wisconsin Hospital and Clinics. PATIENTS: Five hundred sixty-nine consecutive patients (212 with cancer and 357 with benign masses) provided the data for the diagnostic algorithm, and an additional 118 (31 with malignant masses and 87 with benign masses) consecutive, new patients tested the algorithm. One hundred ninety of these patients with invasive cancer and without distant metastases were used for prognosis. INTERVENTIONS: Surgical biopsy specimens were taken from all cancers and some benign masses. The remaining cytologically benign masses were followed up for a year and surgical biopsy specimens were taken if they changed in size or character. Patients with cancer received standard treatment. OUTCOME MEASURES: Cross validation was used to project the accuracy of the diagnostic algorithm and to determine the importance of prognostic features. In addition, the mean errors were calculated between the actual times of distant disease occurrence and the times predicted using various prognostic features. Statistical analyses were also done. RESULTS: The predicted diagnostic accuracy was 97% and the actual diagnostic accuracy on 118 new samples was 100%. Tumor size and lymph node status were weak prognosticators compared with nuclear features, in particular those measuring nuclear size. Compared with the actual time for recurrence, the mean error of predicted times for recurrence with the nuclear features was 17.9 months and was 20.1 months with tumor size and lymph node status (P = .11). CONCLUSION: Computer technology will improve breast fine-needle aspirate accuracy and prognostic estimations.

Biopsy, Needle↗

Recent progresses in the application of machine learning approach for predicting protein functional class independent of sequence similarity.

Protein sequence contains clues to its function. Functional prediction from sequence presents a challenge particularly for proteins that have low or no sequence similarity to proteins of known function. Recently, machine learning methods have been explored for predicting functional class of proteins from sequence-derived properties independent of sequence similarity, which showed promising potential for low- and non-homologous proteins. These methods can thus be explored as potential tools to complement alignment- and clustering-based methods for predicting protein function. This article reviews the strategies, current progresses, and underlying difficulties in using machine learning methods for predicting the functional class of proteins. The relevant software and web-servers are described. The reported prediction performances in the application of these methods are also presented, which need to be interpreted with caution as they are dependent on such factors as datasets used and choice of parameters.

Algorithms↗

Prediction of CTL epitopes using QM, SVM and ANN techniques.

Cytotoxic T lymphocyte (CTL) epitopes are potential candidates for subunit vaccine design for various diseases. Most of the existing T cell epitope prediction methods are indirect methods that predict MHC class I binders instead of CTL epitopes. In this study, a systematic attempt has been made to develop a direct method for predicting CTL epitopes from an antigenic sequence. This method is based on quantitative matrix (QM) and machine learning techniques such as Support Vector Machine (SVM) and Artificial Neural Network (ANN). This method has been trained and tested on non-redundant dataset of T cell epitopes and non-epitopes that includes 1137 experimentally proven MHC class I restricted T cell epitopes. The accuracy of QM-, ANN- and SVM-based methods was 70.0, 72.2 and 75.2%, respectively. The performance of these methods has been evaluated through Leave One Out Cross-Validation (LOOCV) at a cutoff score where sensitivity and specificity was nearly equal. Finally, both machine-learning methods were used for consensus and combined prediction of CTL epitopes. The performances of these methods were evaluated on blind dataset where machine learning-based methods perform better than QM-based method. We also demonstrated through subgroup analysis that our methods can discriminate between T-cell epitopes and MHC binders (non-epitopes). In brief this method allows prediction of CTL epitopes using QM, SVM, ANN approaches. The method also facilitates prediction of MHC restriction in predicted T cell epitopes.

Algorithms↗

Transcriptomics-based exploration of ubiquitination-related biomarkers and potential molecular mechanisms in laryngeal squamous cell carcinoma.

BACKGROUND: One of the most common and prevalent cancers is laryngeal squamous cell carcinoma (LSCC), which poses a great threat to the life and health of the patient. Nonetheless, it has been demonstrated that ubiquitination is crucial for the development and course of LSCC. Therefore, it is particularly important to identify biomarkers for ubiquitination-related genes (UbRGs) in LSCC. METHODS: Differentially expressed genes (DEGs) in the LSCC versus controls were obtained by differential expression analysis. Also, key modular genes associated with LSCC were obtained using weighted gene co-expression network analysis (WGCNA). Next, DEGs, key module genes, and UbRGs were taken to intersect to obtain candidate genes. And then machine algorithms were to screen potential biomarkers, further their diagnostic value were analyzed and validated. Then, therapeutic agents for biomarkers were predict. In addition, the regulatory networks of the biomarkers were mapped. The expression levels of biomarkers were detected in clinical samples using reverse transcription-quantitative PCR (RT-qPCR). RESULTS: A total of eight candidate genes were acquired by the overlap 1,911 DEGs, the key modular genes of WGCNA, and 1,393 UbRGs. A sum of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) were identified by two machine learning, then these four biomarkers were validated in GSE127165 and the expression trend was consistent with TCGA-LSCC, they were recorded as biomarkers. Moreover, the accuracy of the biomarkers in predicting clinical aspects of LSCC was confirmed by the receiver operating characteristic (ROC) curves. Subsequently, cancers such as malignant neoplasms, colorectal cancers, tumors, and primary malignant neoplasms were significantly associated with the biomarkers, which further suggests that these four biomarkers were strongly associated with cancer. Meanwhile, the drugs garcinol, cocaine, and triazolam, among others, used for LSCC treatment were predicted. Finally, transcription factors (TFs) (BRD4, MYC, AR, and CTCF) were predicted to regulate the biomarkers. RT-qPCR assays illustrated that the expression trends of KAT2B, LNX1 and NBEAL2 remained consistent with the dataset. CONCLUSION: The identification of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) associated with UbRGs could ultimately serve as a predictive clinical diagnosis of LSCC and provide insight into the molecular mechanisms of LSCC.

Humans↗

Tumour class prediction and discovery by microarray-based DNA methylation analysis.

Aberrant DNA methylation of CpG sites is among the earliest and most frequent alterations in cancer. Several studies suggest that aberrant methylation occurs in a tumour type-specific manner. However, large-scale analysis of candidate genes has so far been hampered by the lack of high throughput assays for methylation detection. We have developed the first microarray-based technique which allows genome-wide assessment of selected CpG dinucleotides as well as quantification of methylation at each site. Several hundred CpG sites were screened in 76 samples from four different human tumour types and corresponding healthy controls. Discriminative CpG dinucleotides were identified for different tissue type distinctions and used to predict the tumour class of as yet unknown samples with high accuracy using machine learning techniques. Some CpG dinucleotides correlate with progression to malignancy, whereas others are methylated in a tissue-specific manner independent of malignancy. Our results demonstrate that genome-wide analysis of methylation patterns combined with supervised and unsupervised machine learning techniques constitute a powerful novel tool to classify human cancers.

Algorithms↗

Data processing and classification analysis of proteomic changes: a case study of oil pollution in the mussel, Mytilus edulis.

BACKGROUND: Proteomics may help to detect subtle pollution-related changes, such as responses to mixture pollution at low concentrations, where clear signs of toxicity are absent. The challenges associated with the analysis of large-scale multivariate proteomic datasets have been widely discussed in medical research and biomarker discovery. This concept has been introduced to ecotoxicology only recently, so data processing and classification analysis need to be refined before they can be readily applied in biomarker discovery and monitoring studies. RESULTS: Data sets obtained from a case study of oil pollution in the Blue mussel were investigated for differential protein expression by retentate chromatography-mass spectrometry and decision tree classification. Different tissues and different settings were used to evaluate classifiers towards their discriminatory power. It was found that, due the intrinsic variability of the data sets, reliable classification of unknown samples could only be achieved on a broad statistical basis (n > 60) with the observed expression changes comprising high statistical significance and sufficient amplitude. The application of stringent criteria to guard against overfitting of the models eventually allowed satisfactory classification for only one of the investigated data sets and settings. CONCLUSION: Machine learning techniques provide a promising approach to process and extract informative expression signatures from high-dimensional mass-spectrometry data. Even though characterisation of the proteins forming the expression signatures would be ideal, knowledge of the specific proteins is not mandatory for effective class discrimination. This may constitute a new biomarker approach in ecotoxicology, where working with organisms, which do not have sequenced genomes render protein identification by database searching problematic. However, data processing has to be critically evaluated and statistical constraints have to be considered before supervised classification algorithms are employed.

Journal Article↗

BCI Competition 2003--Data set IIb: support vector machines for the P300 speller paradigm.

We propose an approach to analyze data from the P300 speller paradigm using the machine-learning technique support vector machines. In a conservative classification scheme, we found the correct solution after five repetitions. While the classification within the competition is designed for offline analysis, our approach is also well-suited for a real-world online solution: It is fast, requires only 10 electrode positions and demands only a small amount of preprocessing.

Algorithms↗

Temporal sequence learning, prediction, and control: a review of different models and their relation to biological mechanisms.

In this review, we compare methods for temporal sequence learning (TSL) across the disciplines machine-control, classical conditioning, neuronal models for TSL as well as spike-timing-dependent plasticity (STDP). This review introduces the most influential models and focuses on two questions: To what degree are reward-based (e.g., TD learning) and correlation-based (Hebbian) learning related? and How do the different models correspond to possibly underlying biological mechanisms of synaptic plasticity? We first compare the different models in an open-loop condition, where behavioral feedback does not alter the learning. Here we observe that reward-based and correlation-based learning are indeed very similar. Machine control is then used to introduce the problem of closed-loop control (e.g., actor-critic architectures). Here the problem of evaluative (rewards) versus nonevaluative (correlations) feedback from the environment will be discussed, showing that both learning approaches are fundamentally different in the closed-loop condition. In trying to answer the second question, we compare neuronal versions of the different learning architectures to the anatomy of the involved brain structures (basal-ganglia, thalamus, and cortex) and the molecular biophysics of glutamatergic and dopaminergic synapses. Finally, we discuss the different algorithms used to model STDP and compare them to reward-based learning rules. Certain similarities are found in spite of the strongly different timescales. Here we focus on the biophysics of the different calcium-release mechanisms known to be involved in STDP.

Forecasting↗

Biomarkers that discriminate multiple myeloma patients with or without skeletal involvement detected using SELDI-TOF mass spectrometry and statistical and machine learning tools.

Multiple Myeloma (MM) is a severely debilitating neoplastic disease of B cell origin, with the primary source of morbidity and mortality associated with unrestrained bone destruction. Surface enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF MS) was used to screen for potential biomarkers indicative of skeletal involvement in patients with MM. Serum samples from 48 MM patients, 24 with more than three bone lesions and 24 with no evidence of bone lesions were fractionated and analyzed in duplicate using copper ion loaded immobilized metal affinity SELDI chip arrays. The spectra obtained were compiled, normalized, and mass peaks with mass-to-charge ratios (m/z) between 2000 and 20,000 Da identified. Peak information from all fractions was combined together and analyzed using univariate statistics, as well as a linear, partial least squares discriminant analysis (PLS-DA), and a non-linear, random forest (RF), classification algorithm. The PLS-DA model resulted in prediction accuracy between 96-100%, while the RF model was able to achieve a specificity and sensitivity of 87.5% each. Both models as well as multiple comparison adjusted univariate analysis identified a set of four peaks that were the most discriminating between the two groups of patients and hold promise as potential biomarkers for future diagnostic and/or therapeutic purposes.

Adult↗

SVM-BALSA: remote homology detection based on Bayesian sequence alignment.

Biopolymer sequence comparison to identify evolutionarily related proteins, or homologs, is one of the most common tasks in bioinformatics. Support vector machines (SVMs) represent a new approach to the problem in which statistical learning theory is employed to classify proteins into families, thus identifying homologous relationships. Current SVM approaches have been shown to outperform iterative profile methods, such as PSI-BLAST, for protein homology classification. In this study, we demonstrate that the utilization of a Bayesian alignment score, which accounts for the uncertainty of all possible alignments, in the SVM construction improves sensitivity compared to the traditional dynamic programming implementation over a benchmark dataset consisting of 54 unique protein families. The SVM-BALSA algorithms returns a higher area under the receiver operating characteristic (ROC) curves for 37 of the 54 families and achieves an improved overall performance curve at a significance level of 0.07.

Bayes Theorem↗

Prediction of antimicrobial minimum inhibitory concentration from bacterial genomes using a scalable and interpretable machine learning approach.

Although machine learning models can predict antimicrobial susceptibility from bacterial whole genome sequencing (WGS), state-of-the-art approaches are computationally demanding or dependent on knowledge of genetic resistance determinants. Here, we describe an efficient data-driven approach to predicting minimum inhibitory concentration (MIC) by progressively extending and refining predictive genome segments, independent of prior knowledge of resistance determinants. Resultant models had high interpretability - known and potentially novel resistance determinants were captured. Using 762 clinical E. coli strains, 71.6% of predictions were within one dilution of the measured MIC. Models trained with this algorithm generalised better onto external data (F1 score = 0.85) compared with alternative models trained on annotated resistance determinants (F1 = 0.82) or k-mer counts (F1 = 0.74). Computational demands were low (RAM usage 23.6GB vs 38.8GB for k-mer model). These advantages represent an important advance in predicting antimicrobial susceptibility from WGS, with potential applications for clinical diagnostics, drug development, and surveillance.

Journal Article↗

Local irritation/corrosion testing strategies: extending a decision support system by applying self-learning classifiers.

Procedures have been established and tested for the extension of a decision support system (DSS) for the prediction of the local irritation/corrosion potential of chemicals by using self-learning classifiers. The different approaches (decision trees, distances examinations in a multidimensional space, k-nearest-neighbour method) have been implemented, tested and evaluated independently. A combination of all of the established extension approaches was also developed and tested. Self-learning classifiers are constructed "automatically" by a computer, i.e. they are not derived by a human expert, and thus they can be constructed with minimal effort. The classifiers presented here extend the existing DSS in a manner that increased significantly the predictive power of the extended system. However, automatically calculated results of self-learning classifiers are produced by a machine, and a machine is incapable of explaining the toxicological relevance of the results obtained. Thus, these results must be accepted, despite an inability to prove their reliability. Only the mathematical correctness of the method and the prediction rates for suitable test cases can lend some credibility to predictions produced by a computer calculating on a self-learning basis. This may not be adequate for regulatory hazard assessment purposes.

Algorithms↗

Development of neural mechanisms for machine learning.

The goal of this work is to develop a humanoid robot's perceptual mechanisms through the use of learning aids. We describe methods to enable learning on a humanoid robot using learning aids such as books, drawing materials, boards, educational videos or other children toys. Visual properties of objects are learned and inserted into a recognition scheme, which is then applied to acquire new object representations - we propose learning through developmental stages. Inspired in infant development, we will also boost the robot's perceptual capabilities by having a human caregiver performing educational and play activities with the robot (such as drawing, painting or playing with a toy train on a railway). We describe original algorithms to extract meaningful percepts from such learning experiments. Experimental evaluation of the algorithms corroborates the theoretical framework.

Algorithms↗

Machine-learning techniques for macromolecular crystallization data.

Systematizing belief systems regarding macromolecular crystallization has two major advantages: automation and clarification. In this paper, methodologies are presented for systematizing and representing knowledge about the chemical and physical properties of additives used in crystallization experiments. A novel autonomous discovery program is introduced as a method to prune rule-based models produced from crystallization data augmented with such knowledge. Computational experiments indicate that such a system can retain and present informative rules pertaining to protein crystallization that warrant further confirmation via experimental techniques.

Algorithms↗

Transductive machine learning for reliable medical diagnostics.

In the past decades Machine Learning tools have been successfully used in several medical diagnostic problems. While they often significantly outperform expert physicians (in terms of diagnostic accuracy, sensitivity, and specificity), they are mostly not being used in practice. One reason for this is that it is difficult to obtain an unbiased estimation of diagnose's reliability. We discuss how reliability of diagnoses is assessed in medical decision making and propose a general framework for reliability estimation in Machine Learning, based on transductive inference. We compare our approach with a usual (Machine Learning) probabilistic approach as well as with classical stepwise diagnostic process where reliability of diagnose is presented as its posttest probability. The proposed transductive approach is evaluated on several medical data sets from the UCI (University of California, Irvine) repository as well as on a practical problem of clinical diagnosis of the coronary artery disease. In all cases significant improvements over existing techniques are achieved.

Adult↗

Extracting synonymous gene and protein terms from biological literature.

MOTIVATION: Genes and proteins are often associated with multiple names. More names are added as new functional or structural information is discovered. Because authors can use any one of the known names for a gene or protein, information retrieval and extraction would benefit from identifying the gene and protein terms that are synonyms of the same substance. RESULTS: We have explored four complementary approaches for extracting gene and protein synonyms from text, namely the unsupervised, partially supervised, and supervised machine-learning techniques, as well as the manual knowledge-based approach. We report results of a large scale evaluation of these alternatives over an archive of biological journal articles. Our evaluation shows that our extraction techniques could be a valuable supplement to resources such as SWISSPROT, as our systems were able to capture gene and protein synonyms not listed in the SWISSPROT database.

Abstracting and Indexing↗

Identifying genes related to chemosensitivity using support vector machine.

In an effort to identify genes involved in chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine (SVM) is used to associate genes with any of five predefined anticancer drug mechanistic categories. The drug activity profiles are used as training examples to train the SVM and then the gene expression profiles are used as test examples to predict their associated mechanistic categories. This method of correlating drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Algorithms↗

Closed-loop, multiobjective optimization of analytical instrumentation: gas chromatography/time-of-flight mass spectrometry of the metabolomes of human serum and of yeast fermentations.

The number of instrumental parameters controlling modern analytical apparatus can be substantial, and varying them systematically to optimize a particular chromatographic separation, for example, is out of the question because of the astronomical number of combinations that are possible (i.e., the "search space" is very large). However, heuristic methods, such as those based on evolutionary computing, can be used to explore such search spaces efficiently. We here describe the implementation of an entirely automated (closed-loop) strategy for doing this and apply it to the optimization of gas chromatographic separations of the metabolomes of human serum and of yeast fermentation broths. Without human intervention, the Robot Chromatographer system (i) initializes the settings on the instrument, (ii) controls the analytical run, (iii) extracts the variables defining the analytical performance (specifically the number of peaks, signal/noise ratio, and run time), (iv) chooses (via the PESA-II multiobjective genetic algorithm), and (v) programs the next series of instrumental settings, the whole continuing in an iterative cycle until suitable sets of optimal conditions have been established. Genetic programming was used to remove noise peaks and to establish the basis for the improvements observed. The system showed that the number of peaks observable depended enormously on the conditions used and served to increase them by as much as 3-fold (e.g., to over 950 in human serum) while in many cases maintaining or reducing the run time and preserving excellent signal/noise ratios. The evolutionary closed-loop machine learning strategy we describe is generic to any type of analytical optimization.

Fermentation↗