PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Recognizing names in biomedical texts: a machine learning approach.

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective and efficient literature mining and knowledge discovery that can help biologists to gather and make use of the knowledge encoded in text documents. In order to make organized and structured information available, automatically recognizing biomedical entity names becomes critical and is important for information retrieval, information extraction and automated knowledge acquisition. RESULTS: In this paper, we present a named entity recognition system in the biomedical domain, called PowerBioNE. In order to deal with the special phenomena of naming conventions in the biomedical domain, we propose various evidential features: (1) word formation pattern; (2) morphological pattern, such as prefix and suffix; (3) part-of-speech; (4) head noun trigger; (5) special verb trigger and (6) name alias feature. All the features are integrated effectively and efficiently through a hidden Markov model (HMM) and a HMM-based named entity recognizer. In addition, a k-Nearest Neighbor (k-NN) algorithm is proposed to resolve the data sparseness problem in our system. Finally, we present a pattern-based post-processing to automatically extract rules from the training data to deal with the cascaded entity name phenomenon. From our best knowledge, PowerBioNE is the first system which deals with the cascaded entity name phenomenon. Evaluation shows that our system achieves the F-measure of 66.6 and 62.2 on the 23 classes of GENIA V3.0 and V1.1, respectively. In particular, our system achieves the F-measure of 75.8 on the "protein" class of GENIA V3.0. For comparison, our system outperforms the best published result by 7.8 on GENIA V1.1, without help of any dictionaries. It also shows that our HMM and the k-NN algorithm outperform other models, such as back-off HMM, linear interpolated HMM, support vector machines, C4.5, C4.5 rules and RIPPER, by effectively capturing the local context dependency and resolving the data sparseness problem. Moreover, evaluation on GENIA V3.0 shows that the post-processing for the cascaded entity name phenomenon improves the F-measure by 3.9. Finally, error analysis shows that about half of the errors are caused by the strict annotation scheme and the annotation inconsistency in the GENIA corpus. This suggests that our system achieves an acceptable F-measure of 83.6 on the 23 classes of GENIA V3.0 and in particular 86.2 on the "protein" class, without help of any dictionaries. We think that a F-measure of 90 on the 23 classes of GENIA V3.0 and in particular 92 on the "protein" class, can be achieved through refining of the annotation scheme in the GENIA corpus, such as flexible annotation scheme and annotation consistency, and inclusion of a reasonable biomedical dictionary. AVAILABILITY: A demo system is available at http://textmining.i2r.a-star.edu.sg/NLS/demo.htm. Technology license is available upon the bilateral agreement.

Abstracting and Indexing↗

Mapping of neural networks onto the memory-processor integrated architecture.

In this paper, an effective memory-processor integrated architecture, called memory-based processor array for artificial neural networks (MPAA), is proposed. The MPAA can be easily integrated into any host system via memory interface. Specifically, the MPA system provides an efficient mechanism for its local memory accesses allowed by row and column bases, using hybrid row and column decoding, which is suitable for computation models of ANNs such as the accessing and alignment patterns given for matrix-by-vector operations. Mapping algorithms to implement the multilayer perceptron with backpropagation learning on the MPAA system are also provided. The proposed algorithms support both neuron and layer level parallelisms which allow the MPAA system to operate the learning phase as well as the recall phase in the pipelined fashion. Performance evaluation is provided by detailed comparison in terms of two metrics such as the cost and number of computation steps. The results show that the performance of the proposed architecture and algorithms is superior to those of the previous approaches, such as one-dimensional single-instruction multiple data (SIMD) arrays, two-dimensional SIMD arrays, systolic ring structures, and hypercube machines.

Journal Article↗

STAR: predicting recombination sites from amino acid sequence.

BACKGROUND: Designing novel proteins with site-directed recombination has enormous prospects. By locating effective recombination sites for swapping sequence parts, the probability that hybrid sequences have the desired properties is increased dramatically. The prohibitive requirements for applying current tools led us to investigate machine learning to assist in finding useful recombination sites from amino acid sequence alone. RESULTS: We present STAR, Site Targeted Amino acid Recombination predictor, which produces a score indicating the structural disruption caused by recombination, for each position in an amino acid sequence. Example predictions contrasted with those of alternative tools, illustrate STAR'S utility to assist in determining useful recombination sites. Overall, the correlation coefficient between the output of the experimentally validated protein design algorithm SCHEMA and the prediction of STAR is very high (0.89). CONCLUSION: STAR allows the user to explore useful recombination sites in amino acid sequences with unknown structure and unknown evolutionary origin. The predictor service is available from http://pprowler.itee.uq.edu.au/star.

Algorithms↗

Comparative experiments on learning information extractors for proteins and their interactions.

OBJECTIVE: Automatically extracting information from biomedical text holds the promise of easily consolidating large amounts of biological knowledge in computer-accessible form. This strategy is particularly attractive for extracting data relevant to genes of the human genome from the 11 million abstracts in Medline. However, extraction efforts have been frustrated by the lack of conventions for describing human genes and proteins. We have developed and evaluated a variety of learned information extraction systems for identifying human protein names in Medline abstracts and subsequently extracting information on interactions between the proteins. METHODS AND MATERIAL: We used a variety of machine learning methods to automatically develop information extraction systems for extracting information on gene/protein name, function and interactions from Medline abstracts. We present cross-validated results on identifying human proteins and their interactions by training and testing on a set of approximately 1000 manually-annotated Medline abstracts that discuss human genes/proteins. RESULTS: We demonstrate that machine learning approaches using support vector machines and maximum entropy are able to identify human proteins with higher accuracy than several previous approaches. We also demonstrate that various rule induction methods are able to identify protein interactions with higher precision than manually-developed rules. CONCLUSION: Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus.

Algorithms↗

Oligonucleotide microarray identification of Bacillus anthracis strains using support vector machines.

The capability of a custom microarray to discriminate between closely related DNA samples is demonstrated using a set of Bacillus anthracis strains. The microarray was developed as a universal fingerprint device consisting of 390 genome-independent 9mer probes. The genomes of B. anthracis strains are monomorphic and therefore, typically difficult to distinguish using conventional molecular biology tools or microarray data clustering techniques. Using support vector machines (SVMs) as a supervised learning technique, we show that a low-density fingerprint microarray contains enough information to discriminate between B. anthracis strains with 90% sensitivity using a reference library constructed from six replicate arrays and three replicates for new isolates.

Algorithms↗

Multiple exemplar-based facial image retrieval using independent component analysis.

In this paper, we design a content-based image retrieval system where multiple query examples can be used to indicate the need to retrieve not only images similar to the individual examples, but also those images which actually represent a combination of the content of query images. We propose a scheme for representing content of an image as a combination of features from multiple examples. This scheme is exploited for developing a multiple example-based retrieval engine. We have explored the use of machine learning techniques for generating the most appropriate feature combination scheme for a given class of images. The combination scheme can be used for developing purposive query engines for specialized image databases. Here, we have considered facial image databases. The effectiveness of the image retrieval system is experimentally demonstrated on different databases.

Algorithms↗

Links between PPCA and subspace methods for complete Gaussian density estimation.

High-dimensional density estimation is a fundamental problem in pattern recognition and machine learning areas. In this letter, we show that, for complete high-dimensional Gaussian density estimation, two widely used methods, probabilistic principal component analysis and a typical subspace method using eigenspace decomposition, actually give the same results. Additionally, we present a unified view from the aspect of robust estimation of the covariance matrix.

Algorithms↗

Feature subset selection for support vector machines through discriminative function pruning analysis.

In many pattern classification applications, data are represented by high dimensional feature vectors, which induce high computational cost and reduce classification speed in the context of support vector machines (SVMs). To reduce the dimensionality of pattern representation, we develop a discriminative function pruning analysis (DFPA) feature subset selection method in the present study. The basic idea of the DFPA method is to learn the SVM discriminative function from training data using all input variables available first, and then to select feature subset through pruning analysis. In the present study, the pruning is implement using a forward selection procedure combined with a linear least square estimation algorithm, taking advantage of linear-in-the-parameter structure of the SVM discriminative function. The strength of the DFPA method is that it combines good characters of both filter and wrapper methods. Firstly, it retains the simplicity of the filter method avoiding training of a large number of SVM classifier. Secondly, it inherits the good performance of the wrapper method by taking the SVM classification algorithm into account.

Journal Article↗

Ensemble machine learning on gene expression data for cancer classification.

Whole genome RNA expression studies permit systematic approaches to understanding the correlation between gene expression profiles to disease states or different developmental stages of a cell. Microarray analysis provides quantitative information about the complete transcription profile of cells that facilitate drug and therapeutics development, disease diagnosis, and understanding in the basic cell biology. One of the challenges in microarray analysis, especially in cancerous gene expression profiles, is to identify genes or groups of genes that are highly expressed in tumour cells but not in normal cells and vice versa. Previously, we have shown that ensemble machine learning consistently performs well in classifying biological data. In this paper, we focus on three different supervised machine learning techniques in cancer classification, namely C4.5 decision tree, and bagged and boosted decision trees. We have performed classification tasks on seven publicly available cancerous microarray data and compared the classification/prediction performance of these methods. We have observed that ensemble learning (bagged and boosted decision trees) often performs better than single decision trees in this classification task.

Algorithms↗

Descriptor-based protein remote homology identification.

Here, we report a novel protein sequence descriptor-based remote homology identification method, able to infer fold relationships without the explicit knowledge of structure. In a first phase, we have individually benchmarked 13 different descriptor types in fold identification experiments in a highly diverse set of protein sequences. The relevant descriptors were related to the fold class membership by using simple similarity measures in the descriptor spaces, such as the cosine angle. Our results revealed that the three best-performing sets of descriptors were the sequence-alignment-based descriptor using PSI-BLAST e-values, the descriptors based on the alignment of secondary structural elements (SSEA), and the descriptors based on the occurrence of PROSITE functional motifs. In a second phase, the three top-performing descriptors were combined to obtain a final method with improved performance, which we named DescFold. Class membership was predicted by Support Vector Machine (SVM) learning. In comparison with the individual PSI-BLAST-based descriptor, the rate of remote homology identification increased from 33.7% to 46.3%. We found out that the composite set of descriptors was able to identify the true remote homolog for nearly every sixth sequence at the 95% confidence level, or some 10% more than a single PSI-BLAST search. We have benchmarked the DescFold method against several other state-of-the-art fold recognition algorithms for the 172 LiveBench-8 targets, and we concluded that it was able to add value to the existing techniques by providing a confident hit for at least 10% of the sequences not identifiable by the previously known methods.

Algorithms↗

[DNA chip data mining].

DNA chip data routinely contain gene expression levels of thousands of genes and the analysis should be supported by various computational tools. To be brief, the analysis procedure consists of four steps including image scanning, image processing, mathematical interpretation and biological interpretation. In image processing step, we should detect the spots and measure the signals of the spots and the background. In mathematical interpretation step, first of all we should massage the measured signals to make them appropriate for further mathematical analysis. The massaged data could be analyzed by various computational methods especially when the data were generated for multiple samples comparisons. The clustering techniques including hierarchical clustering, k-means clustering, SOTA, SOM are the most popular methods in this step. Various other multivariate statistics and related machine learning techniques are being introduced and applied to DNA chip data analysis recently. And finally the most important step we should tackle is the biological interpretation task. Although the depth of the domain knowledge about the biological situation under which the data were generated is the most important factor to elucidate the biological context, it could be supported by various bioinformatics tools including MEDLINE abstract processing by NLP techniques or genetic network models constructed by Boolean networks algorithms.

Algorithms↗

On the asymptotic distribution of the least-squares estimators in unidentifiable models.

In order to analyze the stochastic property of multilayered perceptrons or other learning machines, we deal with simpler models and derive the asymptotic distribution of the least-squares estimators of their parameters. In the case where a model is unidentified, we show different results from traditional linear models: the well-known property of asymptotic normality never holds for the estimates of redundant parameters.

Algorithms↗

The application of rule-based methods to class prediction problems in genomics.

We propose a method for constructing classifiers using logical combinations of elementary rules. The method is a form of rule-based classification, which has been widely discussed in the literature. In this work we focus specifically on issues that arise in the context of classifying cell samples based on RNA or protein expression measurements. The basic idea is to specify elementary rules that exhibit a locally strong pattern in favor of a single class. Strict admissibility criteria are imposed to produce a manageable universe of elementary rules. Then the elementary rules are combined using a set covering algorithm to form a composite rule that achieves a perfect fit to the training data. The user has explicit control over a parameter that determines the composite rule's level of redundancy and parsimony. This built-in control, along with the simplicity of interpreting the rules, makes the method particularly useful for classification problems in genomics. We demonstrate the new method using several microarray datasets and examine its generalization performance. We also draw comparisons to other machine-learning strategies such as CART, ID3, and C4.5.

Breast Neoplasms↗

Evaluating transmembrane topology prediction methods for the effect of signal peptide in topology prediction.

Reported performance of existing transmembrane (TM) topology prediction methods were often based on evaluations which neglected the risk of signal peptides (SP) being predicted as putative TM as well. Here, we evaluated 12 selected TM topology prediction methods (TMpred, TopPred II, DAS, TMAP, MEMSAT 2, SOSUI, PRED-TMR2, TMHMM 2.0, HMMTOP 2.0, SPLIT 3.5, TM Finder, and MPEx) for the effect of SP in prediction performance considering three SP treatments, namely: "remain" (untreated), "removed first", and "removed later". The results showed that the presence of SP significantly affected the prediction performance of the 12 selected TM topology prediction methods for all three predicted attributes (the number of transmembrane segments (TMSs), the number of TMSs plus position, and the N-tail location) and for the predicted topology (combined predictions of three attributes) by causing a reduction in prediction accuracy. In particular, lower prediction accuracies were obtained if SP is left untreated (remain) while significant increases were observed if SP is removed either first or later. However, between "removed first" and "removed later" SP treatments, the difference was statistically insignificant. In addition, we found that machine learning-based prediction methods were less affected by the presence of SP than hydropathy-based methods, but still the potential risk of degrading the prediction performance is there however to a lesser degree. Thus, when performing genome-wide analysis, the SP issue should be addressed during TM topology prediction.

Algorithms↗

Onvergence and application of online active sampling using orthogonal pillar vectors.

The analysis of convergence and its application is shown for the Active Sampling-at-the-Boundary method applied to multidimensional space using orthogonal pillar vectors. Active learning method facilitates identifying an optimal decision boundary for pattern classification in machine learning. The result of this method is compared with the standard active learning method that uses random sampling on the decision boundary hyperplane. The comparison is done through simulation and application to the real-world data from the UCI benchmark data set. The boundary is modeled as a nonseparable linear decision hyperplane in multidimensional space with a stochastic oracle.

Algorithms↗

Machine learning for an expert system to predict preterm birth risk.

OBJECTIVE: Develop a prototype expert system for preterm birth risk assessment of pregnant women. Normal gestation involves a term of 40 weeks, but because 8-12% of the newborns in the United States are delivered prior to 37 weeks' gestation, problems associated with prematurity continue to plague individuals, families, and the health care system. DESIGN: A knowledge-base development methodology used machine learning, statistical analysis, and validation techniques to analyze three large datasets (18,890 subjects and 214 variables). The dependent (i.e., decision) variable studied was weeks of gestation at delivery, with dichotomous coding of preterm delivery (prior to 37 weeks) and full-term delivery (37+ weeks). RESULTS: Machine learning with a program named Learning from Examples using Rough Sets (LERS) induced 520 usable rules that were entered into a prototype expert system. The prototype expert system was 53-88% accurate in predicting preterm delivery for 9,419 patients. CONCLUSION: The prototype expert system was more accurate than traditional manual techniques in predicting preterm birth.

Adult↗

A case study where biology inspired a solution to a computer science problem.

This paper describes how the biological theory of gene duplication described in Susumu Ohno's provocative book, Evolution by Means of Gene Duplication, was brought to bear on a vexatious problem from the domain of automated machine learning, namely the problem of architecture discovery. Six new architecture-altering operations for genetic programming were motivated by the way that new biological structures, functions, and behaviors arise in nature using gene duplication. Genetic programming with the new architecture-altering operations was then applied to the transmembrane protein segment identification problem. The out-of-sample error rate for the best genetically-evolved program achieved was slightly better than that of previously-reported human-written algorithms for this problem.

Amino Acid Sequence↗

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation↗