PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Specialized microbial databases for inductive exploration of microbial genome sequences.

BACKGROUND: The enormous amount of genome sequence data asks for user-oriented databases to manage sequences and annotations. Queries must include search tools permitting function identification through exploration of related objects. METHODS: The GenoList package for collecting and mining microbial genome databases has been rewritten using MySQL as the database management system. Functions that were not available in MySQL, such as nested subquery, have been implemented. RESULTS: Inductive reasoning in the study of genomes starts from "islands of knowledge", centered around genes with some known background. With this concept of "neighborhood" in mind, a modified version of the GenoList structure has been used for organizing sequence data from prokaryotic genomes of particular interest in China. GenoChore http://bioinfo.hku.hk/genochore.html, a set of 17 specialized end-user-oriented microbial databases (including one instance of Microsporidia, Encephalitozoon cuniculi, a member of Eukarya) has been made publicly available. These databases allow the user to browse genome sequence and annotation data using standard queries. In addition they provide a weekly update of searches against the world-wide protein sequences data libraries, allowing one to monitor annotation updates on genes of interest. Finally, they allow users to search for patterns in DNA or protein sequences, taking into account a clustering of genes into formal operons, as well as providing extra facilities to query sequences using predefined sequence patterns. CONCLUSION: This growing set of specialized microbial databases organize data created by the first Chinese bacterial genome programs (ThermaList, Thermoanaerobacter tencongensis, LeptoList, with two different genomes of Leptospira interrogans and SepiList, Staphylococcus epidermidis) associated to related organisms for comparison.

Algorithms↗

Combination of automated high throughput platforms, flow cytometry, and hierarchical clustering to detect cell state.

BACKGROUND: This study examined whether hierarchical clustering could be used to detect cell states induced by treatment combinations that were generated through automation and high-throughput (HT) technology. Data-mining techniques were used to analyze the large experimental data sets to determine whether nonlinear, non-obvious responses could be extracted from the data. METHODS: Unary, binary, and ternary combinations of pharmacological factors (examples of stimuli) were used to induce differentiation of HL-60 cells using a HT automated approach. Cell profiles were analyzed by incorporating hierarchical clustering methods on data collected by flow cytometry. Data-mining techniques were used to explore the combinatorial space for nonlinear, unexpected events. Additional small-scale, follow-up experiments were performed on cellular profiles of interest. RESULTS: Multiple, distinct cellular profiles were detected using hierarchical clustering of expressed cell-surface antigens. Data-mining of this large, complex data set retrieved cases of both factor dominance and cooperativity, as well as atypical cellular profiles. Follow-up experiments found that treatment combinations producing "atypical cell types" made those cells more susceptible to apoptosis. CONCLUSIONS Hierarchical clustering and other data-mining techniques were applied to analyze large data sets from HT flow cytometry. From each sample, the data set was filtered and used to define discrete, usable states that were then related back to their original formulations. Analysis of resultant cell populations induced by a multitude of treatments identified unexpected phenotypes and nonlinear response profiles.

Algorithms↗

Segmentation and classification of EEG during epileptic seizures.

We present a method for the automatic comparison of epileptic seizures in EEG, allowing the grouping of seizures having similar overall patterns. Each channel of the EEG is first broken down into segments having relatively stationary characteristics. Features are then calculated for each segment and all segments of all channels of the seizures of one patient are grouped into clusters of similar morphology. This clustering allows labeling of every EEG segment. Methods derived from string matching procedures are then used to obtain an overall edit distance between two seizures, a distance that represents how the two seizures, taken in their entirety and including the channels not actually involved in the discharge, resemble each other. Examples from 5 patients, 3 with intracerebral electrodes and two with scalp electrodes, illustrate the ability of the method to group seizures of similar morphology.

Algorithms↗

Statistically rigorous automated protein annotation.

MOTIVATION: Assignment of putative protein functional annotation by comparative analysis using pre-defined experimental annotations is performed routinely by molecular biologists. The number and statistical significance of these assignments remains a challenge in this era of high-throughput proteomics. A combined statistical method that enables robust, automated protein annotation by reliably expanding existing annotation sets is described. An existing clustering scheme, based on relevant experimental information (e.g. sequence identity, keywords or gene expression data) is required. The method assigns new proteins to these clusters with a measure of reliability. It can also provide human reviewers with a reliability score for both new and previously classified proteins. RESULTS: A dataset of 27 000 annotated Protein Data Bank (PDB) polypeptide chains (of 36 000 chains currently in the PDB) was generated from 23 000 chains classified a priori. AVAILABILITY: PDB annotations and sample software implementation are freely accessible on the Web at http://pmr.sdsc.edu/go

Abstracting and Indexing↗

Iterative cluster analysis of protein interaction data.

MOTIVATION: Generation of fast tools of hierarchical clustering to be applied when distances among elements of a set are constrained, causing frequent distance ties, as happens in protein interaction data. RESULTS: We present in this work the program UVCLUSTER, that iteratively explores distance datasets using hierarchical clustering. Once the user selects a group of proteins, UVCLUSTER converts the set of primary distances among them (i.e. the minimum number of steps, or interactions, required to connect two proteins) into secondary distances that measure the strength of the connection between each pair of proteins when the interactions for all the proteins in the group are considered. We show that this novel strategy has advantages over conventional clustering methods to explore protein-protein interaction data. UVCLUSTER easily incorporates the information of the largest available interaction datasets to generate comprehensive primary distance tables. The versatility, simplicity of use and high speed of UVCLUSTER on standard personal computers suggest that it can be a benchmark analytical tool for interactome data analysis. AVAILABILITY: The program is available upon request from the authors, free for academic users. Additional information available at http://www.uv.es/genomica/UVCLUSTER.

Actins↗

An adaptive meta-clustering approach: combining the information from different clustering results.

With the development of microarray techniques, there is an increasing need of information processing methods to analyze the high throughput data. Clustering is one of the most promising candidates because of its simplicity, flexibility and robustness. However, there is no "perfect" clustering approach outperforming its counterparts, and it is hard to evaluate and combine the results from different techniques, especially in a field without much prior knowledge, such as bioinformatics. This paper proposes a meta-clustering approach to extract the information from results of different clustering techniques, so that a better interpretation of the data distribution can be obtained. A special distance measure is defined to represent the statistical "signal" of each cluster produced by various clustering techniques. The algorithm is applied on both artificial and real data Simulations show that the proposed approach is able to extract the information efficiently and accurately from the input clustering structure.

Algorithms↗

Using supervised fuzzy clustering to predict protein structural classes.

Prediction of protein classification is both an important and a tempting topic in protein science. This is because of not only that the knowledge thus obtained can provide useful information about the overall structure of a query protein, but also that the practice itself can technically stimulate the development of novel predictors that may be straightforwardly applied to many other relevant areas. In this paper, a novel approach, the so-called "supervised fuzzy clustering approach" is introduced that is featured by utilizing the class label information during the training process. Based on such an approach, a set of "if-then" fuzzy rules for predicting the protein structural classes are extracted from a training dataset. It has been demonstrated through two different working datasets that the overall success prediction rates obtained by the supervised fuzzy clustering approach are all higher than those by the unsupervised fuzzy c-means introduced by the previous investigators [C.T. Zhang, K.C. Chou, G.M. Maggiora. Protein Eng. (1995) 8, 425-435]. It is anticipated that the current predictor may play an important complementary role to other existing predictors in this area to further strengthen the power in predicting the structural classes of proteins and their other characteristic attributes.

Algorithms↗

Analysis of topological and nontopological structural similarities in the PDB: new examples with old structures.

We have developed a new method and program, SARF2, for fast comparison of protein structures, which can detect topological as well as nontopological similarities. The method searches for large ensembles of secondary structure elements, which are mutually compatible in two proteins. These ensembles consist of small fragments of C alpha-trace, similarly arranged in three-dimensional space in two proteins, but not necessarily equally-ordered along the polypeptide chains. The program SARF2 is available for everyone through the World-Wide Web (WWW). We have performed an exhaustive pairwise comparison of all the entries from a recent issue of the Protein Data Bank (PDB) and report here on the results of an automated hierarchical cluster analysis. In addition, we report on several new cases of significant structural resemblance between proteins. To this end, a new definition of the significance of structural similarity is introduced, which effectively distinguishes the biologically meaningful equivalences from those occurring by chance. Analyzing the distribution of sequence similarity in significant structural matches, we show that sequence similarity as low as 20% in structurally-prealigned proteins can be a strong indication for the biological relevance of structural similarity.

Algorithms↗

AMDA: an R package for the automated microarray data analysis.

BACKGROUND: Microarrays are routinely used to assess mRNA transcript levels on a genome-wide scale. Large amount of microarray datasets are now available in several databases, and new experiments are constantly being performed. In spite of this fact, few and limited tools exist for quickly and easily analyzing the results. Microarray analysis can be challenging for researchers without the necessary training and it can be time-consuming for service providers with many users. RESULTS: To address these problems we have developed an automated microarray data analysis (AMDA) software, which provides scientists with an easy and integrated system for the analysis of Affymetrix microarray experiments. AMDA is free and it is available as an R package. It is based on the Bioconductor project that provides a number of powerful bioinformatics and microarray analysis tools. This automated pipeline integrates different functions available in the R and Bioconductor projects with newly developed functions. AMDA covers all of the steps, performing a full data analysis, including image analysis, quality controls, normalization, selection of differentially expressed genes, clustering, correspondence analysis and functional evaluation. Finally a LaTEX document is dynamically generated depending on the performed analysis steps. The generated report contains comments and analysis results as well as the references to several files for a deeper investigation. CONCLUSION: AMDA is freely available as an R package under the GPL license. The package as well as an example analysis report can be downloaded in the Services/Bioinformatics section of the Genopolis http://www.genopolis.it/.

Algorithms↗

Computationally efficient cluster representation in molecular sequence megaclassification.

Molecular sequence megaclassification is a technique for automated protein sequence analysis and annotation. Implementation of the method has been limited by the need to store and randomly access a database of all the sequence pair similarities. More than 80,000 protein sequences are now present in the public databases, and the pair similarity data table for the full protein sequence database requires over 1 gigabyte of storage. In this paper we present a computationally efficient representation of groups based on a graph theory approach where sequence clusters are described by a minimal spanning tree of highest scoring similarity pairs. This representation allows a classification of N proteins to be stored in order(N) memory. The use of this minimal spanning tree representation simplifies analysis of groups, the description of group characteristics and the manual correction of artifacts resulting from false hits. The new tree representation also introduces new possibilities for artifact generation in sequence classification. Methods for detecting and removing these artifacts are discussed.

Algorithms↗

A tree-based decision rule for identifying profile groups of cases without predefined classes: application in diffuse large B-cell lymphomas.

In this paper, we examined the utility of a forward growing classification tree as a supplement to cluster analysis for deriving a decision rule for the identification of profile groups when the cases do not belong to predefined classes. The technique was applied for the identification of low and high proliferation profile groups of diffuse large B-cell lymphomas according to the immunohistochemical expression levels of proliferation proteins. In a forward growing classification tree method, the size of the tree is controlled by the improvement (threshold value) in the apparent misclassification rate after each split. The classes used in the tree were defined using k-means clustering. The decision rule consisted of the splitting points of the split variables used. The methodology was applied to the histology data from 79 cases of diffuse large B-cell lymphomas. Ten classes of individual cases were derived from k-means clustering. Then, a classification tree with a threshold of 2% was used to derive the decision rule. Branches at the left side of the tree consisted of individuals with a low proliferation profile and branches at the right side of the tree consisted of cases with a high proliferation profile. The classification tree, as a supplement method, not only identified but also provided decision rules for identifying profile groups. Finally, it also allowed for exploration of the data structure.

Algorithms↗

Detecting interspecific recombination with a pruned probabilistic divergence measure.

MOTIVATION: A promising sliding-window method for the detection of interspecific recombination in DNA sequence alignments is based on the monitoring of changes in the posterior distribution of tree topologies with a probabilistic divergence measure. However, as the number of taxa in the alignment increases or the sliding-window size decreases, the posterior distribution becomes increasingly diffuse. This diffusion blurs the probabilistic divergence signal and adversely affects the detection accuracy. The present study investigates how this shortcoming can be redeemed with a pruning method based on post-processing clustering, using the Robinson-Foulds distance as a metric in tree topology space. RESULTS: An application of the proposed scheme to three synthetic and two real-world DNA sequence alignments illustrates the amount of improvement that can be obtained with the pruning method. The study also includes a comparison with two established recombination detection methods: Recpars and the DSS (difference of sum of squares) method. AVAILABILITY: Software, data and further supplementary material are available at the following website: http://www.bioss.sari.ac.uk/~dirk/Supplements/

Algorithms↗

Extension neural network-type 2 and its applications.

A supervised learning pattern classifier, called the extension neural network (ENN), has been described in a recent paper. In this sequel, the unsupervised learning pattern clustering sibling called the extension neural network type 2 (ENN-2) is proposed. This new neural network uses an extension distance (ED) to measure the similarity between data and the cluster center. It does not require an initial guess of the cluster center coordinates, nor of the initial number of clusters. The clustering process is controlled by a distanced parameter and by a novel extension distance. It shows the same capability as human memory systems to keep stability and plasticity characteristics at the same time, and it can produce meaningful weights after learning. Moreover, the structure of the proposed ENN-2 is simpler and the learning time is shorter than traditional neural networks. Experimental results from five different examples, including three benchmark data sets and two practical applications, verify the effectiveness and applicability of the proposed work.

Algorithms↗

EMD: an ensemble algorithm for discovering regulatory motifs in DNA sequences.

BACKGROUND: Understanding gene regulatory networks has become one of the central research problems in bioinformatics. More than thirty algorithms have been proposed to identify DNA regulatory sites during the past thirty years. However, the prediction accuracy of these algorithms is still quite low. Ensemble algorithms have emerged as an effective strategy in bioinformatics for improving the prediction accuracy by exploiting the synergetic prediction capability of multiple algorithms. RESULTS: We proposed a novel clustering-based ensemble algorithm named EMD for de novo motif discovery by combining multiple predictions from multiple runs of one or more base component algorithms. The ensemble approach is applied to the motif discovery problem for the first time. The algorithm is tested on a benchmark dataset generated from E. coli RegulonDB. The EMD algorithm has achieved 22.4% improvement in terms of the nucleotide level prediction accuracy over the best stand-alone component algorithm. The advantage of the EMD algorithm is more significant for shorter input sequences, but most importantly, it always outperforms or at least stays at the same performance level of the stand-alone component algorithms even for longer sequences. CONCLUSION: We proposed an ensemble approach for the motif discovery problem by taking advantage of the availability of a large number of motif discovery programs. We have shown that the ensemble approach is an effective strategy for improving both sensitivity and specificity, thus the accuracy of the prediction. The advantage of the EMD algorithm is its flexibility in the sense that a new powerful algorithm can be easily added to the system.

Algorithms↗

Rule extraction from a mutagenicity data set using adaptively grown phylogenetic-like trees.

A public bacterial mutagenicity database was classified into 2-D structural families using a set of specific algorithms and clustering techniques that find overlapping classes of compounds based upon chemical substructures. Structure-activity relationships were learned from the biological activity of the compounds within each class and used to identify rules that define substructures potentially responsible for mutagenic activity. In addition, this method of analysis was used to compare the pharmacologically relevant substructure of test compounds with their potential toxic substructures making this a potentially valuable in silico profiling tool for lead selection and optimization.

Algorithms↗

Evidence for radical anion formation during liquid secondary ion mass spectrometry analysis of oligonucleotides and synthetic oligomeric analogues: a deconvolution algorithm for molecular ion region clusters.

It is shown that one-electron reduction is a common process that occurs in negative ion liquid secondary ion mass spectrometry (LSIMS) of oligonucleotides and synthetic oligonucleosides and that this process is in competition with proton loss. Deconvolution of the molecular anion cluster reveals contributions from (M-2H).-, (M-H)-, M.-, and (M + H)-. A model based on these ionic species gives excellent agreement with the experimental data. A correlation between the concentration of species arising via one-electron reduction [M.- and (M + H)-] and the electron affinity of the matrix has been demonstrated. The relative intensity of M.- is mass-dependent; this is rationalized on the basis of base-stacking. Base sequence ion formation is theorized to arise from M.- radical anion among other possible pathways.

Algorithms↗

2HAPI: a microarray data analysis system.

SUMMARY: 2HAPI (version 2 of High density Array Pattern Interpreter) is a web-based, publicly-available analytical tool designed to aid researchers in microarray data analysis. 2HAPI includes tools for searching, manipulating, visualizing, and clustering the large sets of data generated by microarray experiments. Other features include association of genes with NCBI information and linkage to external data resources. Unique to 2HAPI is the ability to retrieve upstream sequences of co-regulated genes for promoter analysis using MEME (Multiple Expectation-maximization for Motif Elicitation) AVAILABILITY: 2HAPI is freely available at http://array.sdsc.edu. Users can try 2HAPI anonymously with pre-loaded data or they can register as a 2HAPI user and upload their data.

Algorithms↗

Protein profiling in brain tumors using mass spectrometry: feasibility of a new technique for the analysis of protein expression.

PURPOSE: The purpose of this research was to perform a preliminary assessment of protein patterns in primary brain tumors using a direct-tissue mass spectrometric technique to profile and map biomolecules. EXPERIMENTAL DESIGN: We examined 20 prospectively collected, snap-frozen normal brain and brain tumor specimens using matrix-assisted laser desorption/ionization (MALDI) mass spectrometry (MS), and compared peptide and protein expression in primary brain tumor and nontumor brain tissues. RESULTS: MS can be used to identify protein expression patterns in human brain tissue and tumor specimens. The mass spectral patterns can reliably identify glial neoplasms of similar histological grade and differentiate them from tumors of different histological grades as well as from nontumor brain tissues. Initial bioinformatics cluster analysis algorithms classified tumor and nontumor tissues into similar groups comparable with their histological grade. CONCLUSIONS: We describe a novel tool for the analysis of protein expression patterns in human glial neoplasms. Initial results demonstrate that MALDI-MS technology can significantly aid in the process of unraveling and understanding the molecular complexities of gliomas. MALDI-MS accurately and reliably identified normal and neoplastic tissues, and could be used to discriminate between tumors of increasing grades.

Adult↗