PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Clustering gene-expression data with repeated measurements.

Clustering is a common methodology for the analysis of array data, and many research laboratories are generating array data with repeated measurements. We evaluated several clustering algorithms that incorporate repeated measurements, and show that algorithms that take advantage of repeated measurements yield more accurate and more stable clusters. In particular, we show that the infinite mixture model-based approach with a built-in error model produces superior results.

Algorithms↗

Genomic comparison using data mining techniques based on a possibilistic fuzzy sets model.

Current copiousness of genomic information stored in biological databases [Mar Albà, M., Lee, M., Pearl, D., Shepherd, F.M.G., Martin, A.J., Orengo, N., Kellam, C.A., 2001. P. VIDA: a virus database system for the organisation of virus genome open reading frames. Nuleic Acids Res. 133-136] makes ultimately feasible the proposal for an application of knowledge management aimed to discover general rules in subcellular phenomena. The goal of this work is primarily to discover relationships between genes by microarray analysis. The tools exploited come from clustering techniques and are mainly based on Knowledge Discovery in Databases (KDD) concepts [Fayyad, U., Piatetsky-Shapiro, G., Smyth, P., 1996. From data mining to knowledge discovery in databases. AI Magazine 17(3), 37-54]. Starting from a data set, each element can be represented by a characteristic matrix, which sums up all data attributes. In this case data mining is oriented to perform a Pattern Recognition of related sequences, hidden in databases [Hand, D.J., Nicholas, A., 2005. Heard finding groups in gene expression data. J. Biomed. Biotechnol. 215-225]. Following a bottom up approach, the next refinement is to compare retrieved data to gather similar features, by dedicated clustering algorithms [Kaufman, L., Rousseeuw, P.J., 1990. Finding groups in data. An Introduction to Cluster Analysis. John Wiley & Sons, New York; Forman, G., Zhang, B., 2000. Distributed Data clustering can be efficient and exact HP. Laboratories Palo Alto HPL-2000, p. 158], driven by fuzzy logic, allowing us to perceive by intuition a common denominator for various genomic families and to anticipate likely future developments.

Algorithms↗

Molecular classification of breast carcinomas using tissue microarrays.

The histopathologic classification of breast cancer stratifies tumors based on tumor grade, stage, and type. Despite an overall correlation with survival, this classification is poorly predictive and tumors with identical grade and stage can have markedly contrasting outcomes. Recently, breast carcinomas have been classified by their gene expression profiles on frozen material. The validation of such a classification on formalin-fixed paraffin-embedded tumor archives linked to clinical information in a high-throughput fashion would have a major impact on clinical practice. The authors tested the ability of tumor tissue microarrays (TMAs) to sub-classify breast cancers using a TMA containing 107 breast cancers. The pattern of expression of 13 different protein biomarkers was assessed by immunohistochemistry and the multidimensional data was analyzed using an unsupervised two-dimensional clustering algorithm. This revealed distinct tumor clusters which divided into two main groups correlating with tumor grade (P<0.001) and nodal status (P = 0.04). None of the protein biomarkers tested could individually identify these groups. The biological significance of this classification is supported by its similarity with one derived from gene expression microarray analysis. Thus, molecular profiling of breast cancer using a limited number of protein biomarkers in TMAs can sub-classify tumors into clinically and biologically relevant subgroups.

Adenocarcinoma↗

Systematic analysis of domain motions in proteins from conformational change: new results on citrate synthase and T4 lysozyme.

Methods developed originally to analyze domain motions from simulation [Proteins 27:425-437, 1997] are adapted and extended for the analysis of X-ray conformers and for proteins with more than two domains. The method can be applied as an automatic procedure to any case where more than one conformation is available. The basis of the methodology is that domains can be recognized from the difference in the parameters governing their quasi-rigid body motion, and in particular their rotation vectors. A clustering algorithm is used to determine clusters of rotation vectors corresponding to main-chain segments that form possible dynamic domains. Domains are accepted for further analysis on the basis of a ratio of interdomain to intradomain fluctuation, and Chasles' theorem is used to determine interdomain screw axes. Finally residues involved in the interdomain motion are identified. The methodology is tested on citrate synthase and the M6I mutant of T4 lysozyme. In both cases new aspects to their conformational change are revealed, as are individual residues intimately involved in their dynamics. For citrate synthase the beta sheet is identified to be part of the hinging mechanism. In the case of T4 lysozyme, one of the four transitions in the pathway from the closed to the open conformation, furnished four dynamic domains rather than the expected two. This result indicates that the number of dynamic domains a protein possesses may not be a constant of the motion.

Algorithms↗

Interlaboratory comparative study of the numerical analysis of one-dimensional sodium dodecyl sulphate-polyacrylamide gel electrophoretic protein patterns of Campylobacter strains.

Twenty-nine bacterial strains of the genus Campylobacter were examined independently at two collaborating institutes, the National Collection of Type Cultures, London and the Laboratorium voor Microbiologie, Rijksuniversiteit Gent. The one-dimensional polyacrylamide gel electrophoretic protein patterns of the strains were analysed using computerised numerical methods which employed a correlation coefficient and a clustering algorithm. The electrophoretic methods used at the two institutes included both major differences such as gel composition and running conditions and minor differences in buffer composition. Although the algorithm on which similarity and clustering were computed were the same, the detailed treatment of scan patterns differed. The resulting protein patterns in the gels differed markedly in appearance but after numerical analysis the two systems were equally effective in their ability to speciate the strains. There were, however, differences in the relationships between the species defined at the two institutes and these were at least partly due to the different background subtraction methods employed. In conclusion, the portability and reproducibility of the two systems for identification was demonstrated but for definitive classification further standardization may be required.

Bacterial Proteins↗

Comparing the similarity of time-series gene expression using signal processing metrics.

Many algorithms have been used to cluster genes measured by microarray across a time series. Instead of clustering, our goal was to compare all pairs of genes to determine whether there was evidence of a phase shift between them. We describe a technique where gene expression is treated as a discrete time-invariant signal, allowing the use of digital signal-processing tools, including power spectral density, coherence, and transfer gain and phase shift. We used these on a public RNA expression set of 2467 genes measured every 7 min for 119 min and found 18 putative associations. Two of these were known in the biomedical literature and may have been missed using correlation coefficients. Digital signal processing tools can be embedded and enhance existing clustering algorithms.

Algorithms↗

Radial basis function neural networks in non-destructive determination of compound aspirin tablets on NIR spectroscopy.

The application of the second most popular artificial neural networks (ANNs), namely, the radial basis function (RBF) networks, has been developed for quantitative analysis of drugs during the last decade. In this paper, the two components (aspirin and phenacetin) were simultaneously determined in compound aspirin tablets by using near-infrared (NIR) spectroscopy and RBF networks. The total database was randomly divided into a training set (50) and a testing set (17). Different preprocessing methods (standard normal variate (SNV), multiplicative scatter correction (MSC), first-derivative and second-derivative) were applied to two sets of NIR spectra of compound aspirin tablets with different concentrations of two active components and compared each other. After that, the performance of RBF learning algorithm adopted the nearest neighbor clustering algorithm (NNCA) and the criterion for selection used a cross-validation technique. Results show that using RBF networks to quantificationally analyze tablets is reliable, and the best RBF model was obtained by first-derivative spectra.

Aspirin↗

Large-scale identification of single-feature polymorphisms in complex genomes.

We have developed a high-throughput genotyping platform by hybridizing genomic DNA from Arabidopsis thaliana accessions to an RNA expression GeneChip (AtGenome1). Using newly developed analytical tools, a large number of single-feature polymorphisms (SFPs) were identified. A comparison of two accessions, the reference strain Columbia (Col) and the strain Landsberg erecta (Ler), identified nearly 4000 SFPs, which could be reliably scored at a 5% error rate. Ler sequence was used to confirm 117 of 121 SFPs and to determine the sensitivity of array hybridization. Features containing sequence repeats, as well as those from high copy genes, showed greater polymorphism rates. A linear clustering algorithm was developed to identify clusters of SFPs representing potential deletions in 111 genes at a 5% false discovery rate (FDR). Among the potential deletions were transposons, disease resistance genes, and genes involved in secondary metabolism. The applicability of this technique was demonstrated by genotyping a recombinant inbred line. Recombination break points could be clearly defined, and in one case delimited to an interval of 29 kb. We further demonstrate that array hybridization can be combined with bulk segregant analysis to quickly map mutations. The extension of these tools to organisms with complex genomes, such as Arabidopsis, will greatly increase our ability to map and clone quantitative trait loci (QTL).

Arabidopsis↗

Network analysis of online bidding activity.

With the advent of digital media, people are increasingly resorting to online channels for commercial transactions. The online auction is a prototypical example. In such online transactions, the pattern of bidding activity is more complex than traditional offline transactions; this is because the number of bidders participating in a given transaction is not bounded and the bidders can also easily respond to the bidding instantaneously. By using the recently developed network theory, we study the interaction patterns between bidders (items) who (that) are connected when they bid for the same item (if the item is bid by the same bidder). The resulting network is analyzed by using the hierarchical clustering algorithm, which is used for clustering analysis for expression data from DNA microarrays. A dendrogram is constructed for the item subcategories; this dendrogram is compared to a traditional classification scheme. The implication of the difference between the two is discussed.

Journal Article↗

Exploring supervised and unsupervised methods to detect topics in biomedical text.

BACKGROUND: Topic detection is a task that automatically identifies topics (e.g., "biochemistry" and "protein structure") in scientific articles based on information content. Topic detection will benefit many other natural language processing tasks including information retrieval, text summarization and question answering; and is a necessary step towards the building of an information system that provides an efficient way for biologists to seek information from an ocean of literature. RESULTS: We have explored the methods of Topic Spotting, a task of text categorization that applies the supervised machine-learning technique naïve Bayes to assign automatically a document into one or more predefined topics; and Topic Clustering, which apply unsupervised hierarchical clustering algorithms to aggregate documents into clusters such that each cluster represents a topic. We have applied our methods to detect topics of more than fifteen thousand of articles that represent over sixteen thousand entries in the Online Mendelian Inheritance in Man (OMIM) database. We have explored bag of words as the features. Additionally, we have explored semantic features; namely, the Medical Subject Headings (MeSH) that are assigned to the MEDLINE records, and the Unified Medical Language System (UMLS) semantic types that correspond to the MeSH terms, in addition to bag of words, to facilitate the tasks of topic detection. Our results indicate that incorporating the MeSH terms and the UMLS semantic types as additional features enhances the performance of topic detection and the naïve Bayes has the highest accuracy, 66.4%, for predicting the topic of an OMIM article as one of the total twenty-five topics. CONCLUSION: Our results indicate that the supervised topic spotting methods outperformed the unsupervised topic clustering; on the other hand, the unsupervised topic clustering methods have the advantages of being robust and applicable in real world settings.

Abstracting and Indexing↗

The P-A-I-N MMPI classification system: a critical review.

The Costello et al. (Pain, 30 (1987) 199-209) literature-based MMPI clustering algorithm was compared to a standard clustering procedure of chronic pain patients' MMPI profiles. Results indicated that the Costello algorithm was too restrictive, failing to classify 69% of the MMPI profiles of the local sample. It was suggested that it may be premature to adopt a literature based clustering method until the validity of empirically derived clusters has been more thoroughly determined. The utility of a given set of clusters may best be determined by using locally derived clusters to predict treatment response. Once the predictive validity of the local clusters has been demonstrated, the Costello approach may prove more useful.

Adolescent↗

Tandem machine learning for the identification of genes regulated by transcription factors.

BACKGROUND: The identification of promoter regions that are regulated by a given transcription factor has traditionally relied upon the identification and distributions of binding sites recognized by the factor. In this study, we have developed a tandem machine learning approach for the identification of regulatory target genes based on these parameters and on the corresponding binding site information contents that measure the affinities of the factor for these cognate elements. RESULTS: This method has been validated using models of DNA binding sites recognized by the xenobiotic-sensitive nuclear receptor, PXR/RXRalpha, for target genes within the human genome. An information theory-based weight matrix was first derived and refined from known PXR/RXRalpha binding sites. The promoter region of candidate genes was scanned with the weight matrix. A novel information density-based clustering algorithm was then used to identify clusters of information rich sites. Finally, transformed data representing metrics of location, strength and clustering of binding sites were used for classification of promoter regions using an ensemble approach involving neural networks, decision trees and Naïve Bayesian classification. The method was evaluated on a set of 24 known target genes and 288 genes known not to be regulated by PXR/RXRalpha. We report an average accuracy (proportion of correctly classified promoter regions) of 71%, sensitivity of 73%, and specificity of 70%, based on multiple cross-validation and the leave-one-out strategy. The performance on a test set of 13 genes showed that 10 were correctly classified. CONCLUSION: We have developed a machine learning approach for the successful detection of gene targets for transcription factors with high accuracy. The method has been validated for the transcription factor PXR/RXRalpha and has the potential to be extended to other transcription factors.

Algorithms↗

Cluster analysis of fMRI data using dendrogram sharpening.

The major disadvantage of hierarchical clustering in fMRI data analysis is that an appropriate clustering threshold needs to be specified. Upon grouping data into a hierarchical tree, clusters are identified either by specifying their number or by choosing an appropriate inconsistency coefficient. Since the number of clusters present in the data is not known beforehand, even a slight variation of the inconsistency coefficient can significantly affect the results. To address these limitations, the dendrogram sharpening method, combined with a hierarchical clustering algorithm, is used in this work to identify modality regions, which are, in essence, areas of activation in the human brain during an fMRI experiment. The objective of the algorithm is to remove data from the low-density regions in order to obtain a clearer representation of the data structure. Once cluster cores are identified, the classification algorithm is run on voxels, set aside during sharpening, attempting to reassign them to the detected groups. When applied to a paced motor paradigm, task-related activations in the motor cortex are detected. In order to evaluate the performance of the algorithm, the obtained clusters are compared to standard activation maps where the expected hemodynamic response function is specified as a regressor. The obtained patterns of both methods have a high concordance (correlation coefficient = 0.91). Furthermore, the dependence of the clustering results on the sharpening parameters is investigated and recommendations on the appropriate choice of these variables are offered. Hum. Brain Mapping 20:201-219, 2003.

Algorithms↗

A visual fuzzy cluster system for patient analysis.

A visual fuzzy cluster (VFC) system is developed to assist physicians in interactively detecting and refining cluster partitions for patients with acute upper respiratory infections. The VFC system assists physicians in discovering relationships among patients by applying a fuzzy cluster algorithm to analyse a case base of patient findings. The algorithm discovers similarities among patients, while at the same time identifying atypical patients. The system visually presents the fuzzy cluster solutions on a three-dimensional animated display. Physicians then interactively manipulate icons representing patients to explore and refine the fuzzy cluster solution. Initial experiences with the VFC prototype are encouraging and support the claims that the system improves physician understanding and allows physicians to take advantage of visual recognition and manipulation skills to define and label patient groupings. The resulting labels and cluster centres or prototypes offer insight into the set of patient features that best discriminate between groups.

Algorithms↗

Detection of multiple sclerosis with visual evoked potentials--an unsupervised computational intelligence system.

This paper describes the application of a novel unsupervised pattern recognition system to the classification of the Visual Evoked Potentials (VEP's) of normal and multiple sclerosis (MS) patients. The method combines a traditional statistical feature extractor with a fuzzy clustering method, all implemented in a parallel neural network architecture. The optimization routine, ALOPEX, is used to train the network while decreasing the likelihood of local solutions. The unsupervised system includes a feature extraction and clustering module, trained by the optimization routine ALOPEX. Through maximization of the output variance of each node, and an architecture which excludes redundancy, the feature extraction network retains the most significant Karhunen-Loeve expansion vectors. The clustering module uses a modification to the Fuzzy c-Means (FCM) clustering algorithms, where ALOPEX adjusts a set of cluster centers to minimize an objective error function. The result combines the power of the FCM algorithms with the advantage of a more global solution from ALOPEX. The new pattern recognition system is used to cluster the VEP's of 13 normal and 12 MS subjects. The classification with this technique can, without supervision, separate the patient population into two groups which largely correspond to the MS and control subject groups. A suitable threshold can be chosen so that the recognizer chooses no false negatives. The use of multiple stimulation patterns appears to improve the reliability of the decision. The reasoning of most neural networks in their decision making cannot easily be extracted upon the completion of training. However, due to the linearity of the network nodes, the cluster prototypes of this unsupervised system can be reconstructed to illustrate the reasoning of the system. In this application, this analysis hints at the usefulness of previously unused portions of the VEP in detecting MS. It also indicates a possible use of the system as a training aide.

Adult↗

An On-line agglomerative clustering method for nonstationary data.

An on-line agglomerative clustering algorithm for nonstationary data is described. Three issues are addressed. The first regards the temporal aspects of the data. The clustering of stationary data by the proposed algorithm is comparable to the other popular algorithms tested (batch and on-line). The second issue addressed is the number of clusters required to represent the data. The algorithm provides an efficient framework to determine the natural number of clusters given the scale of the problem. Finally, the proposed algorithm implicitly minimizes the local distortion, a measure that takes into account clusters with relatively small mass. In contrast, most existing on-line clustering methods assume stationarity of the data. When used to cluster nonstationary data, these methods fail to generate a good representation. Moreover, most current algorithms are computationally intensive when determining the correct number of clusters. These algorithms tend to neglect clusters of small mass due to their minimization of the global distortion (Energy).

Algorithms↗

A hierarchical clustering approach for large compound libraries.

A modified version of the k-means clustering algorithm was developed that is able to analyze large compound libraries. A distance threshold determined by plotting the sum of radii of leaf clusters was used as a termination criterion for the clustering process. Hierarchical trees were constructed that can be used to obtain an overview of the data distribution and inherent cluster structure. The approach is also applicable to ligand-based virtual screening with the aim to generate preferred screening collections or focused compound libraries. Retrospective analysis of two activity classes was performed: inhibitors of caspase 1 [interleukin 1 (IL1) cleaving enzyme, ICE] and glucocorticoid receptor ligands. The MDL Drug Data Report (MDDR) and Collection of Bioactive Reference Analogues (COBRA) databases served as the compound pool, for which binary trees were produced. Molecules were encoded by all Molecular Operating Environment 2D descriptors and topological pharmacophore atom types. Individual clusters were assessed for their purity and enrichment of actives belonging to the two ligand classes. Significant enrichment was observed in individual branches of the cluster tree. After clustering a combined database of MDDR, COBRA, and the SPECS catalog, it was possible to retrieve MDDR ICE inhibitors with new scaffolds using COBRA ICE inhibitors as seeds. A Java implementation of the clustering method is available via the Internet (http://www.modlab.de).

Algorithms↗

Automatic document classification of biological literature.

BACKGROUND: Document classification is a wide-spread problem with many applications, from organizing search engine snippets to spam filtering. We previously described Textpresso, a text-mining system for biological literature, which marks up full text according to a shallow ontology that includes terms of biological interest. This project investigates document classification in the context of biological literature, making use of the Textpresso markup of a corpus of Caenorhabditis elegans literature. RESULTS: We present a two-step text categorization algorithm to classify a corpus of C. elegans papers. Our classification method first uses a support vector machine-trained classifier, followed by a novel, phrase-based clustering algorithm. This clustering step autonomously creates cluster labels that are descriptive and understandable by humans. This clustering engine performed better on a standard test-set (Reuters 21578) compared to previously published results (F-value of 0.55 vs. 0.49), while producing cluster descriptions that appear more useful. A web interface allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept. CONCLUSION: We have demonstrated a simple method to classify biological documents that embodies an improvement over current methods. While the classification results are currently optimized for Caenorhabditis elegans papers by human-created rules, the classification engine can be adapted to different types of documents. We have demonstrated this by presenting a web interface that allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept.

Abstracting and Indexing↗