PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “hierarchical clustering”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Introduction to hierarchical clustering.

Hierarchical clustering of spike events is a method of grouping events that are similar in topology, morphology, or both, and it provides a method of efficient, detailed analysis of interictal events. Information about the relative populations of spikes at multiple foci is presented, and artifact events are grouped and eliminated en masse. The process of hierarchical clustering is explained, and a set of simulated traces is used to illustrate the process of hierarchical clustering and the development of a cluster tree to display the relative populations of similar spike events. Using EEG data from long-term monitoring, the use of a "review wizard" is explored as a means of structuring the process of hierarchical clustering and traversing the cluster tree. This aid is also used to streamline the process of determining the similarity of events within each group and of verifying that events exhibiting clinically important differences are not hidden within the groups comprising the average traces.

Cluster Analysis↗

Selection of informative clusters from hierarchical cluster tree with gene classes.

BACKGROUND: A common clustering method in the analysis of gene expression data has been hierarchical clustering. Usually the analysis involves selection of clusters by cutting the tree at a suitable level and/or analysis of a sorted gene list that is obtained with the tree. Cutting of the hierarchical tree requires the selection of a suitable level and it results in the loss of information on the other level. Sorted gene lists depend on the sorting method of the joined clusters. Author proposes that the clusters should be selected using the gene classifications. RESULTS: This article presents a simple method for searching for clusters with the strongest enrichment of gene classes from a cluster tree. The clusters found are presented in the estimated order of importance. The method is demonstrated with a yeast gene expression data set and with two database classifications. The obtained clusters demonstrated a very strong enrichment of functional classes. The obtained clusters are also able to present similar gene groups to those that were observed from the data set in the original analysis and also many gene groups that were not reported in the original analysis. Visualization of the results on top of a cluster tree shows that the method finds informative clusters from several levels of the cluster tree and indicates that the clusters found could not have been obtained by simply cutting the cluster tree. Results were also used in the comparison of cluster trees from different clustering methods. CONCLUSION: The presented method should facilitate the exploratory analysis of big data sets when the associated categorical data is available.

Cluster Analysis↗

Chaotic cluster itinerancy and hierarchical cluster trees in electrochemical experiments.

Experiments on an array of 64 globally coupled chaotic electrochemical oscillators were carried out. The array is heterogeneous due to small variations in the properties of the electrodes and there is also a small amount of noise. Over some ranges of the coupling parameter, dynamical clustering was observed. The precision-dependent cluster configuration is analyzed using hierarchical cluster trees. The cluster configurations varied with time: spontaneous changes of number of clusters and their configurations were detected. Simple transitions occurred with the switch of a single element or groups of elements. During more complicated transitions subclusters were exchanged among clusters but original cluster configurations were revisited. At weaker coupling the system itinerated among lower-dimensional quasistationary chaotic two-cluster states and higher-dimensional states with many clusters. In this region the transitions showed characteristics of on-off intermittency.

Cluster Analysis↗

A new strategy of cooperativity of biclustering and hierarchical clustering: a case of analyzing yeast genomic microarray datasets.

Hierarchical clustering is difficult to be deployed effectively in finding meaningful subtrees since genes rarely exhibit similar expression pattern across a wide range of conditions. It is also difficult to find a suitable level in cleaving a big hierarchy tree. Biclustering is a promising methodology in the field of the analysis of gene expression data of genechip. Generally it can be employed in identification of gene groups, which show a coherent expression profile across a subset of conditions. But in some cases of biclustering analysis of gene expressions, the genes in one bicluster are involved in more than one functional group, or all genes in one bicluster are involved in unknown functional groups (e.g. pattern VI and VIII in our studies). Then, how to predict the function of genes in these patterns? In the present research, we developed a new strategy of combining both of the clustering methods, hierarchical clustering and biclustering. The reserved conditions in datasets for hierarchical clustering were elicited according to the conditions in biclusters, and after hierarchical clustering, more detailed results in predicting unknown genes in certain patterns were obtained. This strategy of cooperating both of the methods during clustering procedure should be an effective guideline for functional predictions.

Cluster Analysis↗

Serum 1H-nuclear magnetic spectroscopy followed by principal component analysis and hierarchical cluster analysis to demonstrate effects of statins on hyperlipidemic patients.

Use of statins for prevention of coronary heart disease is based on the decrease of serum cholesterol and LDL cholesterol. To better investigate the changes in lipid profile after statin treatment, we propose here to use an analysis of serum by proton nuclear magnetic resonance (NMR) spectroscopy associated with a multivariate analysis of the main spectral components. Sera were obtained from 60 male patients treated for 6 weeks with simvastatin (30 patients) or atorvastatin (30 patients) for who LDL cholesterol decreased by over 45% in all selected patients. Proton nuclear magnetic resonance spectra were obtained and the region of methyl resonance from lipids was separated into six consecutive lines attributed to lipids which were analyzed by principal component analysis (PCA) and clustering by hierarchical cluster analysis (HCA) based on Euclidian distance coupled with the Ward's minimum variance method. PCA and HCA gave a map discriminating the 120 samples into five clusters, three clusters containing samples obtained at baseline and two others containing samples obtained after treatment. Both statins produced a decrease in lower-density lipoprotein components and an increase in higher density lipoprotein components. Patients with a coronary heart disease history could be discriminated after treatment by the increase in the component containing the highest proportion of HDL. Proton NMR spectroscopy of sera coupled with a PCA and an HCA was able to detect variations in the metabolism of lipids resulting from statin treatments.

Aged↗

Efficient determination of cluster boundaries for analysis of gene expression profile data using hierarchical clustering and wavelet transform.

The existing methods for clustering of gene expression profile data either require manual inspection and other biological knowledge or require some cut-off value which can not be directly calculated from the given data set. Thus, the problem of systematic and efficient determination of cluster boundaries of clusters in gene expression profile data still remains demanding. In this context, we have developed a procedure for automatic and systematic determination of the boundaries of clusters in the hierarchical clustering of gene expression data based on the ratio of with-in class variance and between-class variance, which can be fully calculated from the given expression data. After the determination of dendrogram based on agglomerative hierarchical clustering, this ratio is used to determine the cluster boundary. Except this ratio which can be completely calculated from the given expression profile data, unlike other existing approaches, our approach does not require any manual inspection or biological knowledge. Our results are favorably comparable and in some of cases better than existing method which does not utilize prior information or manual inspection. Moreover, gene expression profile data are often contaminated with various type of noise and in order to reduce this noise content, we have also applied image enhancing technique called discrete wavelet transform. We tested a number of mother wavelet functions to smooth the noise in the gene expression data set and obtained some improvements in the quality of the results.

Algorithms↗

Fast optimal leaf ordering for hierarchical clustering.

We present the first practical algorithm for the optimal linear leaf ordering of trees that are generated by hierarchical clustering. Hierarchical clustering has been extensively used to analyze gene expression data, and we show how optimal leaf ordering can reveal biological structure that is not observed with an existing heuristic ordering method. For a tree with n leaves, there are 2(n-1) linear orderings consistent with the structure of the tree. Our optimal leaf ordering algorithm runs in time O(n(4)), and we present further improvements that make the running time of our algorithm practical.

Algorithms↗

Expression profiling of 68 glycosyltransferase genes in 27 different human tissues by the systematic multiplex reverse transcription-polymerase chain reaction method revealed clustering of sexually related tissues in hierarchical clustering algorithm analysis.

We have developed an experimental system to study the expression of 68 human glycosyltransferase genes. Using this system, we examined the expression of those genes in 27 different tissues by the technique which we named systematic multiplex reverse transcription-polymerase chain reaction (SM RT-PCR). The panoramic view of a total of 1836 (68 x 27) expression data demonstrates that some glycosyltransferase genes are differentially expressed whereas some others are ubiquitously expressed. The data gathered provide more information on glycosyltransferase gene expression in tissues than any other paper published, and surpass in quantity all the information combined from previous publications. Although the expression profiling of glycosyltransferase genes alone may not directly explain the repertoires of oligosaccharides synthesized, it is an important step toward a better understanding of the gene expression network involved in oligosaccharide synthesis/degradation. Our modestly high-throughput gene expression study and the data analysis using a hierarchical clustering algorithm have allowed us to investigate the correlation between tissues and glycosyltransferase gene expression. Similar patterns of glycosyltransferase gene expression were observed in functionally and anatomically related tissues. All, but one, sexually related tissues formed a cluster in a tissue dendrogram, suggesting the involvement of sex hormones in the transcriptional control of many glycosyltransferase genes. Once established, the SM RT-PCR is cost- and time-efficient and requires small amounts of RNA as template. It is especially useful for the simultaneous analyses of multiple samples. Because of its simple design, the SM RT-PCR may offer an easy alternative in studying the expression of many other families of genes, as well as groups of related/unrelated genes, in various biological phenomena.

Algorithms↗

A hierarchical clustering algorithm for MIMD architecture.

Hierarchical clustering is the most often used method for grouping similar patterns of gene expression data. A fundamental problem with existing implementations of this clustering method is the inability to handle large data sets within a reasonable time and memory resources. We propose a parallelized algorithm of hierarchical clustering to solve this problem. Our implementation on a multiple instruction multiple data (MIMD) architecture shows considerable reduction in computational time and inter-node communication overhead, especially for large data sets. We use the standard message passing library, message passing interface (MPI) for any MIMD systems.

Algorithms↗

Influence of microarrays experiments missing values on the stability of gene groups by hierarchical clustering.

BACKGROUND: Microarray technologies produced large amount of data. The hierarchical clustering is commonly used to identify clusters of co-expressed genes. However, microarray datasets often contain missing values (MVs) representing a major drawback for the use of the clustering methods. Usually the MVs are not treated, or replaced by zero or estimated by the k-Nearest Neighbor (kNN) approach. The topic of the paper is to study the stability of gene clusters, defined by various hierarchical clustering algorithms, of microarrays experiments including or not MVs. RESULTS: In this study, we show that the MVs have important effects on the stability of the gene clusters. Moreover, the magnitude of the gene misallocations is depending on the aggregation algorithm. The most appropriate aggregation methods (e.g. complete-linkage and Ward) are highly sensitive to MVs, and surprisingly, for a very tiny proportion of MVs (e.g. 1%). In most of the case, the MVs must be replaced by expected values. The MVs replacement by the kNN approach clearly improves the identification of co-expressed gene clusters. Nevertheless, we observe that kNN approach is less suitable for the extreme values of gene expression. CONCLUSION: The presence of MVs (even at a low rate) is a major factor of gene cluster instability. In addition, the impact depends on the hierarchical clustering algorithm used. Some methods should be used carefully. Nevertheless, the kNN approach constitutes one efficient method for restoring the missing expression gene values, with a low error level. Our study highlights the need of statistical treatments in microarray data to avoid misinterpretation.

Cluster Analysis↗

Combination of automated high throughput platforms, flow cytometry, and hierarchical clustering to detect cell state.

BACKGROUND: This study examined whether hierarchical clustering could be used to detect cell states induced by treatment combinations that were generated through automation and high-throughput (HT) technology. Data-mining techniques were used to analyze the large experimental data sets to determine whether nonlinear, non-obvious responses could be extracted from the data. METHODS: Unary, binary, and ternary combinations of pharmacological factors (examples of stimuli) were used to induce differentiation of HL-60 cells using a HT automated approach. Cell profiles were analyzed by incorporating hierarchical clustering methods on data collected by flow cytometry. Data-mining techniques were used to explore the combinatorial space for nonlinear, unexpected events. Additional small-scale, follow-up experiments were performed on cellular profiles of interest. RESULTS: Multiple, distinct cellular profiles were detected using hierarchical clustering of expressed cell-surface antigens. Data-mining of this large, complex data set retrieved cases of both factor dominance and cooperativity, as well as atypical cellular profiles. Follow-up experiments found that treatment combinations producing "atypical cell types" made those cells more susceptible to apoptosis. CONCLUSIONS Hierarchical clustering and other data-mining techniques were applied to analyze large data sets from HT flow cytometry. From each sample, the data set was filtered and used to define discrete, usable states that were then related back to their original formulations. Analysis of resultant cell populations induced by a multitude of treatments identified unexpected phenotypes and nonlinear response profiles.

Algorithms↗

Unsupervised multistage image classification using hierarchical clustering with a Bayesian similarity measure.

A new multistage method using hierarchical clustering for unsupervised image classification is presented. In the first phase, the multistage method performs segmentation using a hierarchical clustering procedure which confines merging to spatially adjacent clusters and generates an image partition such that no union of any neighboring segments has homogeneous intensity values. In the second phase, the segments resulting from the first stage are classified into a small number of distinct states by a sequential merging operation. The region-merging procedure in the first phase makes use of spatial contextual information by characterizing the geophysical connectedness of a digital image structure with a Markov random field, while the second phase employs a context-free similarity measure in the clustering process. The segmentation procedure of region merging is implemented as a hierarchical clustering algorithm whereby a multiwindow approach using a pyramid-like structure is employed to increase computational efficiency while maintaining spatial connectivity in merging. From experiments with both simulated and remotely sensed data, the proposed method was determined to be quite effective for unsupervised analysis. In particular, the region-merging approach based on spatial contextual information was shown to provide more accurate classification of images with smooth spatial patterns.

Algorithms↗

A dynamically growing self-organizing tree (DGSOT) for hierarchical clustering gene expression profiles.

MOTIVATION: The increasing use of microarray technologies is generating large amounts of data that must be processed in order to extract useful and rational fundamental patterns of gene expression. Hierarchical clustering technology is one method used to analyze gene expression data, but traditional hierarchical clustering algorithms suffer from several drawbacks (e.g. fixed topology structure; mis-clustered data which cannot be reevaluated). In this paper, we introduce a new hierarchical clustering algorithm that overcomes some of these drawbacks. RESULT: We propose a new tree-structure self-organizing neural network, called dynamically growing self-organizing tree (DGSOT) algorithm for hierarchical clustering. The DGSOT constructs a hierarchy from top to bottom by division. At each hierarchical level, the DGSOT optimizes the number of clusters, from which the proper hierarchical structure of the underlying dataset can be found. In addition, we propose a new cluster validation criterion based on the geometric property of the Voronoi partition of the dataset in order to find the proper number of clusters at each hierarchical level. This criterion uses the Minimum Spanning Tree (MST) concept of graph theory and is computationally inexpensive for large datasets. A K-level up distribution (KLD) mechanism, which increases the scope of data distribution in the hierarchy construction, was used to improve the clustering accuracy. The KLD mechanism allows the data misclustered in the early stages to be reevaluated at a later stage and increases the accuracy of the final clustering result. The clustering result of the DGSOT is easily displayed as a dendrogram for visualization. Based on a yeast cell cycle microarray expression dataset, we found that our algorithm extracts gene expression patterns at different levels. Furthermore, the biological functionality enrichment in the clusters is considerably high and the hierarchical structure of the clusters is more reasonable. AVAILABILITY: DGSOT is available upon request from the authors.

Algorithms↗

Automatic synthesis of synergies for control of reaching--hierarchical clustering.

In this paper we describe a novel method for determining synergies between joint motions in reaching movements by hierarchical clustering. A set of recorded elbow and shoulder trajectories is used in a learning algorithm to determine the relationships between angular velocities at elbow and shoulder joints. The learning algorithm is based on optimal criteria for obtaining the hierarchy of descriptions of movement trajectories. We show that this method finds complex synergism between optimal joint trajectories for a given set of data and angular velocities at the shoulder and elbow joints. Three other machine learning techniques (ML) are used for comparison with our method of hierarchical clustering of trajectories. These MLs are: (1) radial basis functions (RBF), (2) inductive learning (IL), and (3) adaptive-network-based fuzzy inference system (ANFIS). Better error characteristics were obtained using the method of hierarchical clustering in comparison with the other techniques. The advantage of the method of hierarchical clustering with respect to the other MLs is in integrating the spatial and temporal elements of reaching movements. Determination and analysis of spatio-temporal events of movement trajectories is a useful tool in designing control systems for functional electrical stimulation (FES) assisted manipulation.

Algorithms↗

Hierarchical clustering analysis of tissue microarray immunostaining data identifies prognostically significant groups of breast carcinoma.

Prognostically relevant cluster groups, based on gene expression profiles, have been recently identified for breast cancers, lung cancers, and lymphoma. Our aim was to determine whether hierarchical clustering analysis of multiple immunomarkers (protein expression profiles) improves prognostication in patients with invasive breast cancer. A cohort of 438 sequential cases of invasive breast cancer with median follow-up of 15.4 years was selected for tissue microarray construction. A total of 31 biomarkers were tested by immunohistochemistry on these tissue arrays. The prognostic significance of individual markers was assessed by using Kaplan-Meier survival estimates and log-rank tests. Seventeen of 31 markers showed prognostic significance in univariate analysis (P < or = 0.05) and 4 markers showed a trend toward significance (P < or = 0.2). Unsupervised hierarchical clustering analysis was done by using these 21 immunomarkers, and this resulted in identification of three cluster groups with significant differences in clinical outcome. chi2 analysis showed that expression of 11 markers significantly correlated with membership in one of the three cluster groups. Unsupervised hierarchical clustering analysis with this set of 11 markers reproduced the same three prognostically significant cluster groups identified by using the larger set of markers. These cluster groups were of prognostic significance independent of lymph node metastasis, tumor size, and tumor grade in multivariate analysis (P=0.0001). The cluster groups were as powerful a prognostic indicator as lymph node status. This work demonstrates that hierarchical clustering of immunostaining data by using multiple markers can group breast cancers into classes with clinical relevance and is superior to the use of individual prognostic markers.

Adult↗

A memetic-aided approach to hierarchical clustering from distance matrices: application to gene expression clustering and phylogeny.

We propose a heuristic approach to hierarchical clustering from distance matrices based on the use of memetic algorithms (MAs). By using MAs to solve some variants of the Minimum Weight Hamiltonian Path Problem on the input matrix, a sequence of the individual elements to be clustered (referred to as patterns) is first obtained. While this problem is also NP-hard, a probably optimal sequence is easy to find with the current advances for this problem and helps to prune the space of possible solutions and/or to guide the search performed by an actual clustering algorithm. This technique has been successfully applied to both a Branch-and-Bound algorithm, and to evolutionary algorithms and MAs. Experimental results are given in the context of phylogenetic inference and in the hierarchical clustering of gene expression data.

Algorithms↗

Evaluation of immunohistochemical markers in non-small cell lung cancer by unsupervised hierarchical clustering analysis: a tissue microarray study of 284 cases and 18 markers.

This study has investigated a panel of immunomarkers in non-small cell lung carcinoma (NSCLC). Unsupervised hierarchical clustering analysis was used to investigate the possibility of identifying different subgroups in NSCLC based on their molecular expression profile rather than morphological features. A tissue microarray consisting of 284 cases of NSCLC was constructed. Immunohistochemistry was used to detect the presence of 18 biomarkers including synaptophysin, chromogranin, bombesin, NSE, GFI1, ASH-1, p53, p63, p21, p27, E2F-1, cyclin D1, Bcl-2, TTF-1, CEA, HER2/neu, cytokeratin 5/6, and pancytokeratin. Univariate analysis of all 18 markers for prognostic significance was performed. Immunohistochemical scoring data for NSCLC were analysed by unsupervised hierarchical clustering analysis. Kaplan-Meier survival curves were plotted for the different cluster groups of lung tumours identified by this method. Analysis of the three different World Health Organization (WHO) subtypes (adenocarcinoma, squamous cell carcinoma, large cell carcinoma) of NSCLC individually showed that different markers were significant in different subtypes. For example, p53 and p63 were significant for squamous cell carcinoma (p = 0.007 and p = 0.03, respectively), whereas cyclin D1 and HER2/neu were significant prognostic markers for adenocarcinoma (p = 0.025 and p = 0.015, respectively). These markers were not significant prognostic predictors for NSCLC as a group. Hierarchical clustering analysis of NSCLC produced four separate cluster groups, although the vast majority of cases were found in two cluster groups, one dominated by squamous cell carcinoma and the other by adenocarcinoma. The clinical outcomes of cases from the four cluster groups were not significantly different. Prognostic indicators vary between different morphological subtypes of NSCLC. Unsupervised hierarchical clustering analysis, based on an extended immunoprofile, identifies two main cluster groups corresponding to adenocarcinoma and squamous cell carcinoma; cases of large cell carcinomas are assigned to one of these two groups based on their molecular phenotype.

Adenocarcinoma↗

A hierarchical clustering approach for large compound libraries.

A modified version of the k-means clustering algorithm was developed that is able to analyze large compound libraries. A distance threshold determined by plotting the sum of radii of leaf clusters was used as a termination criterion for the clustering process. Hierarchical trees were constructed that can be used to obtain an overview of the data distribution and inherent cluster structure. The approach is also applicable to ligand-based virtual screening with the aim to generate preferred screening collections or focused compound libraries. Retrospective analysis of two activity classes was performed: inhibitors of caspase 1 [interleukin 1 (IL1) cleaving enzyme, ICE] and glucocorticoid receptor ligands. The MDL Drug Data Report (MDDR) and Collection of Bioactive Reference Analogues (COBRA) databases served as the compound pool, for which binary trees were produced. Molecules were encoded by all Molecular Operating Environment 2D descriptors and topological pharmacophore atom types. Individual clusters were assessed for their purity and enrichment of actives belonging to the two ligand classes. Significant enrichment was observed in individual branches of the cluster tree. After clustering a combined database of MDDR, COBRA, and the SPECS catalog, it was possible to retrieve MDDR ICE inhibitors with new scaffolds using COBRA ICE inhibitors as seeds. A Java implementation of the clustering method is available via the Internet (http://www.modlab.de).

Algorithms↗