PubMed · 10876334
Measuring cluster similarity across methods.
Abstract
Cluster analysis techniques delineate groupings or categories of observations based on some shared commonality over a set of variables. If such groupings can be formed, their commonality may be investigated to define relationships that may otherwise go undetected given their complexity. However, the cluster analyses are inappropriate unless the results can be replicated. A number of clustering techniques are available, differing mostly in the technical criteria used to judge the similarity of the observations. There is added validity to the cluster structure when different methods produce similar groupings; however, in most cases, different clustering techniques will not produce identical clusters and the extent of cluster similarity becomes an important measure. In this paper the hypergeometric distribution is used to gauge cluster similarity across different methods, providing an appropriate measure of consistency. This measure is used to validate reproducibility of the clusters.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
A J Kos, C Psenicka. 2000. Measuring cluster similarity across methods.. https://doi.org/10.2466/pr0.2000.86.3.858
Cite the original work for its findings. Save a collection to share your selection of sources.