PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Atlas-level single-cell integration and clustering-free differential expression analysis with GEDI 2.0.

MOTIVATION: GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS: We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98× mean single-threaded speedups, and up to 11.5× acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.

Single-Cell Analysis↗

Detecting clusters of different geometrical shapes in microarray gene expression data.

MOTIVATION: Clustering has been used as a popular technique for finding groups of genes that show similar expression patterns under multiple experimental conditions. Many clustering methods have been proposed for clustering gene-expression data, including the hierarchical clustering, k-means clustering and self-organizing map (SOM). However, the conventional methods are limited to identify different shapes of clusters because they use a fixed distance norm when calculating the distance between genes. The fixed distance norm imposes a fixed geometrical shape on the clusters regardless of the actual data distribution. Thus, different distance norms are required for handling the different shapes of clusters. RESULTS: We present the Gustafson-Kessel (GK) clustering method for microarray gene-expression data. To detect clusters of different shapes in a dataset, we use an adaptive distance norm that is calculated by a fuzzy covariance matrix (F) of each cluster in which the eigenstructure of F is used as an indicator of the shape of the cluster. Moreover, the GK method is less prone to falling into local minima than the k-means and SOM because it makes decisions through the use of membership degrees of a gene to clusters. The algorithmic procedure is accomplished by the alternating optimization technique, which iteratively improves a sequence of sets of clusters until no further improvement is possible. To test the performance of the GK method, we applied the GK method and well-known conventional methods to three recently published yeast datasets, and compared the performance of each method using the Saccharomyces Genome Database annotations. The clustering results of the GK method are more significantly relevant to the biological annotations than those of the other methods, demonstrating its effectiveness and potential for clustering gene-expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at http://dragon.kaist.ac.kr/gk.

Algorithms↗

Applying watershed algorithms to the segmentation of clustered nuclei.

Cluster division is a critical issue in fluorescence microscopy-based analytical cytology when preparation protocols do not provide appropriate separation of objects. Overlooking clustered nuclei and analyzing only isolated nuclei may dramatically increase analysis time or affect the statistical validation of the results. Automatic segmentation of clustered nuclei requires the implementation of specific image segmentation tools. Most algorithms are inspired by one of the two following strategies: 1) cluster division by the detection of internuclei gradients; or 2) division by definition of domains of influence (geometrical approach). Both strategies lead to completely different implementations, and usually algorithms based on a single view strategy fail to correctly segment most clustered nuclei, or perform well just for a specific type of sample. An algorithm based on morphological watersheds has been implemented and tested on the segmentation of microscopic nuclei clusters. This algorithm provides a tool that can be used for the implementation of both gradient- and domain-based algorithms, and, more importantly, for the implementation of mixed (gradient- and shape-based) algorithms. Using this algorithm, almost 90% of the test clusters were correctly segmented in peripheral blood and bone marrow preparations. The algorithm was valid for both types of samples, using the appropriate markers and transformations.

Algorithms↗

Cluster hybrid Monte Carlo simulation algorithms.

We show that addition of Metropolis single spin flips to the Wolff cluster-flipping Monte Carlo procedure leads to a dramatic increase in performance for the spin-1/2 Ising model. We also show that adding Wolff cluster flipping to the Metropolis or heat bath algorithms in systems where just cluster flipping is not immediately obvious (such as the spin-3/2 Ising model) can substantially reduce the statistical errors of the simulations. A further advantage of these methods is that systematic errors introduced by the use of imperfect random-number generation may be largely healed by hybridizing single spin flips with cluster flipping.

Journal Article↗

The fuzzy clustering analysis based on AFS theory.

In the framework of axiomatic fuzzy sets theory, we first study how to impersonally and automatically determine the membership functions for fuzzy sets according to original data and facts, and a new algorithmic framework of determining membership functions and their logic operations for fuzzy sets has been proposed. Then, we apply the proposed algorithmic framework to give a new clustering algorithm and show that the algorithm is feasible. A number of illustrative examples show that this approach offers a far more flexible and effective means for the intelligent systems in real-world applications. Compared with popular fuzzy clustering algorithms, such as c-means fuzzy algorithm and k-nearest-neighbor fuzzy algorithm, the new fuzzy clustering algorithm is more simple and understandable, the data types of the attributes can be various data types or subpreference relations, even descriptions of human intuition, and the distance function and the class number need not be given beforehand.

Algorithms↗

A scan statistic with a variable window.

Given N points or events occurring according to some probability distribution in the unit interval (0, 1), the simple scan statistic is defined to be the maximum number of points in any sub-interval of length d. In many areas, as in epidemiology, it is used to test the null hypothesis that the events are random, against the alternative that they cluster within some window of fixed width d. Since d must be chosen without snooping at the data, the test restricts the alternative to clusters of a fixed size. In this paper, we propose a scan statistic with a variable window, whose size does not need to be chosen a priori. This test is the generalized likelihood ratio test for a uniform null distribution against an alternative of non-random clustering which allows for clusters of variable width. A simple algorithm for the implementation of the method is given and applied to birth defects data previously analysed by a simple scan statistic.

Algorithms↗

Polychromatic flow cytometry: a rapid method for the reduction and analysis of complex multiparameter data.

BACKGROUND: Recent advances in flow cytometry have resulted in the development of reliable techniques for performing polychromatic (5-17 color) flow cytometry analysis. However, the data reduction and analysis involved in the resolution of hundreds of possible cellular subphenotypes identified, using a single polychromatic flow cytometry staining panel, presents a major obstacle to the successful application of this technology. METHODS: To generate two distinct collections of T cell populations with differentially expressed surface markers, cryopreserved lymph node cells from 5 melanoma patients vaccinated with the modified gp100(209-2M) melanoma peptide were stimulated with cognate peptide and cultured in either IL-21 + low-dose IL-2 or IL-15 + low-dose IL-2. In vitro stimulated (IVS) cells were interrogated using 8-color flow cytometry. Data were analyzed using Winlist Hyperlog and FCOM software, and 32 T cell subsets were resolved for each culture condition. Hierarchical clustering analysis was applied to the relative percentages of each subphenotype for both IVS conditions to determine if unique cell surface marker expression signatures were produced for each IVS culture. RESULTS: Sequential data analysis using Hyperlog and FCOM demonstrated that lymphocytes cultured in IL-21 + IL-2 had a distinctively different set of subphenotype signatures compared to cells grown in IL-15 + IL-2 for all 5 patients. Importantly, subsequent cluster analysis of all 32 subphenotype frequencies in each IVS test condition for all 5 patients reproducibly demonstrated that cellular subphenotypes produced after IL-21 + IL-2 IVS partitioned separately from subphenotypes produced by IL-15 + IL-2 IVS. CONCLUSIONS: The integrated sequential use of Hyperlog and FCOM software with cluster analysis algorithms for the reduction and analysis of polychromatic flow cytometry data produces an effective, rapid technique for the assessment of complex patterns of subphenotype expression between and within multiple test samples. This approach to data analysis may enhance the use of polychromatic flow cytometry for both research and clinical applications.

Algorithms↗

Automatic identification of significant graphoelements in multichannel EEG recordings by adaptive segmentation and fuzzy clustering.

A new approach to visual evaluation of long-term EEG recordings is proposed. The method is based on multichannel adaptive segmentation, subsequent feature extraction, automatic classification of the acquired segments by fuzzy cluster analysis (fuzzy c-means algorithm), and on the distinguishing of thus identified EEG segments by colour directly in the EEG record. The black and white variant of the described automatic system is presented. The method was evaluated by applying it to simulated artificial data and to real EEG recordings; some of the illustrative results are shown. In addition, the performance of this system is evaluated and the first experience with its application to routine EEG recordings is discussed.

Algorithms↗

Rotational diffusivity of fractal clusters.

The rotational diffusion behavior of fractal clusters generated through an off-lattice cluster-cluster aggregation algorithm in both diffusion-limited cluster aggregation and reaction-limited cluster aggregation conditions is investigated. The extended Kirkwood-Riseman theory (Garcia de la Torre et al., Macromolecules, 1987) is used to estimate the cluster rotational diffusion tensor. The three eigenvalues of this tensor, which correspond to the three main rotational diffusivity values of the cluster, have been computed for each generated cluster. Once the eigenvalues have been sorted in ascending order, each of them has been averaged over several thousands of clusters. It is found that one of the three main average rotational diffusivities is substantially larger than the other two, indicating significant anisotropy of fractal clusters. Moreover, a rotational hydrodynamic radius Rh,r has been determined on the basis of the mean value of the three average rotational diffusivities, which is about 25% larger than the mean translational hydrodynamic radius Rh calculated through the same Kirkwood-Riseman theory. Finally, the obtained Rh,r values have been applied to interpret dynamic light scattering data from aggregating colloidal systems and to investigate the reliability of the assumption, Rh = Rh,r, typically made in the literature.

Journal Article↗

SEAN: SNP prediction and display program utilizing EST sequence clusters.

SEAN is an application that predicts single nucleotide polymorphisms (SNPs) using multiple sequence alignments produced from expressed sequence tag (EST) clusters. The algorithm uses rules of sequence identity and SNP abundance to determine the quality of the prediction. A Java viewer is provided to display the EST alignments and predicted SNPs.

Algorithms↗

Incorporating gene functions as priors in model-based clustering of microarray gene expression data.

MOTIVATION: Cluster analysis of gene expression profiles has been widely applied to clustering genes for gene function discovery. Many approaches have been proposed. The rationale is that the genes with the same biological function or involved in the same biological process are more likely to co-express, hence they are more likely to form a cluster with similar gene expression patterns. However, most existing methods, including model-based clustering, ignore known gene functions in clustering. RESULTS: To take advantage of accumulating gene functional annotations, we propose incorporating known gene functions as prior probabilities in model-based clustering. In contrast to a global mixture model applicable to all the genes in the standard model-based clustering, we use a stratified mixture model: one stratum corresponds to the genes of unknown function while each of the other ones corresponding to the genes sharing the same biological function or pathway; the genes from the same stratum are assumed to have the same prior probability of coming from a cluster while those from different strata are allowed to have different prior probabilities of coming from the same cluster. We derive a simple EM algorithm that can be used to fit the stratified model. A simulation study and an application to gene function prediction demonstrate the advantage of our proposal over the standard method. CONTACT: weip@biostat.umn.edu

Algorithms↗

Cluster analysis and related techniques in medical research.

In this paper we review methods of cluster analysis in the context of classifying patients on the basis of clinical and/or laboratory type observations. Both hierarchical and non-hierarchical methods of clustering are considered, although the emphasis is on the latter type, with particular attention devoted to the mixture likelihood-based approach. For the purposes of dividing a given data set into g clusters, this approach fits a mixture model of g components, using the method of maximum likelihood. It thus provides a sound statistical basis for clustering. The important but difficult question of how many clusters are there in the data can be addressed within the framework of standard statistical theory, although theoretical and computational difficulties still remain. Two case studies, involving the cluster analysis of some haemophilia and diabetes data respectively, are reported to demonstrate the mixture likelihood-based approach to clustering.

Algorithms↗

A software tool for creating simulated outbreaks to benchmark surveillance systems.

BACKGROUND: Evaluating surveillance systems for the early detection of bioterrorism is particularly challenging when systems are designed to detect events for which there are few or no historical examples. One approach to benchmarking outbreak detection performance is to create semi-synthetic datasets containing authentic baseline patient data (noise) and injected artificial patient clusters, as signal. METHODS: We describe a software tool, the AEGIS Cluster Creation Tool (AEGIS-CCT), that enables users to create simulated clusters with controlled feature sets, varying the desired cluster radius, density, distance, relative location from a reference point, and temporal epidemiological growth pattern. AEGIS-CCT does not require the use of an external geographical information system program for cluster creation. The cluster creation tool is an open source program, implemented in Java and is freely available under the Lesser GNU Public License at its Sourceforge website. Cluster data are written to files or can be appended to existing files so that the resulting file will include both existing baseline and artificially added cases. Multiple cluster file creation is an automated process in which multiple cluster files are created by varying a single parameter within a user-specified range. To evaluate the output of this software tool, sets of test clusters were created and graphically rendered. RESULTS: Based on user-specified parameters describing the location, properties, and temporal pattern of simulated clusters, AEGIS-CCT created clusters accurately and uniformly. CONCLUSION: AEGIS-CCT enables the ready creation of datasets for benchmarking outbreak detection systems. It may be useful for automating the testing and validation of spatial and temporal cluster detection algorithms.

Algorithms↗

Exploring cross-category relationships between symptoms in people with hypermobile EDS (hEDS) to identify disability patterns.

BACKGROUND: Hypermobile Ehlers-Danlos Syndrome (hEDS) is a connective tissue disorder with variable symptom presentation across multiple organ systems and significant morbidity. Little is known about hEDS etiology and identifying patterns of symptom co-occurrence can reveal previously unidentified relationships between phenotypes and inform studies of underlying disease pathophysiology for symptoms that may share functional biological pathways. In this exploratory analysis, we specifically assessed the distribution of symptoms in case and controls to identify clusters of co-occurring symptoms. METHODS: We have interrogated clinically relevant symptom areas in 47 females with hEDS, 36 age-matched female controls and 8 hypermobile patients without chronic pain. Studied symptoms include general health, mental health, body pain, vitality and energy, autonomic symptoms, bleeding, and gastrointestinal symptoms. We conducted hierarchal clustering on principle components (HCPC) to identify groups and compared the groups for the previously described symptoms. Radial plots were used to identify relationships between severe symptom categories. RESULTS: Our analysis reveals statistically significantly more severe symptoms in all categories in people with hEDS compared with age- and sex-matched controls and asymptomatic hypermobile patients. HCPC identified clearly separated Low, Moderate, and High symptom groups within participants. The Low dysfunction groups include nearly all controls and hypermobile patients without chronic pain. The High dysfunction group includes ~60% of people with hEDS, while around 40% are in the Moderate dysfunction cluster. Cluster solutions for all participants were stable with moderate fit (silhouette 0.64; Jaccard boot mean 0.91). Group level radial plots showed high bleeding severity across all symptom clusters, while disproportional severity of general health, physical function, limitation of role due to physical symptoms, pain, and social functioning deficits differentiates the High from Moderate and Low Dysfunction clusters. CONCLUSION: Using this analysis at the group level has revealed patterns suggesting a progression of disease symptoms. People with hypermobility do not uniformly have severe symptoms but instead have some symptoms that differentiate from non-hypermobile individuals. While exploratory, using a radar multi-symptom analysis may be used to evaluate disproportionately severe symptoms contributing to the patterns of global symptom severity. These include pain but also ability to perform roles, suggesting strong utility of physical and occupational therapies to emphasize coping. This may also allow better targeting of etiological studies and may have additional utility at an individual level to develop symptom management strategies.

Humans↗

Analysis of gene expression profiles: an application of memetic algorithms to the minimum sum-of-squares clustering problem.

Microarrays have become a key technology in experimental molecular biology since they allow monitoring of gene expression for more than 10,000 genes in parallel producing huge amounts of data. In the exploration of transcriptional regulatory networks, an important task is to cluster gene expression data to identify groups of genes with similar patterns and hence similar function. In this paper, memetic algorithms (MAs)-evolutionary algorithms incorporating local search-are proposed for minimum sum-of-squares clustering (MSSC). In a fitness landscape analysis, it is shown that the MSSC problem has correlation structure exploitable by MAs. The proposed MAs are shown to be superior to multi-start k-means as well as five other clustering algorithms from the bioinformatics literature including hierarchical algorithms and self-organizing maps. Although the fitness values of the different clustering solutions lie close together, it is shown that the solutions differ significantly from each other in terms of cluster memberships which is extremely important for the biological interpretation of the clustering results.

Algorithms↗

Mixture modelling of gene expression data from microarray experiments.

MOTIVATION: Hierarchical clustering is one of the major analytical tools for gene expression data from microarray experiments. A major problem in the interpretation of the output from these procedures is assessing the reliability of the clustering results. We address this issue by developing a mixture model-based approach for the analysis of microarray data. Within this framework, we present novel algorithms for clustering genes and samples. One of the byproducts of our method is a probabilistic measure for the number of true clusters in the data. RESULTS: The proposed methods are illustrated by application to microarray datasets from two cancer studies; one in which malignant melanoma is profiled (Bittner et al., Nature, 406, 536-540, 2000), and the other in which prostate cancer is profiled (Dhanasekaran et al., 2001, submitted).

Algorithms↗

A novel functional module detection algorithm for protein-protein interaction networks.

BACKGROUND: The sparse connectivity of protein-protein interaction data sets makes identification of functional modules challenging. The purpose of this study is to critically evaluate a novel clustering technique for clustering and detecting functional modules in protein-protein interaction networks, termed STM. RESULTS: STM selects representative proteins for each cluster and iteratively refines clusters based on a combination of the signal transduced and graph topology. STM is found to be effective at detecting clusters with a diverse range of interaction structures that are significant on measures of biological relevance. The STM approach is compared to six competing approaches including the maximum clique, quasi-clique, minimum cut, betweeness cut and Markov Clustering (MCL) algorithms. The clusters obtained by each technique are compared for enrichment of biological function. STM generates larger clusters and the clusters identified have p-values that are approximately 125-fold better than the other methods on biological function. An important strength of STM is that the percentage of proteins that are discarded to create clusters is much lower than the other approaches. CONCLUSION: STM outperforms competing approaches and is capable of effectively detecting both densely and sparsely connected, biologically relevant functional modules with fewer discards.

Journal Article↗

Structure discovery in medical databases: a conceptual clustering approach.

Clustering is an important data analysis tool for discovering structure in data sets. Although research on conceptual clustering has produced algorithms showing significant advantages over earlier numerical ones, existing methods still present some limitations regarding applicability to biomedical domains. In this paper we describe ADAGIO, a conceptual clustering algorithm combining a low-cost preordering process with a breadth-first incremental control strategy that incorporates merging and splitting operators. Experimental evaluation indicated that the algorithm achieves a good balance between structure discovery performance and computational efficiency, and demonstrated the comparative effectiveness of its missing information handling process. ADAGIO is able to handle qualitative, quantitative and mixed-type data. An application example to a cancer domain is given, where the algorithm was able to suggest interesting epidemiological interpretations.

Algorithms↗