PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Effect of data normalization on fuzzy clustering of DNA microarray data.

BACKGROUND: Microarray technology has made it possible to simultaneously measure the expression levels of large numbers of genes in a short time. Gene expression data is information rich; however, extensive data mining is required to identify the patterns that characterize the underlying mechanisms of action. Clustering is an important tool for finding groups of genes with similar expression patterns in microarray data analysis. However, hard clustering methods, which assign each gene exactly to one cluster, are poorly suited to the analysis of microarray datasets because in such datasets the clusters of genes frequently overlap. RESULTS: In this study we applied the fuzzy partitional clustering method known as Fuzzy C-Means (FCM) to overcome the limitations of hard clustering. To identify the effect of data normalization, we used three normalization methods, the two common scale and location transformations and Lowess normalization methods, to normalize three microarray datasets and three simulated datasets. First we determined the optimal parameters for FCM clustering. We found that the optimal fuzzification parameter in the FCM analysis of a microarray dataset depended on the normalization method applied to the dataset during preprocessing. We additionally evaluated the effect of normalization of noisy datasets on the results obtained when hard clustering or FCM clustering was applied to those datasets. The effects of normalization were evaluated using both simulated datasets and microarray datasets. A comparative analysis showed that the clustering results depended on the normalization method used and the noisiness of the data. In particular, the selection of the fuzzification parameter value for the FCM method was sensitive to the normalization method used for datasets with large variations across samples. CONCLUSION: Lowess normalization is more robust for clustering of genes from general microarray data than the two common scale and location adjustment methods when samples have varying expression patterns or are noisy. In particular, the FCM method slightly outperformed the hard clustering methods when the expression patterns of genes overlapped and was advantageous in finding co-regulated genes. Thus, the FCM approach offers a convenient method for finding subsets of genes that are strongly associated to a given cluster.

Algorithms↗

Context-specific infinite mixtures for clustering gene expression profiles across diverse microarray dataset.

MOTIVATION: Identifying groups of co-regulated genes by monitoring their expression over various experimental conditions is complicated by the fact that such co-regulation is condition-specific. Ignoring the context-specific nature of co-regulation significantly reduces the ability of clustering procedures to detect co-expressed genes due to additional 'noise' introduced by non-informative measurements. RESULTS: We have developed a novel Bayesian hierarchical model and corresponding computational algorithms for clustering gene expression profiles across diverse experimental conditions and studies that accounts for context-specificity of gene expression patterns. The model is based on the Bayesian infinite mixtures framework and does not require a priori specification of the number of clusters. We demonstrate that explicit modeling of context-specificity results in increased accuracy of the cluster analysis by examining the specificity and sensitivity of clusters in microarray data. We also demonstrate that probabilities of co-expression derived from the posterior distribution of clusterings are valid estimates of statistical significance of created clusters. AVAILABILITY: The open-source package gimm is available at http://eh3.uc.edu/gimm.

Algorithms↗

Classification and subtype prediction of adult soft tissue sarcoma by functional genomics.

Adult soft tissue sarcomas are a heterogeneous group of tumors, including well-described subtypes by histological and genotypic criteria, and pleomorphic tumors typically characterized by non-recurrent genetic aberrations and karyotypic heterogeneity. The latter pose a diagnostic challenge, even to experienced pathologists. We proposed that gene expression profiling in soft tissue sarcoma would identify a genomic-based classification scheme that is useful in diagnosis. RNA samples from 51 pathologically confirmed cases, representing nine different histological subtypes of adult soft tissue sarcoma, were examined using the Affymetrix U95A GeneChip. Statistical tests were performed on experimental groups identified by cluster analysis, to find discriminating genes that could subsequently be applied in a support vector machine algorithm. Synovial sarcomas, round-cell/myxoid liposarcomas, clear-cell sarcomas and gastrointestinal stromal tumors displayed remarkably distinct and homogenous gene expression profiles. Pleomorphic tumors were heterogeneous. Notably, a subset of malignant fibrous histiocytomas, a controversialhistological subtype, was identified as a distinct genomic group. The support vector machine algorithm supported a genomic basis for diagnosis, with both high sensitivity and specificity. In conclusion, we showed gene expression profiling to be useful in classification and diagnosis, providing insights into pathogenesis and pointing to potential new therapeutic targets of soft tissue sarcoma.

Adult↗

Mining gene expression data for positive and negative co-regulated gene clusters.

MOTIVATION: Analysis of gene expression data can provide insights into the positive and negative co-regulation of genes. However, existing methods such as association rule mining are computationally expensive and the quality and quantities of the rules are sensitive to the support and confidence values. In this paper, we introduce the concept of positive and negative co-regulated gene cluster (PNCGC) that more accurately reflects the co-regulation of genes, and propose an efficient algorithm to extract PNCGCs. RESULTS: We experimented with the Yeast dataset and compared our resulting PNCGCs with the association rules generated by the Apriori mining algorithm. Our results show that our PNCGCs identify some missing co-regulations of association rules, and our algorithm greatly reduces the large number of rules involving uncorrelated genes generated by the Apriori scheme. AVAILABILITY: The software is available upon request.

Algorithms↗

Oxidative stress response of tumor cells: microarray-based comparison between artemisinins and anthracyclines.

The antimalarial artemisinins also reveal profound cytotoxic activity against tumor cells. Artemisinins harbor an endoperoxide bridge whose cleavage results in the generation of reactive oxygen species (ROS) and/or artemisinin carbon-centered free radicals. Established cancer drugs such as anthracyclines also form ROS and free radicals that are responsible for the cardiotoxicity of anthracyclines. In contrast, artemisinins do not reveal cardiotoxicity. In the present investigation, we compared the cytotoxic activities of different artemisinins (artemisinin, artesunate, arteether, artemether, artemisitene, dihydroartemisinylester stereoisomers) in 60 cell lines of the National Cancer Institute (NCI), USA, with those of anthracyclines (doxorubicin, daunorubicin, 4'-epirubicin, idarubicin, deoxydoxorubicin, trifluoroacetyl-doxorubicin-14-valerate). The inhibition concentration 50% (IC(50)) values of artemisinins and anthracyclines were correlated with the mRNA expression of 170 genes involved in oxygen stress response and metabolism as recently determined by microarray analysis and deposited in the NCI database (http://dtp.nci.nih.gov). The genes whose expression was significantly linked to cellular drug response in Kendall's tau tests were subjected to hierarchical cluster analysis and cluster image mapping. Mathematical correction for false-positive correlations was done by a false discovery rate algorithm. One cluster contained predominantly genes with a relationship to artemisinins and another one genes with a relationship to anthracyclines. In a third cluster, genes correlating to both drug classes were assembled. This indicates that different sets of genes involved in oxidative stress response and metabolism may contribute to the cytotoxic and differing toxic side effects of these drug classes.

Animals↗

Effective Memetic Algorithms for VLSI design = Genetic Algorithms + local search + multi-level clustering.

Combining global and local search is a strategy used by many successful hybrid optimization approaches. Memetic Algorithms (MAs) are Evolutionary Algorithms (EAs) that apply some sort of local search to further improve the fitness of individuals in the population. Memetic Algorithms have been shown to be very effective in solving many hard combinatorial optimization problems. This paper provides a forum for identifying and exploring the key issues that affect the design and application of Memetic Algorithms. The approach combines a hierarchical design technique, Genetic Algorithms, constructive techniques and advanced local search to solve VLSI circuit layout in the form of circuit partitioning and placement. Results obtained indicate that Memetic Algorithms based on local search, clustering and good initial solutions improve solution quality on average by 35% for the VLSI circuit partitioning problem and 54% for the VLSI standard cell placement problem.

Algorithms↗

Likelihood inference for exchangeable binary data with varying cluster sizes.

This article investigates maximum likelihood estimation with saturated and unsaturated models for correlated exchangeable binary data, when a sample of independent clusters of varying sizes is available. We discuss various parameterizations of these models, and propose using the EM algorithm to obtain maximum likelihood estimates. The methodology is illustrated by applications to a study of familial disease aggregation and to the design of a proposed group randomized cancer prevention trial.

Algorithms↗

A simulation study of odds ratio estimation for binary outcomes from cluster randomized trials.

We used simulation to compare accuracy of estimation and confidence interval coverage of several methods for analysing binary outcomes from cluster randomized trials. The following methods were used to estimate the population-averaged intervention effect on the log-odds scale: marginal logistic regression models using generalized estimating equations with information sandwich estimates of standard error (GEE); unweighted cluster-level mean difference (CL/U); weighted cluster-level mean difference (CL/W) and cluster-level random effects linear regression (CL/RE). Methods were compared across trials simulated with different numbers of clusters per trial arm, numbers of subjects per cluster, intraclass correlation coefficients (rho), and intervention versus control arm proportions. Two thousand data sets were generated for each combination of design parameter values. The results showed that the GEE method has generally acceptable properties, including close to nominal levels of confidence interval coverage, when a simple adjustment is made for data with relatively few clusters. CL/U and CL/W have good properties for trials where the number of subjects per cluster is sufficiently large and rho is sufficiently small. CL/RE also has good properties in this situation provided a t-distribution multiplier is used for confidence interval calculation in studies with small numbers of clusters. For studies where the number of subjects per cluster is small and rho is large all cluster-level methods may perform poorly for studies with between 10 and 50 clusters per trial arm.

Algorithms↗

Comparison of chemical clustering methods using graph- and fingerprint-based similarity measures.

This paper compares several published methods for clustering chemical structures, using both graph- and fingerprint-based similarity measures. The clusterings from each method were compared to determine the degree of cluster overlap. Each method was also evaluated on how well it grouped structures into clusters possessing a non-trivial substructural commonality. The methods which employ adjustable parameters were tested to determine the stability of each parameter for datasets of varying size and composition. Our experiments suggest that both graph- and fingerprint-based similarity measures can be used effectively for generating chemical clusterings; it is also suggested that the CAST and Yin-Chen methods, suggested recently for the clustering of gene expression patterns, may also prove effective for the clustering of 2D chemical structures.

Algorithms↗

Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.

Spatial omics technologies have revolutionized the study of tissue architecture and cellular heterogeneity by integrating molecular profiles with spatial localization. In spatially resolved transcriptomics, delineating higher-order anatomical structures is critical for understanding how cellular organization affects function. However, the reliability of current benchmarks of spatially aware clustering (SAC) methods is undermined by their narrow focus on Visium and brain tissue datasets and the incorrect interpretation of manual annotation as ground truth. Here we present SACCELERATOR, a community-driven, extensible framework that standardizes data formatting, method integration and metric evaluation, enabling rapid inclusion of new methods and datasets. Our analysis revealed substantial limitations in the generalizability and reproducibility of SAC methods and shows that anatomical labels commonly used as ground truths are often biased, error prone and unsuitable for benchmarking. Rather than ranking methods, we propose a consensus-guided workflow where descriptive spatial metrics highlight high-entropy regions of method disagreement, enabling targeted feedback for tissue experts. Applied to brain and cancer datasets, this approach uncovered biologically meaningful patterns overlooked by individual SAC methods and manual annotations, highlighting the need for iterative, expert-in-the-loop evaluation.

Benchmarking↗

Shifting and scaling patterns from gene expression data.

MOTIVATION: During the last years, the discovering of biclusters in data is becoming more and more popular. Biclustering aims at extracting a set of clusters, each of which might use a different subset of attributes. Therefore, it is clear that the usefulness of biclustering techniques is beyond the traditional clustering techniques, especially when datasets present high or very high dimensionality. Also, biclustering considers overlapping, which is an interesting aspect, algorithmically and from the point of view of the result interpretation. Since the Cheng and Church's works, the mean squared residue has turned into one of the most popular measures to search for biclusters, which ideally should discover shifting and scaling patterns. RESULTS: In this work, we identify both types of patterns (shifting and scaling) and demonstrate that the mean squared residue is very useful to search for shifting patterns, but it is not appropriate to find scaling patterns because even when we find a perfect scaling pattern the mean squared residue is not zero. In addition, we provide an interesting result: the mean squared residue is highly dependent on the variance of the scaling factor, which makes possible that any algorithm based on this measure might not find these patterns in data when the variance of gene values is high. The main contribution of this paper is to prove that the mean squared residue is not precise enough from the mathematical point of view in order to discover shifting and scaling patterns at the same time. CONTACT: aguilar@lsi.us.es.

Algorithms↗

Comments on "The multisynapse neural network and its application to fuzzy clustering".

In the above-mentioned paper, Wei and Fahn proposed a neural architecture, the multisynapse neural network, to solve constrained optimization problems including high-order, logarithmic, and sinusoidal forms, etc. As one of its main applications, a fuzzy bidirectional associative clustering network (FBACN) was proposed for fuzzy-partition clustering according to the objective-functional method. The connection between the objective-functional-based fuzzy c-partition algorithms and FBACN is the Lagrange multiplier approach. Unfortunately, the Lagrange multiplier approach was incorrectly applied so that FBACN does not equivalently minimize its corresponding constrained objective-function. Additionally, Wei and Fahn adopted traditional definition of fuzzy c-partition, which is not satisfied by FBACN. Therefore, FBACN can not solve constrained optimization problems, either.

Algorithms↗

AMADA: analysis of microarray data.

SUMMARY: AMADA is a Windows program for identifying co-expressed genes from microarray data. It performs data transformation, principal component analysis, a variety of cluster analyses and extensive graphic functions for visualizing expression profiles.

Algorithms↗

Evaluation and optimization of clustering in gene expression data analysis.

MOTIVATION: A measurement of cluster quality is needed to choose potential clusters of genes that contain biologically relevant patterns of gene expression. This is strongly desirable when a large number of gene expression profiles have to be analyzed and proper clusters of genes need to be identified for further analysis, such as the search for meaningful patterns, identification of gene functions or gene response analysis. RESULTS: We propose a new cluster quality method, called stability, by which unsupervised learning of gene expression data can be performed efficiently. The method takes into account a cluster's stability on partition. We evaluate this method and demonstrate its performance using four independent, real gene expression and three simulated datasets. We demonstrate that our method outperforms other techniques listed in the literature. The method has applications in evaluating clustering validity as well as identifying stable clusters. AVAILABILITY: Please contact the first author.

Algorithms↗

Complex networks approach to gene expression driven phenotype imaging.

MOTIVATION: The need is to visualize and quantify gene expression spatial patterns. Because of their generality for representation of interaction among several elements, complex networks are used to measure the spatial interactions and adjacencies defined by gene expression patterns. RESULTS: Enhanced visualization of spatial interactions between elements where genes are expressed is possible, allowing the identification of structures which would go unnoticed by using conventional imaging. The quantification of the expression intensity in terms of the node degree and clustering coefficient allows the identification of different types of interactions, yielding insights about cell signaling and differentiation, and providing the basis for comparison and discrimination of the patterns along the developmental stages. AVAILABILITY: Supplementary Material, including visualizations as well as the basic routines for translating gene expression images into complex networks and obtaining node degree and clustering coefficient measurements, are provided. CONTACT: luciano@if.sc.usp.br; diambra@univap.br.

Algorithms↗

Structure of the Na(x)Cl(x+1) (-) (x=1-4) clusters via ab initio genetic algorithm and photoelectron spectroscopy.

The application of the ab initio genetic algorithm with an embedded gradient has been carried out for the elucidation of global minimum structures of a series of anionic sodium chloride clusters, Na(x)Cl(x+1) (-) (x=1-4), produced in the gas phase using electrospray ionization and studied by photoelectron spectroscopy. These are all superhalogen species with extremely high electron binding energies. The vertical electron detachment energies for Na(x)Cl(x+1) (-) were measured to be 5.6, 6.46, 6.3, and 7.0 eV, for x=1-4, respectively. Our ab initio gradient embedded genetic algorithm program detected the linear global minima for NaCl(2) (-) and Na(2)Cl(3) (-) and three-dimensional structures for the larger species. Na(3)Cl(4) (-) was found to have C(3v) symmetry, which can be viewed as a Na(4)Cl(4) cube missing a corner Na(+) cation, whereas Na(4)Cl(5) (-) was found to have C(4v) symmetry, close to a 3x3 planar structure. Excellent agreement between the theoretically calculated and the experimental spectra was observed, confirming the obtained structures and demonstrating the power of the developed genetic algorithm technique.

Computer Simulation↗

Cluster analysis of gene expression data based on self-splitting and merging competitive learning.

Cluster analysis of gene expression data from a cDNA microarray is useful for identifying biologically relevant groups of genes. However, finding the natural clusters in the data and estimating the correct number of clusters are still two largely unsolved problems. In this paper, we propose a new clustering framework that is able to address both these problems. By using the one-prototype-take-one-cluster (OPTOC) competitive learning paradigm, the proposed algorithm can find natural clusters in the input data, and the clustering solution is not sensitive to initialization. In order to estimate the number of distinct clusters in the data, we propose a cluster splitting and merging strategy. We have applied the new algorithm to simulated gene expression data for which the correct distribution of genes over clusters is known a priori. The results show that the proposed algorithm can find natural clusters and give the correct number of clusters. The algorithm has also been tested on real gene expression changes during yeast cell cycle, for which the fundamental patterns of gene expression and assignment of genes to clusters are well understood from numerous previous studies. Comparative studies with several clustering algorithms illustrate the effectiveness of our method.

Algorithms↗