PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Clustering gene expression data based on predicted differential effects of GV interaction.

Microarray has become a popular biotechnology in biological and medical research. However, systematic and stochastic variabilities in microarray data are expected and unavoidable, resulting in the problem that the raw measurements have inherent "noise" within microarray experiments. Currently, logarithmic ratios are usually analyzed by various clustering methods directly, which may introduce bias interpretation in identifying groups of genes or samples. In this paper, a statistical method based on mixed model approaches was proposed for microarray data cluster analysis. The underlying rationale of this method is to partition the observed total gene expression level into various variations caused by different factors using an ANOVA model, and to predict the differential effects of GV (gene by variety) interaction using the adjusted unbiased prediction (AUP) method. The predicted GV interaction effects can then be used as the inputs of cluster analysis. We illustrated the application of our method with a gene expression dataset and elucidated the utility of our approach using an external validation.

Algorithms↗

A mixed factors model for dimension reduction and extraction of a group structure in gene expression data.

When we cluster tissue samples on the basis of genes, the number of observations to be grouped is much smaller than the dimension of feature vector. In such a case, the applicability of conventional model-based clustering is limited since the high dimensionality of feature vector leads to overfitting during the density estimation process. To overcome such difficulty, we attempt a methodological extension of the factor analysis. Our approach enables us not only to prevent from the occurrence of overfitting, but also to handle the issues of clustering, data compression and extracting a set of genes to be relevant to explain the group structure. The potential usefulness are demonstrated with the application to the leukemia dataset.

Algorithms↗

Graph visualization techniques for web clustering engines.

One of the most challenging issues in mining information from the World Wide Web is the design of systems that present the data to the end user by clustering them into meaningful semantic categories. We show that the analysis of the results of a clustering engine can significantly take advantage of enhanced graph drawing and visualization techniques. We propose a graph-based user interface for Web clustering engines that makes it possible for the user to explore and visualize the different semantic categories and their relationships at the desired level of detail.

Algorithms↗

Coupled two-way clustering server.

UNLABELLED: The CTWC server provides access to the software, CTWC1.00, that implements Coupled Two Way Clustering (Getz et al., 2000), a method designed to mine gene expression data AVAILABILITY: Free, at http://ctwc.weizmann.ac.il. SUPPLEMENTARY INFORMATION: The site has a link to an example which provides figures and detailed explanations

Algorithms↗

Transformation of expression intensities across generations of Affymetrix microarrays using sequence matching and regression modeling.

The utility of previously generated microarray data is severely limited owing to small study size, leading to under-powered analysis, and failure of replication. Multiplicity of platforms and various sources of systematic noise limit the ability to compile existing data from similar studies. We present a model for transformation of data across different generations of Affymetrix arrays, developed using previously published datasets describing technical replicates performed with two generations of arrays. The transformation is based upon a probe set-specific regression model, generated from replicate measurements across platforms, performed using correlation coefficients. The model, when applied to the expression intensities of 5069 shared, sequence-matched probe sets in three different generations of Affymetrix Human oligonucleotide arrays, showed significant improvement in inter generation correlations between sample-wide means and individual probe set pairs. The approach was further validated by an observed reduction in Euclidean distance between signal intensities across generations for the predicted values. Finally, application of the model to independent, but related datasets resulted in improved clustering of samples based upon their biological, as opposed to technical, attributes. Our results suggest that this transformation method is a valuable tool for integrating microarray datasets from different generations of arrays.

Algorithms↗

Analysis of IR spectra of carbon clusters trapped in noble gas matrices using algorithms based on digital filtering techniques.

IR spectra of carbon clusters C(n) (n>/=3) trapped in noble gas matrices (Ar and Kr) at temperatures of 13, 30 and 35 K are analyzed using algorithms based on digital filtering techniques and non-linear least-squares fitting. The spectral features at different temperatures and in different matrices are discussed. The majority of the resolved lines are assigned to C(n) cluster species based on data obtained via quantum chemical computations and from tunable laser IR spectrometry of C(n) species in the gas phase. A complete analysis of the IR spectra is available upon request.

Journal Article↗

Comparison of microarray designs for class comparison and class discovery.

MOTIVATION: Two-color microarray experiments in which an aliquot derived from a common RNA sample is placed on each array are called reference designs. Traditionally, microarray experiments have used reference designs, but designs without a reference have recently been proposed as alternatives. RESULTS: We develop a statistical model that distinguishes the different levels of variation typically present in cancer data, including biological variation among RNA samples, experimental error and variation attributable to phenotype. Within the context of this model, we examine the reference design and two designs which do not use a reference, the balanced block design and the loop design, focusing particularly on efficiency of estimates and the performance of cluster analysis. We calculate the relative efficiency of designs when there are a fixed number of arrays available, and when there are a fixed number of samples available. Monte Carlo simulation is used to compare the designs when the objective is class discovery based on cluster analysis of the samples. The number of discrepancies between the estimated clusters and the true clusters were significantly smaller for the reference design than for the loop design. The efficiency of the reference design relative to the loop and block designs depends on the relation between inter- and intra-sample variance. These results suggest that if cluster analysis is a major goal of the experiment, then a reference design is preferable. If identification of differentially expressed genes is the main concern, then design selection may involve a consideration of several factors.

Algorithms↗

HMMGEP: clustering gene expression data using hidden Markov models.

SUMMARY: The package HMMGEP performs cluster analysis on gene expression data using hidden Markov models. AVAILABILITY: HMMGEP, including the source code, documentation and sample data files, is available at http://www.bioinfo.tsinghua.edu.cn:8080/~rich/hmmgep_download/index.html.

Algorithms↗

A geometric invariant-based framework for the analysis of protein conformational space.

MOTIVATION: Characterization of the restricted nature of the protein local conformational space has remained a challenge, thereby necessitating a computationally expensive conformational search in protein modeling. Moreover, owing to the lack of unilateral structural descriptors, conventional data mining techniques, such as clustering and classification, have not been applied in protein structure analysis. RESULTS: We first map the local conformations in a fixed dimensional space by using a carefully selected suite of geometric invariants (GIs) and then reduce the number of dimensions via principal component analysis (PCA). Distribution of the conformations in the space spanned by the first four PCs is visualized as a set of conditional bivariate probability distribution plots, where the peaks correspond to the preferred conformations. The locations of the different canonical structures in the PC-space have been interpreted in the context of the weights of the GIs to the first four PCs. Clustering of the available conformations reveals that the number of preferred local conformations is several orders of magnitude smaller than that suggested previously. SUPPLEMENTARY INFORMATION: www.it.iitb.ac.in/~ashish/bioinfo2005/.

Algorithms↗

The topological relationship between the large-scale attributes and local interaction patterns of complex networks.

Recent evidence indicates that the abundance of recurring elementary interaction patterns in complex networks, often called subgraphs or motifs, carry significant information about their function and overall organization. Yet, the underlying reasons for the variable quantity of different subgraph types, their propensity to form clusters, and their relationship with the networks' global organization remain poorly understood. Here we show that a network's large-scale topological organization and its local subgraph structure mutually define and predict each other, as confirmed by direct measurements in five well studied cellular networks. We also demonstrate the inherent existence of two distinct classes of subgraphs, and show that, in contrast to the low-density type II subgraphs, the highly abundant type I subgraphs cannot exist in isolation but must naturally aggregate into subgraph clusters. The identified topological framework may have important implications for our understanding of the origin and function of subgraphs in all complex networks.

Algorithms↗

ArrayQuest: a web resource for the analysis of DNA microarray data.

BACKGROUND: Numerous microarray analysis programs have been created through the efforts of Open Source software development projects. Providing browser-based interfaces that allow these programs to be executed over the Internet enhances the applicability and utility of these analytic software tools. RESULTS: Here we present ArrayQuest, a web-based DNA microarray analysis process controller. Key features of ArrayQuest are that (1) it is capable of executing numerous analysis programs such as those written in R, BioPerl and C++; (2) new analysis programs can be added to ArrayQuest Methods Library at the request of users or developers; (3) input DNA microarray data can be selected from public databases (i.e., the Medical University of South Carolina (MUSC) DNA Microarray Database or Gene Expression Omnibus (GEO)) or it can be uploaded to the ArrayQuest center-point web server into a password-protected area; and (4) analysis jobs are distributed across computers configured in a backend cluster. To demonstrate the utility of ArrayQuest we have populated the methods library with methods for analysis of Affymetrix DNA microarray data. CONCLUSION: ArrayQuest enables browser-based implementation of DNA microarray data analysis programs that can be executed on a Linux-based platform. Importantly, ArrayQuest is a platform that will facilitate the distribution and implementation of new analysis algorithms and is therefore of use to both developers of analysis applications as well as users. ArrayQuest is freely available for use at http://proteogenomics.musc.edu/arrayquest.html.

Algorithms↗

Identification of the binding sites of regulatory proteins in bacterial genomes.

We present an algorithm that extracts the binding sites (represented by position-specific weight matrices) for many different transcription factors from the regulatory regions of a genome, without the need for delineating groups of coregulated genes. The algorithm uses the fact that many DNA-binding proteins in bacteria bind to a bipartite motif with two short segments more conserved than the intervening region. It identifies all statistically significant patterns of the form W(1)N(x)W(2), where W(1) and W(2) are two short oligonucleotides separated by x arbitrary bases, and groups them into clusters of similar patterns. These clusters are then used to derive quantitative recognition profiles of putative regulatory proteins. For a given cluster, the algorithm finds the matching sequences plus the flanking regions in the genome and performs a multiple sequence alignment to derive position-specific weight matrices. We have analyzed the Escherichia coli genome with this algorithm and found approximately 1,500 significant patterns, which give rise to approximately 160 distinct position-specific weight matrices. A fraction of these matrices match the binding sites of one-third of the approximately 60 characterized transcription factors with high statistical significance. Many of the remaining matrices are likely to describe binding sites and regulons of uncharacterized transcription factors. The significance of these matrices was evaluated by their specificity, the location of the predicted sites, and the biological functions of the corresponding regulons, allowing us to suggest putative regulatory functions. The algorithm is efficient for analyzing newly sequenced bacterial genomes for which little is known about transcriptional regulation.

Algorithms↗

Co-clustering of biological networks and gene expression data.

MOTIVATION: Large scale gene expression data are often analysed by clustering genes based on gene expression data alone, though a priori knowledge in the form of biological networks is available. The use of this additional information promises to improve exploratory analysis considerably. RESULTS: We propose constructing a distance function which combines information from expression data and biological networks. Based on this function, we compute a joint clustering of genes and vertices of the network. This general approach is elaborated for metabolic networks. We define a graph distance function on such networks and combine it with a correlation-based distance function for gene expression measurements. A hierarchical clustering and an associated statistical measure is computed to arrive at a reasonable number of clusters. Our method is validated using expression data of the yeast diauxic shift. The resulting clusters are easily interpretable in terms of the biochemical network and the gene expression data and suggest that our method is able to automatically identify processes that are relevant under the measured conditions.

Algorithms↗

[Research on gene expression data based on clustering/classification technology].

As the work of sequencing the genome of the human and many model organisms has been partially or fully finished, the "postgenomic era" has begun. Scientists are turning their focus toward identifying gene function from sequencing. Clustering technology, as one of the important tools of analyzing gene expression data and identifying gene function, has been used widely. In this paper we discuss main clustering technology about gene expression data at present, analyze their advantages and disadvantages, present the methods to solve the problems and give new approaches to study gene expression data.

Algorithms↗

Lip image segmentation using fuzzy clustering incorporating an elliptic shape function.

Recently, lip image analysis has received much attention because its visual information is shown to provide improvement for speech recognition and speaker authentication. Lip image segmentation plays an important role in lip image analysis. In this paper, a new fuzzy clustering method for lip image segmentation is presented. This clustering method takes both the color information and the spatial distance into account while most of the current clustering methods only deal with the former. In this method, a new dissimilarity measure, which integrates the color dissimilarity and the spatial distance in terms of an elliptic shape function, is introduced. Because of the presence of the elliptic shape function, the new measure is able to differentiate the pixels having similar color information but located in different regions. A new iterative algorithm for the determination of the membership and centroid for each class is derived, which is shown to provide good differentiation between the lip region and the nonlip region. Experimental results show that the new algorithm yields better membership distribution and lip shape than the standard fuzzy c-means algorithm and four other methods investigated in the paper.

Algorithms↗

A mixture model-based approach to the clustering of microarray expression data.

MOTIVATION: This paper introduces the software EMMIX-GENE that has been developed for the specific purpose of a model-based approach to the clustering of microarray expression data, in particular, of tissue samples on a very large number of genes. The latter is a nonstandard problem in parametric cluster analysis because the dimension of the feature space (the number of genes) is typically much greater than the number of tissues. A feasible approach is provided by first selecting a subset of the genes relevant for the clustering of the tissue samples by fitting mixtures of t distributions to rank the genes in order of increasing size of the likelihood ratio statistic for the test of one versus two components in the mixture model. The imposition of a threshold on the likelihood ratio statistic used in conjunction with a threshold on the size of a cluster allows the selection of a relevant set of genes. However, even this reduced set of genes will usually be too large for a normal mixture model to be fitted directly to the tissues, and so the use of mixtures of factor analyzers is exploited to reduce effectively the dimension of the feature space of genes. RESULTS: The usefulness of the EMMIX-GENE approach for the clustering of tissue samples is demonstrated on two well-known data sets on colon and leukaemia tissues. For both data sets, relevant subsets of the genes are able to be selected that reveal interesting clusterings of the tissues that are either consistent with the external classification of the tissues or with background and biological knowledge of these sets. AVAILABILITY: EMMIX-GENE is available at http://www.maths.uq.edu.au/~gjm/emmix-gene/

Algorithms↗

Differential coexpression analysis using microarray data and its application to human cancer.

MOTIVATION: Microarrays have been used to identify differential expression of individual genes or cluster genes that are coexpressed over various conditions. However, alteration in coexpression relationships has not been studied. Here we introduce a model for finding differential coexpression from microarrays and test its biological validity with respect to cancer. RESULTS: We collected 10 published gene expression datasets from cancers of 13 different tissues and constructed 2 distinct coexpression networks: a tumor network and normal network. Comparison of the two networks showed that cancer affected many coexpression relationships. Functional changes such as alteration in energy metabolism, promotion of cell growth and enhanced immune activity were accompanied with coexpression changes. Coregulation of collagen genes that may control invasion and metastatic spread of tumor cells was also found. Cluster analysis in the tumor network identified groups of highly interconnected genes related to ribosomal protein synthesis, the cell cycle and antigen presentation. Metallothionein expression was also found to be clustered, which may play a role in apoptosis control in tumor cells. Our results show that this model would serve as a novel method for analyzing microarrays beyond the specific implications for cancer.

Algorithms↗

Design of long oligonucleotide probes for functional gene detection in a microbial community.

MOTIVATION: Analysis of the functions of microorganisms and their dynamics in the environment is essential for understanding microbial ecology. For analysis of highly similar sequences of a functional gene family using microarrays, the previous long oligonucleotide probe design strategies have not been useful in generating probes. RESULTS: We developed a Hierarchical Probe Design (HPD) program that designs both sequence-specific probes and hierarchical cluster-specific probes from sequences of a conserved functional gene based on the clustering tree of the genes, specifically for analyses of functional gene diversity in environmental samples. HPD was tested on datasets for the nirS and pmoA genes. Our results showed that HPD generated more sequence-specific probes than several popular oligonucleotide design programs. With a combination of sequence-specific and cluster-specific probes, HPD generated a probe set covering all the sequences of each test set. AVAILABILITY: http://brcapp.kribb.re.kr/HPD/

Algorithms↗