PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Clustering under the line graph transformation: application to reaction network.

BACKGROUND: Many real networks can be understood as two complementary networks with two kind of nodes. This is the case of metabolic networks where the first network has chemical compounds as nodes and the second one has nodes as reactions. In general, the second network may be related to the first one by a technique called line graph transformation (i.e., edges in an initial network are transformed into nodes). Recently, the main topological properties of the metabolic networks have been properly described by means of a hierarchical model. While the chemical compound network has been classified as hierarchical network, a detailed study of the chemical reaction network had not been carried out. RESULTS: We have applied the line graph transformation to a hierarchical network and the degree-dependent clustering coefficient C(k) is calculated for the transformed network. C(k) indicates the probability that two nearest neighbours of a vertex of degree k are connected to each other. While C(k) follows the scaling law C(k) approximately k(-1.1) for the initial hierarchical network, C(k) scales weakly as k0.08 for the transformed network. This theoretical prediction was compared with the experimental data of chemical reactions from the KEGG database finding a good agreement. CONCLUSIONS: The weak scaling found for the transformed network indicates that the reaction network can be identified as a degree-independent clustering network. By using this result, the hierarchical classification of the reaction network is discussed.

Algorithms↗

'Gene shaving' as a method for identifying distinct sets of genes with similar expression patterns.

BACKGROUND: Large gene expression studies, such as those conducted using DNA arrays, often provide millions of different pieces of data. To address the problem of analyzing such data, we describe a statistical method, which we have called 'gene shaving'. The method identifies subsets of genes with coherent expression patterns and large variation across conditions. Gene shaving differs from hierarchical clustering and other widely used methods for analyzing gene expression studies in that genes may belong to more than one cluster, and the clustering may be supervised by an outcome measure. The technique can be 'unsupervised', that is, the genes and samples are treated as unlabeled, or partially or fully supervised by using known properties of the genes or samples to assist in finding meaningful groupings. RESULTS: We illustrate the use of the gene shaving method to analyze gene expression measurements made on samples from patients with diffuse large B-cell lymphoma. The method identifies a small cluster of genes whose expression is highly predictive of survival. CONCLUSIONS: The gene shaving method is a potentially useful tool for exploration of gene expression data and identification of interesting clusters of genes worth further investigation.

Algorithms↗

Recent advances in gene expression data clustering: a case study with comparative results.

Several advanced techniques have been proposed for data clustering and many of them have been applied to gene expression data, with partial success. The high dimensionality and the multitude of admissible perspectives for data analysis of gene expression require additional computational resources, such as hierarchical structures and dynamic allocation of resources. We present an immune-inspired hierarchical clustering device, called hierarchical artificial immune network (HaiNet), especially devoted to the analysis of gene expression data. This technique was applied to a newly generated data set, involving maize plants exposed to different aluminum concentrations. The performance of the algorithm was compared with that of a self-organizing map, which is commonly adopted to deal with gene expression data sets. More consistent and informative results were obtained with HaiNet.

Algorithms↗

Integrating contextual information to enhance SOM-based text document clustering.

Exploration of text corpora using self-organizing maps has shown promising results in recent years. Topographic map approaches usually use the original vector space model known from Information Retrieval for text document representation. In this paper I present a two stage model using features based on sentence categories as alternative approach which includes contextual information. Algorithmic optimizations required by this computationally expensive model are shown and evaluated. Also a method for model independent comparison of document maps by evaluation of document distribution on maps is introduced and used to compare results obtained with both the new model and the vector space model.

Algorithms↗

Clustering the annotation space of proteins.

BACKGROUND: Current protein clustering methods rely on either sequence or functional similarities between proteins, thereby limiting inferences to one of these areas. RESULTS: Here we report a new approach, named CLAN, which clusters proteins according to both annotation and sequence similarity. This approach is extremely fast, clustering the complete SwissProt database within minutes. It is also accurate, recovering consistent protein families agreeing on average in more than 97% with sequence-based protein families from Pfam. Discrepancies between sequence- and annotation-based clusters were scrutinized and the reasons reported. We demonstrate examples for each of these cases, and thoroughly discuss an example of a propagated error in SwissProt: a vacuolar ATPase subunit M9.2 erroneously annotated as vacuolar ATP synthase subunit H. CLAN algorithm is available from the authors and the CLAN database is accessible at http://maine.ebi.ac.uk:8000/cgi-bin/clan/ClanSearch.pl CONCLUSIONS: CLAN creates refined function-and-sequence specific protein families that can be used for identification and annotation of unknown family members. It also allows easy identification of erroneous annotations by spotting inconsistencies between similarities on annotation and sequence levels.

Adenosine Triphosphatases↗

Network constrained clustering for gene microarray data.

UNLABELLED: Many bioinformatics problems can be tackled from a fresh angle offered by the network perspective. Directly inspired by metabolic network structural studies, we propose an improved gene clustering approach for inferring gene signaling pathways from gene microarray data. Based on the construction of co-expression networks that consists of both significantly linear and non-linear gene associations together with controlled biological and statistical significance, our approach tends to group functionally related genes into tight clusters despite their expression dissimilarities. We illustrate our approach and compare it to the traditional clustering approaches on a yeast galactose metabolism dataset and a retinal gene expression dataset. Our approach greatly outperforms the traditional approach in rediscovering the relatively well known galactose metabolism pathway in yeast and in clustering genes of the photoreceptor differentiation pathway. AVAILABILITY: The clustering method has been implemented in an R package "GeneNT" that is freely available from: http://www.cran.org.

Algorithms↗

Sample size calculator for cluster randomized trials.

Cluster randomized trials, where individuals are randomized in groups are increasingly being used in healthcare evaluation. The adoption of a clustered design has implications for design, conduct and analysis of studies. In particular, standard sample sizes have to be inflated for cluster designs, as outcomes for individuals within clusters may be correlated; inflation can be achieved either by increasing the cluster size or by increasing the number of clusters in the study. A sample size calculator is presented for calculating appropriate sample sizes for cluster trials, whilst allowing the implications of both methods of inflation to be considered.

Algorithms↗

Machaon CVE: cluster validation for gene expression data.

UNLABELLED: This paper presents a cluster validation tool for gene expression data. Machaon CVE (Clustering and Validation Environment) system aims to partition samples or genes into groups characterized by similar expression patterns, and to evaluate the quality of the clusters obtained. AVAILABILITY: The program is freely available for non-profit use on request at http://www.cs.tcd.ie/Nadia.Bolshakova/Machaon.html SUPPLEMENTARY INFORMATION: http://www.cs.tcd.ie/Nadia.Bolshakova/Machaon.html

Algorithms↗

ClutrFree: cluster tree visualization and interpretation.

UNLABELLED: ClutrFree facilitates the visualization and interpretation of clusters or patterns computed from microarray data through a graphical user interface that displays patterns, membership information of the genes and annotation statistics simultaneously. ClutrFree creates a tree linking the patterns based on similarity, permitting the navigation among patterns identified by different algorithms or by the same algorithm with different parameters, and aids the inferring of conclusions from a microarray experiment. AVAILABILITY: The ClutrFree Java source code and compiled bytecode are available as a package under the GNU General Public License at http://bioinformatics.fccc.edu

Algorithms↗

A semiparametric mixture model for analyzing clustered competing risks data.

A very general class of multivariate life distributions is considered for analyzing failure time clustered data that are subject to censoring and multiple modes of failure. Conditional on cluster-specific quantities, the joint distribution of the failure time and event indicator can be expressed as a mixture of the distribution of time to failure due to a certain type (or specific cause), and the failure type distribution. We assume here the marginal probabilities of various failure types are logistic functions of some covariates. The cluster-specific quantities are subject to some unknown distribution that causes frailty. The unknown frailty distribution is modeled nonparametrically using a Dirichlet process. In such a semiparametric setup, a hybrid method of estimation is proposed based on the i.i.d. Weighted Chinese Restaurant algorithm that helps us generate observations from the predictive distribution of the frailty. The Monte Carlo ECM algorithm plays a vital role for obtaining the estimates of the parameters that assess the extent of the effects of the causal factors for failures of a certain type. A simulation study is conducted to study the consistency of our methodology. The proposed methodology is used to analyze a real data set on HIV infection of a cohort of female prostitutes in Senegal.

AIDS Vaccines↗

A neural network-based similarity index for clustering DNA microarray data.

A common approach to the analysis of gene expression data is to define clusters of genes that have similar expression. A critical step in cluster analysis is the determination of similarity between the expression levels of two genes. We introduce a neural network-based similarity index as a non-linear similarity index and compare the results with other proximity measures for Saccharomyces cerevisiae gene expression data. We show that the clusters obtained using Euclidean distance, correlation coefficients, and mutual information were not significantly different. The clusters formed with the neural network-based index were more in agreement with those defined by functional categories and common regulatory motifs.

Algorithms↗

Clustering ensembles of neural network models.

We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models' outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. A parallel application is the prediction of newspaper sales, where we learn a series of parallel tasks. The results indicate that it is not necessary to store all samples in the ensembles: a small number of representative models generally matches, or even surpasses, the performance of the full ensemble. The clustered representation of the ensemble obtained thus is much better suitable for qualitative analysis, and will be shown to yield new insights into the data.

Algorithms↗

Study of sleep-wakefulness states by computer graphics and cluster analysis before and after lesions of the pontine tegmentum in the cat.

A computerized method of quantification, graphic representation and classification of sleep-wakefulness data in the cat before and after pontine tegmental lesions has been presented. Electrophysiological signal features including average EEG amplitude, average EMG amplitude and PGO spike rate which are particularly important for the definition of sleep-wakefulness states have been quantified for each 1 min epoch in the day. The data were presented in a projected 3-dimensional data display, in which they formed clusters that are considered to be analogous to sleep-wakefulness states. A cluster analysis algorithm was employed for the automatic classification of these data, and this automatic classification was compared graphically and with contingency table analyses to traditional visual assessment of state from polygraphic records. Although there were systematic differences in the locations of state boundaries, total percent agreement between cluster analysis classification and traditional human classification was comparable to the percent agreement between any two human classifiers (about 90%). After pontine tegmental lesions involving both the gigantocellular and lateral tegmental fields, paradoxical sleep was eliminated, and the characteristics of slow wave sleep and wakefulness were altered. The elimination of the state of paradoxical sleep was evident in the computer display by the absence of the paradoxical sleep cluster, and alterations of the other states were indicated in the display by shifts in the positions of their respective clusters. Automatic classification of slow wave sleep and wakefulness after such lesions compared well with traditional classification, attesting to the validity of this approach.

Animals↗

Ultra-violet resonance Raman spectroscopy for the rapid discrimination of urinary tract infection bacteria.

The ability to identify pathogenic organisms rapidly provides significant benefits to clinicians; in particular, with respect to best prescription practices and tracking of recurrent infections. Conventional bioassays require 3-5 days before identification of an organism can be made, thus compromising the effectiveness with which patients can be treated for bacterial infections. We analysed 20 clinical isolates of urinary tract infections (UTI) by ultra-violet resonance Raman (UVRR) spectroscopy, utilising 244 nm excitation delivering approximately 0.1 mW laser power at the sample, with typical spectral collection times of 120 s. UVRR results in resonance-enhanced Raman signals for certain chromophoric segments of macromolecules, intensifying those selected bands above what would otherwise be observed for a normal Raman experiment. Utilising the whole-organism 'fingerprints' obtained by UVRR we were able to discriminate successfully between UTI pathogens using chemometric cluster analyses. This work demonstrates significant improvements in the speed with which spectra can be obtained by Raman spectroscopic techniques for the discrimination of clinical bacterial samples.

Algorithms↗

Region growing method for the analysis of functional MRI data.

Existing analytical techniques for functional magnetic resonance imaging (fMRI) data always need some specific assumptions on the time series. In this article, we present a new approach for fMRI activation detection, which can be implemented without any assumptions on the time series. Our method is based on a region growing method, which is very popular for image segmentation. A comparison of performance on fMRI activation detection is made between the proposed method and the deconvolution method and the fuzzy clustering method with receiver operating characteristic (ROC) methodology. In addition, we examine the effectiveness and usefulness of our method on real experimental data. Experimental results show that our method outperforms over the deconvolution method and the fuzzy clustering method on a number of aspects. These results suggest that our region growing method can serve as a reliable analysis of fMRI data.

Acoustic Stimulation↗

Target space for structural genomics revisited.

MOTIVATION: Structural genomics eventually aims at determining structures for all proteins. However, in the beginning experimentalists are likely to focus on globular proteins to achieve a rapid basic coverage of protein sequence space. How many proteins will structural genomics have to target? How many proteins will be excluded since we already have structural information for these or since they are not globular? We have to answer these questions in the context of our target selection for the North-East Structural Genomics Consortium (NESG). RESULTS: We estimated that structural information is available for about 6-38% of all proteins; 6% if we require high accuracy in comparative modelling, 38% if we are satisfied with having a rough idea about the fold. Excluding all regions that are not globular, we found that structural genomics may have to target about 48% of all proteins. This corresponded to a similar percentage of residues of the entire proteomes (52%). We explored a number of different strategies to cluster protein space in order to find the number of families representing these 48% of structurally unknown proteins. For the subset of all entirely sequenced eukaryotes, we found over 18 000 fragment clusters each of which may be a suitable target for structural genomics. AVAILABILITY: All data are available from the authors, most results are summarized at: http://cubic.bioc.columbia.edu/genomes/RES/2002_bioinformatics/

Algorithms↗

Gender differences in the organization of sexual information.

One widely used model of knowledge representation is that of a network in which concepts are portrayed as nodes with links between the nodes representing associations. Schvaneveldt (1990) developed a method (Pathfinder) that generates associative networks from individual's ratings of similarity of word pairs. We had 51 female and 47 male undergraduates rate for similarity all paired combinations of 16 words judged as relevant to the domain of sexuality. Using a measure of network similarity, we found that each gender's networks were more similar to each other than they were to the other gender's. Using the number of links on words as the dependent variable there were gender differences in the number of links within clusters of words, between clusters of words, and on specific words. These differences, for the most part, are consistent with gender stereotypes and prior research showing gender differences in the processing of sexual information.

Algorithms↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗