PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Study of protein-protein interaction using conformational space annealing.

We apply conformational space annealing (CSA), an efficient global optimization method, to the study of protein-protein interaction. The CSA is incorporated into the Tinker molecular modeling package along with a B-spline method for CAPRI Round 5 experiments. We have used an energy function for the protein-protein interaction that consists of electrostatic interaction, van der Waals interaction, and solvation energy terms represented by the occupancy desolvation method. The parameters of the AMBER94 all-atom empirical force field are used. Each energy term is calculated by precalculated grid potentials and B-spline method approximation. The ligand protein is placed inside a sphere of 50 A radius centered at an appropriate location, and the CSA rigid docking studies are carried out to find stable complexes. Up to 10 complexes are selected using the K-mean clustering method and biological information when available. These complexes are energy-minimized for further refinement by considering the flexibility of interacting proteins. The results show that the CSA method has a potential for the study of protein-protein interaction.

Algorithms↗

MatArray: a Matlab toolbox for microarray data.

The microarray technology allows the high-throughput quantification of the mRNA level of thousands of genes under dozens of conditions, generating a wealth of data which must be analyzed using some form of computational means. A popular framework for such analysis is Matlab, a powerful computing language for which many functions have been written. However, although complex topics like neural networks or principal component analysis are freely available in Matlab, functions to perform more basic tasks like data normalization or hierarchical clustering in an efficient manner are not. The MatArray toolbox aims at filling this gap by offering efficient implementations of the most needed functions for microarray analysis. The functions in the toolbox are command-line only, since it is geared toward seasoned Matlab users.

Algorithms↗

Towards clustering of incomplete microarray data without the use of imputation.

MOTIVATION: Clustering technique is used to find groups of genes that show similar expression patterns under multiple experimental conditions. Nonetheless, the results obtained by cluster analysis are influenced by the existence of missing values that commonly arise in microarray experiments. Because a clustering method requires a complete data matrix as an input, previous studies have estimated the missing values using an imputation method in the preprocessing step of clustering. However, a common limitation of these conventional approaches is that once the estimates of missing values are fixed in the preprocessing step, they are not changed during subsequent processes of clustering; badly estimated missing values obtained in data preprocessing are likely to deteriorate the quality and reliability of clustering results. Thus, a new clustering method is required for improving missing values during iterative clustering process. RESULTS: We present a method for Clustering Incomplete data using Alternating Optimization (CIAO) in which a prior imputation method is not required. To reduce the influence of imputation in preprocessing, we take an alternative optimization approach to find better estimates during iterative clustering process. This method improves the estimates of missing values by exploiting the cluster information such as cluster centroids and all available non-missing values in each iteration. To test the performance of the CIAO, we applied the CIAO and conventional imputation-based clustering methods, e.g. k-means based on KNNimpute, for clustering two yeast incomplete data sets, and compared the clustering result of each method using the Saccharomyces Genome Database annotations. The clustering results of the CIAO method are more significantly relevant to the biological gene annotations than those of other methods, indicating its effectiveness and potential for clustering incomplete gene expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request.

Algorithms↗

MatrixExplorer: a dual-representation system to explore social networks.

MatrixExplorer is a network visualization system that uses two representations: node-link diagrams and matrices. Its design comes from a list of requirements formalized after several interviews and a participatory design session conducted with social science researchers. Although matrices are commonly used in social networks analysis, very few systems support the matrix-based representations to visualize and analyze networks. MatrixExplorer provides several novel features to support the exploration of social networks with a matrix-based representation, in addition to the standard interactive filtering and clustering functions. It provides tools to reorder (layout) matrices, to annotate and compare findings across different layouts and find consensus among several clusterings. MatrixExplorer also supports Node-link diagram views which are familiar to most users and remain a convenient way to publish or communicate exploration results. Matrix and node-link representations are kept synchronized at all stages of the exploration process.

Algorithms↗

Improved gene selection for classification of microarrays.

In this paper we derive a method for evaluating and improving techniques for selecting informative genes from microarray data. Genes of interest are typically selected by ranking genes according to a test-statistic and then choosing the top k genes. A problem with this approach is that many of these genes are highly correlated. For classification purposes it would be ideal to have distinct but still highly informative genes. We propose three different pre-filter methods--two based on clustering and one based on correlation--to retrieve groups of similar genes. For these groups we apply a test-statistic to finally select genes of interest. We show that this filtered set of genes can be used to significantly improve existing classifiers.

Algorithms↗

Interaction between tracheal sound and flow rate: a comparison of some different flow evaluations from lung sounds.

We simultaneously recorded tracheal sound and air flow from nine normal subjects (seven males and two females). Sound was picked up at the supra sternal notch with an air-coupled sensitive microphone held in a small airtight probe. Flow was measured at the mouth using a pneumotachograph Fleisch n degrees 2. Both sound and flow were directly digitized at a sampling rate of 5120 Hz and then divided in 128-sample blocks. For each sound block the frequency spectrum was computed using the fast Fourier transform. In order to evaluate instantaneous flow-rate from tracheal sounds we investigated eight methods divided in two groups of four. In the first group (i.e., reference curves methods), we assumed that a relationship existed between sound and flow and was thus reflected by the variations of certain parameters. We chose to use simple straightforward relationships, already known and published. We tested four different parameters. During a calibration phase, we built for each parameter P a reference curve representing the variations of P versus flow and being specific to each subject. Then, an unknown flow was evaluated in calculating P on a 128-sample block, and the reference curve gave the corresponding flow. In the second group, we made a hierarchial clustering analysis of sound spectra for revealing the frequency modifications, induced by the flow. We tested two kinds of spectra as well as two ways of associating a flow to a given cluster. This led us to four other methods for calculating the flow. All the eight methods but one gave a mean uncertainty in the measure of flow of about 15%.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent↗

Analyzing Sub-Classifications of Glaucoma via SOM Based Clustering of Optic Nerve Images.

We present a data mining framework to cluster optic nerve images obtained by Confocal Scanning Laser Tomography (CSLT) in normal subjects and patients with glaucoma. We use self-organizing maps and expectation maximization methods to partition the data into clusters that provide insights into potential sub-classification of glaucoma based on morphological features. We conclude that our approach provides a first step towards a better understanding of morphological features in optic nerve images obtained from glaucoma patients and healthy controls.

Algorithms↗

Tissue segmentation on MR images of the brain by possibilistic clustering on a 3D wavelet representation.

An algorithm for the segmentation of a single sequence of three-dimensional magnetic resonance (MR) images into cerebrospinal fluid, gray matter, and white matter classes is proposed. This new method is a possibilistic clustering algorithm using the fuzzy theory as frame and the wavelet coefficients of the voxels as features to be clustered. Fuzzy logic models the uncertainty and imprecision inherent in MR images of the brain, while the wavelet representation allows for both spatial and textural information. The procedure is fast, unsupervised, and totally independent of any statistical assumptions. The method is tested on a phantom image, then applied to normal and Alzheimer's brains, and finally compared with another classic brain tissue segmentation method, affording a relevant classification of voxels into the different tissue classes.

Adult↗

Feature-guided analysis for reduction of false positives in CAD of polyps for computed tomographic colonography.

We evaluated the effect of our novel technique of feature-guided analysis of polyps on the reduction of false-positive (FP) findings generated by our computer-aided diagnosis (CAD) scheme for the detection of polyps from computed tomography colonographic data sets. The detection performance obtained by use of feature-guided analysis in the segmentation and feature analysis of polyp candidates was compared with that obtained by use of our previously employed fuzzy clustering technique. We also evaluated the effect of a feature called modified gradient concentration (MGC) on the detection performance. A total of 144 data sets, representing prone and supine views of 72 patients that included 14 patients with 21 colorectal polyps 5-25 mm in diameter, were used in the evaluation. At a 100% by-patient (95% by-polyp) detection sensitivity, the FP rate of our CAD scheme with feature-guided analysis based on round-robin evaluation was 1.3 (1.5) FP detections per patient. This corresponds to a 70-75% reduction in the number of FPs obtained by use of fuzzy clustering at the same sensitivity levels. Application of the MGC feature instead of our previously used gradient concentration feature did not improve the detection result. The results indicate that feature-guided analysis is useful for achieving high sensitivity and a low FP rate in our CAD scheme.

Algorithms↗

Generating stochastic dispersed and periodic clustered textures using a composite hybrid screen.

In electrophotographic printing, a periodic clustered-dot halftone pattern is preferred for a smooth and stable result. In addition, the screen frequency should be high enough to minimize the visibility of the halftone textures and to ensure good detail rendition. However, at these frequencies, the halftone cell may contain too few pixels to provide a sufficient number of distinct gray levels. This will result in contouring and posterization. The traditional solution is to grow the clusters asynchronously within a repeating block of clusters known as a supercell. The growth of each individual cluster is governed by a microscreen. The order in which the clusters grow within the supercell is determined by a macroscreen. Typically, the macroscreen is a recursive pattern due to Bayer. In highlights and shadows, this ordering results in visible artifacts. Replacing the Bayer screen by a stochastic macroscreen eliminates these artifacts, but results in new artifacts. In this paper, we propose a new composite screen architecture that employs multiple microscreens and multiple macroscreens in the highlights and shadows. These screens are jointly designed by using the direct binary search (DBS) algorithm.

Algorithms↗

FOUNTAIN: a JAVA open-source package to assist large sequencing projects.

BACKGROUND: Better automation, lower cost per reaction and a heightened interest in comparative genomics has led to a dramatic increase in DNA sequencing activities. Although the large sequencing projects of specialized centers are supported by in-house bioinformatics groups, many smaller laboratories face difficulties managing the appropriate processing and storage of their sequencing output. The challenges include documentation of clones, templates and sequencing reactions, and the storage, annotation and analysis of the large number of generated sequences. RESULTS: We describe here a new program, named FOUNTAIN, for the management of large sequencing projects http://genetics.hpi.uni-hamburg.de/FOUNTAIN.html. FOUNTAIN uses the JAVA computer language and data storage in a relational database. Starting with a collection of sequencing objects (clones), the program generates and stores information related to the different stages of the sequencing project using a web browser interface for user input. The generated sequences are subsequently imported and annotated based on BLAST searches against the public databases. In addition, simple algorithms to cluster sequences and determine putative polymorphic positions are implemented. CONCLUSIONS: A simple, but flexible and scalable software package is presented to facilitate data generation and storage for large sequencing projects. Open source and largely platform and database independent, we wish FOUNTAIN to be improved and extended in a community effort.

Algorithms↗

Estimation of the combined response to treatment in multicenter trials.

Analyses of multicenter trials consider the estimated treatment effect differences of the individual centers and combine them into an estimate of the overall treatment effect. There has been much debate in the literature concerning the best way to combine these treatment effect differences. We emphasize that first of all one should define the combined response to treatment (CRT), the object that has to be estimated from the results of a multicenter clinical trial. It is shown that the choice of CRT determines not only the best estimator, but also the allocation of patients among the centers that minimizes the mean squared error. A new estimator of the CRT is proposed that is based on a preliminary clustering of the centers and the use of a weighted average of the Type I estimators obtained from within each cluster. The clustering aims to minimize the bias of the combined estimator. We show via a simulation study that the simple clustering procedure provides a reasonably improved estimator. The clustering can be done on blinded data, as long as the numbers of patients on each treatment arm in each center are known. The methodology is illustrated by analyzing a multicountry, multicenter trial to compare an active treatment with placebo for the treatment of a psychiatric disorder.

Algorithms↗

On the refinement of time-resolved diffraction data: comparison of the random-distribution and cluster-formation models and analysis of the light-induced increase in the atomic displacement parameters.

Expressions for the random-distribution and cluster-formation models for light-induced changes in crystals studied by time-resolved diffraction are presented. The two models can be distinguished on the basis of differences in the predicted intensities. The light-induced increase in the atomic displacement parameters is analyzed with both simulated and experimental data sets.

Algorithms↗

Single-parent evolution algorithm and the optimization of Si clusters.

We describe a novel method for the structural optimization of molecular systems. Similar to genetic algorithms (GA), our approach involves an evolving population in which new members are formed by cutting and pasting operations on existing members. Unlike previous GA's, however, the population in each generation has a single parent only. This scheme has been used to optimize Si clusters with 13-23 atoms. We have found a number of new isomers that are lower in energy than any previously reported and have properties in much better agreement with experimental data.

Journal Article↗

Chemotaxonomy of mints of genus Mentha by applying Raman spectroscopy.

The characterization of mints is often problematic because Mentha is a taxonomically complex genus. In order to provide a fast and easy characterization method, we use a combination of micro-Raman spectroscopy and hierarchical cluster analysis. A classification trial of different mint taxa is possible for one collection time. For spectra measured at different points during the growing season, a more sophisticated pretreatment of the data is necessary to receive good discrimination between the species, as well as between the subspecies and varieties of the mints.

Algorithms↗

The interactome as a tree--an attempt to visualize the protein-protein interaction network in yeast.

The refinement and high-throughput of protein interaction detection methods offer us a protein-protein interaction network in yeast. The challenge coming along with the network is to find better ways to make it accessible for biological investigation. Visualization would be helpful for extraction of meaningful biological information from the network. However, traditional ways of visualizing the network are unsuitable because of the large number of proteins. Here, we provide a simple but information-rich approach for visualization which integrates topological and biological information. In our method, the topological information such as quasi-cliques or spoke-like modules of the network is extracted into a clustering tree, where biological information spanning from protein functional annotation to expression profile correlations can be annotated onto the representation of it. We have developed a software named PINC based on our approach. Compared with previous clustering methods, our clustering method ADJW performs well both in retaining a meaningful image of the protein interaction network as well as in enriching the image with biological information, therefore is more suitable in visualization of the network.

Algorithms↗

Cluster diversity and entropy on the percolation model: the lattice animal identification algorithm

We present an algorithm to identify and count different lattice animals (LA's) in the site-percolation model. This algorithm allows a definition of clusters based on the distinction of cluster shapes, in contrast with the well-known Hoshen-Kopelman algorithm, in which the clusters are differentiated by their sizes. It consists in coding each unit cell of a cluster according to the nearest neighbors (NN) and ordering the codes in a proper sequence. In this manner, a LA is represented by a specific code sequence. In addition, with some modification the algorithm is capable of differentiating between fixed and free LA's. The enhanced Hoshen-Kopelman algorithm [J. Hoshen, M. W. Berry, and K. S. Minser, Phys. Rev. E 56, 1455 (1997)] is used to compose the set of NN code sequences of each cluster. Using Monte Carlo simulations on planar square lattices up to 2000x2000, we apply this algorithm to the percolation model. We calculate the cluster diversity and cluster entropy of the system, which leads to the determination of probabilities associated with the maximum of these functions. We show that these critical probabilities are associated with the percolation transition and with the complexity of the system.

Journal Article↗

A comparison of cluster analysis methods using DNA methylation data.

MOTIVATION: Aberrant DNA methylation is common in cancer. DNA methylation profiles differ between tumor types and subtypes and provide a powerful diagnostic tool for identifying clusters of samples and/or genes. DNA methylation data obtained with the quantitative, highly sensitive MethyLight technology is not normally distributed; it frequently contains an excess of zeros. Established tools to analyze this type of data do not exist. Here, we evaluate a variety of methods for cluster analysis to determine which is most reliable. RESULTS: We introduce a Bernoulli-lognormal mixture model for clustering DNA methylation data obtained using MethyLight. We model the outcomes using a two-part distribution having discrete and continuous components. It is compared with standard cluster analysis approaches for continuous data and for discrete data. In a simulation study, we find that the two-part model has the lowest classification error rate for mixture outcome data compared with other approaches. The methods are illustrated using DNA methylation data from a study of lung cancer cell lines. Compared with competing hierarchical clustering methods, the mixture model approaches have the lowest cross-validation error for detecting lung cancer subtype (non-small versus small cell). The Bernoulli-lognormal mixture assigns observations to subgroups with the lowest uncertainty. AVAILABILITY: Software is available upon request from the authors. SUPPLEMENTARY INFORMATION: http://www-rcf.usc.edu/~kims/SupplementaryInfo.html

Algorithms↗