PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Numerical classification of some named strains of Nocardia asteroides and related isolates from soil.

One hundred and forty-nine strains of nocardiae, freshly isolated from soil samples obtained from a number of countries with either tropical or temperate climates, and from rubber pipe seals, were compared with appropriate marker cultures in a numerical phenetic study using 156 unit characters. Marker strains were chosen to represent the Nocardia asteroides complex, other Nocardia species and related taxa in an effort both to classify the new soil isolates and, possibly, clarify the structure of the heterogeneous N. asteroides complex. The data were examined using the simple matching (SSM) and pattern (DP) coefficients, and clustering was achieved using both single and average linkage algorithms. Cluster composition was not markedly affected by either of the coefficients or clustering methods. The estimated test error of 7.1% was rather high and could account for a few apparently anomalous results. The 16 defined clusters, containing 185 of the 197 strains studied, were divided into seven major and nine minor clusters, four of which were further subdivided into two subclusters. Marker strains allowed four clusters to be designated as N. asteroides, seven as Nocardia species and one each as Nocardia carnea, Nocardia farcinica, Nocardia autotrophica, Mycobacterium farcinogenes and Rhodococcus species. Twelve strains formed single member clusters including the type strains of Nocardia aerocolonigenes, Nocardia amarae, Nocardia fukuyae, Nocardia orientalis and Nocardia otitidis-caviarum. The majority of the soil and rubber isolates were recovered in the major clusters labelled N. asteroides, N. carnea and Nocardia species and clusters of soil isolates without marker strains seem to represent new centres of variation. The study highlights the need for additional reproducible tests to help both define and determine the status of defined clusters within the N. asteroides complex which would considerably benefit both the ecological and epidemiological study of these organisms.

Nocardia asteroides↗

Personality characteristics of substance abusers: an MCMI cluster typology of recreational drug users treated in a therapeutic community and its relationship to length of stay and outcome.

The purpose of this investigation was to define homogeneous personality subtypes among substance abusers treated in a long-term, inpatient, drug-free therapeutic community and to determine how the resulting typology was related to length of stay and treatment outcome. A hierarchical agglomerative cluster analysis was performed on the Millon Clinical Multiaxial Inventory (MCMI) scale scores of 235 admissions to a therapeutic community. Five cluster types emerged, which were similar to typologies found in studies with alcoholic inpatients. A concordant solution evolved when a different clustering algorithm was used with the same sample and when clustering was done with a different group of substance abusers. As hypothesized, clusters of patients with average MCMI elevations that indicated avoidant, schizoid, and antisocial qualities tended to stay in treatment fewer days and relapsed earlier during the 1-year follow-up. The implications for substance abuse treatment are discussed.

Adult↗

SiMCAL 1 algorithm for analysis of gene expression data related to the phosphatidylserine receptor.

OBJECTIVE: SiMCAL 1 (simple multilevel clustering and linking, version 1) is a novel clustering algorithm for time-series microarray data, presented here with an application to a specific data set. The purpose of the algorithm is to present a complete feature set not found in either Jarvis-Patrick clustering, from which it is derived, or in other popular clustering methods such as hierarchical and k-means. The data concern the activity of the phosphatidylserine receptor (PSR) which is believed to be a crucial molecular switch in the mediation of inflammatory response in apoptosis and lysis. By analyzing the behavior of PSR-related genes in mouse macrophages, we hope to elucidate the mechanisms involved in this important biological process. METHODS AND MATERIALS: SiMCAL 1 is implemented in the Python programming language using the Numerical Python extensions, and the data are stored using the MySQL database management system. The data are derived from exposures of multiple Affymetrix mouse gene microarray chips to elevated levels of PSR antibody and control conditions. Code and data are available at (accessed: 17 January 2005). RESULTS: The algorithm meets its objectives: it is simple, in that it is computationally inexpensive; it is multilevel, in that it provides a small number of clearly defined hierarchical levels of clusters; and it offers linking between clusters at the same level in each hierarchy. Clustering and linking results indicate previously unknown co-regulation for genes expressing PGH synthase (COX2) and PGE2, appear to confirm increased production of proteins for clearance of apoptotic cells in the presence of PSR antibody, and correspond to other findings regarding the temporal relationship between PGE2 production and B cell proliferation and differentiation. These results are promising but should be taken as highly preliminary. CONCLUSION: Both the algorithm and its application to this problem show great potential for future development. We plan to improve and extend the SiMCAL family of algorithms, and to obtain new data so that the algorithm(s) may be further applied to this and other problems of interest.

Algorithms↗

Expression profiling of human renal carcinomas with functional taxonomic analysis.

BACKGROUND: Molecular characterization has contributed to the understanding of the inception, progression, treatment and prognosis of cancer. Nucleic acid array-based technologies extend molecular characterization of tumors to thousands of gene products. To effectively discriminate between tumor sub-types, reliable laboratory techniques and analytic methods are required. RESULTS: We derived mRNA expression profiles from 21 human tissue samples (eight normal kidneys and 13 kidney tumors) and two pooled samples using the Affymetrix GeneChip platform. A panel of ten clustering algorithms combined with four data pre-processing methods identified a consensus cluster dendrogram in 18 of 40 analyses and of these 16 used a logarithmic transformation. Within the consensus dendrogram the expression profiles of the samples grouped according to tissue type; clear cell and chromophobe carcinomas displayed distinctly different gene expression patterns. By using a rigorous statistical selection based method we identified 355 genes that showed significant (p < 0.001) gene expression changes in clear cell renal carcinomas compared to normal kidney. These genes were classified with a tool to conceptualize expression patterns called "Functional Taxonomy". Each tumor type had a distinct "signature," with a high number of genes in the categories of Metabolism, Signal Transduction, and Cellular and Matrix Organization and Adhesion. CONCLUSIONS: Affymetrix GeneChip profiling differentiated clear cell and chromophobe carcinomas from one another and from normal kidney cortex. Clustering methods that used logarithmic transformation of data sets produced dendrograms consistent with the sample biology. Functional taxonomy provided a practical approach to the interpretation of gene expression data.

Adenocarcinoma, Clear Cell↗

A simple clustering technique to improve QSAR model selection and predictivity: application to a receptor independent 4D-QSAR analysis of cyclic urea derived inhibitors of HIV-1 protease.

A training set of 50 tetrahydropyrimidine-2-one based inhibitors of HIV-1 protease, for which the -log K(i) values were measured, was used to construct receptor independent 4D-QSAR models. A novel clustering technique was employed to facilitate and improve model selection as well as test set predictions. Following the manifold model theory, five unique models were chosen by the clustering algorithm (q(2) = 0.81-0.84). The models were used to map the atom type morphology of the inhibitor binding site of HIV-1 protease as well as to predict the potencies (-log K(i)) of 10 test set compounds. The rank-difference correlation coefficient was used to evaluate the quality of the test set predictions, which was improved from 0.39 to 0.68 when the clustering technique was applied. The set of five models, collectively, identify the important binding characteristics of the HIV protease receptor site. This study demonstrates that the selected simple clustering technique provides a discrete algorithm for model selection, as well as improving the quality of test set, or unknown, compound prediction as determined by the rank-difference correlation coefficient.

Algorithms↗

Identification of aperiodic seasonality in non-Gaussian time series.

Time series that arise from biological experimentation can exhibit seasonality where the lengths of the seasons may vary. In addition, such time series may not be stationary with respect to either mean, variance, or autocorrelation, thus making the usual waveform-fitting techniques inappropriate. An agglomerative clustering algorithm for identifying seasons in such series is proposed, consisting of an initialization step, iterative steps where clusters are combined into larger clusters, and a stopping rule for the iteration. The clusters can be associated with seasons or phases, and biological cycles can be identified from the phases. Results of a simulation and an analysis of luteinizing hormone concentrations are presented.

Algorithms↗

Selection of a representative set of structures from Brookhaven Protein Data Bank.

Reliable structural and statistical analyses of three dimensional protein structures should be based on unbiased data. The Protein Data Bank is highly redundant, containing several entries for identical or very similar sequences. A technique was developed for clustering the known structures based on their sequences and contents of alpha- and beta-structures. First, sequences were aligned pairwise. A representative sample of sequences was then obtained by grouping similar sequences together, and selecting a typical representative from each group. The similarity significance threshold needed in the clustering method was found by analyzing similarities of random sequences. Because three dimensional structures for proteins of same structural class are generally more conserved than their sequences, the proteins were clustered also according to their contents of secondary structural elements. The results of these clusterings indicate conservation of alpha- and beta-structures even when sequence similarity is relatively low. An unbiased sample of 103 high resolution structures, representing a wide variety of proteins, was chosen based on the suggestions made by the clustering algorithm. The proteins were divided into structural classes according to their contents and ratios of secondary structural elements. Previous classifications have suffered from subjective view of secondary structures, whereas here the classification was based on backbone geometry. The concise view lead to reclassification of some structures. The representative set of structures facilitates unbiased analyses of relationships between protein sequence, function, and structure as well as of structural characteristics.

Algorithms↗

Identifying spatial relationships in neural processing using a multiple classification approach.

The application of statistical classification methods to in vivo functional neuroimaging data makes it possible to explore spatial patterns in task-related changes in neural processing. Cluster analysis is one group of descriptive statistical procedures that can assist in identifying classes of brain regions that exhibit similar task-related functionality. In practice, a limitation of cluster analysis is that the performances of clustering algorithms rely on unknown characteristics of the data, making it difficult to determine which procedure best suits a particular analysis. We present a multiple classification approach that incorporates numerous algorithms, evaluates the associated classifications, and either selects a plausible partition relative to the others considered or pools the results from the numerous methods. The multiple classification approach utilizes a new performance criterion, called the relative information (RI) measure, to evaluate the quality of the candidate partitions and as the basis for producing a composite classification image. Employing multiple classifications, rather than a single algorithm, our methodology increases the chance of detecting the functional relationships within the data and, therefore, produces more reliable results. We apply our methodology to a PET study to explore spatial relationships in measured brain function associated with increasing blood alcohol concentration levels, and we perform a simulation study to evaluate the performance of RI.

Alcoholic Intoxication↗

Temporal pattern of source activities evoked by different types of motion onset stimuli.

The aim of this study was to compare the time course of motion-related source activities evoked by the onset of different kinds of visual motion stimuli in human subjects. Event-related potentials (ERP) were recorded from 64 scalp electrodes in ten healthy subjects while they were viewing four different types of motion stimuli (translation, rotation, expansion and contraction). Following a new approach combining a current density reconstruction with clustering algorithms, source maxima in the time range from 50 to 400 ms after the onset of the visual stimulus were localized and the time courses of activation were elaborated. Six regions contributed significantly to source activity, half originating in the occipital lobe and half in the right parietal and right temporal cortex. The comparison of their time courses led to the following conclusions: (i) the different kinds of motion stimuli activated about the same areas of the brain but with different temporal patterns. (ii) Mainly parietal and extrastriate areas, but not V1/V2, were significantly involved in the differentiation of different kinds of motion. (iii) Contrasting the different kinds of motion onsets, responses from parietal areas were found mainly before those from lateral occipital areas. (iv) The classically defined N2 and P2 components were significantly different among the four motion conditions, but not P1. The N2 motion-related component was elicited not only by lateral occipital areas and middle temporal areas but also by right parietal areas. (v) The rotation condition evoked a novel component P180, concomitant with an increased activity in the left middle temporal gyrus.

Adult↗

A novel means of using gene clusters in a two-step empirical Bayes method for predicting classes of samples.

MOTIVATION: The classification of samples using gene expression profiles is an important application in areas such as cancer research and environmental health studies. However, the classification is usually based on a small number of samples, and each sample is a long vector of thousands of gene expression levels. An important issue in parametric modeling for so many gene expression levels is the control of the number of nuisance parameters in the model. Large models often lead to intensive or even intractable computation, while small models may be inadequate for complex data. METHODOLOGY: We propose a two-step empirical Bayes classification method as a solution to this issue. At the first step, we use the model-based cluster algorithm with a non-traditional purpose of assigning gene expression levels to form abundance groups. At the second step, by assuming the same variance for all the genes in the same group, we substantially reduce the number of nuisance parameters in our statistical model. RESULTS: The proposed model is more parsimonious, which leads to efficient computation under an empirical Bayes estimation procedure. We consider two real examples and simulate data using our method. Desired low classification error rates are obtained even when a large number of genes are pre-selected for class prediction.

Algorithms↗

Associative clustering for exploring dependencies between functional genomics data sets.

High-throughput genomic measurements, interpreted as cooccurring data samples from multiple sources, open up a fresh problem for machine learning: What is in common in the different data sets, that is, what kind of statistical dependencies are there between the paired samples from the different sets? We introduce a clustering algorithm for exploring the dependencies. Samples within each data set are grouped such that the dependencies between groups of different sets capture as much of pairwise dependencies between the samples as possible. We formalize this problem in a novel probabilistic way, as optimization of a Bayes factor. The method is applied to reveal commonalities and exceptions in gene expression between organisms and to suggest regulatory interactions in the form of dependencies between gene expression profiles and regulator binding patterns.

Algorithms↗

Unsupervised multistage image classification using hierarchical clustering with a Bayesian similarity measure.

A new multistage method using hierarchical clustering for unsupervised image classification is presented. In the first phase, the multistage method performs segmentation using a hierarchical clustering procedure which confines merging to spatially adjacent clusters and generates an image partition such that no union of any neighboring segments has homogeneous intensity values. In the second phase, the segments resulting from the first stage are classified into a small number of distinct states by a sequential merging operation. The region-merging procedure in the first phase makes use of spatial contextual information by characterizing the geophysical connectedness of a digital image structure with a Markov random field, while the second phase employs a context-free similarity measure in the clustering process. The segmentation procedure of region merging is implemented as a hierarchical clustering algorithm whereby a multiwindow approach using a pyramid-like structure is employed to increase computational efficiency while maintaining spatial connectivity in merging. From experiments with both simulated and remotely sensed data, the proposed method was determined to be quite effective for unsupervised analysis. In particular, the region-merging approach based on spatial contextual information was shown to provide more accurate classification of images with smooth spatial patterns.

Algorithms↗

An information theoretic approach for analyzing temporal patterns of gene expression.

MOTIVATION: Arrays allow measurements of the expression levels of thousands of mRNAs to be made simultaneously. The resulting data sets are information rich but require extensive mining to enhance their usefulness. Information theoretic methods are capable of assessing similarities and dissimilarities between data distributions and may be suited to the analysis of gene expression experiments. The purpose of this study was to investigate information theoretic data mining approaches to discover temporal patterns of gene expression from array-derived gene expression data. RESULTS: The Kullback-Leibler divergence, an information-theoretic distance that measures the relative dissimilarity between two data distribution profiles, was used in conjunction with an unsupervised self-organizing map algorithm. Two published, array-derived gene expression data sets were analyzed. The patterns obtained with the KL clustering method were found to be superior to those obtained with the hierarchical clustering algorithm using the Pearson correlation distance measure. The biological significance of the results was also examined. AVAILABILITY: Software code is available by request from the authors. All programs were written in ANSI C and Matlab (Mathworks Inc., Natick, MA).

Algorithms↗

Bayesian algorithms for simultaneous structure from motion estimation of multiple independently moving objects.

In this paper, the problem of simultaneous structure from motion estimation for multiple independently moving objects from a monocular image sequence is addressed. Two Bayesian algorithms are presented for solving this problem using the sequential importance sampling (SIS) technique. The empirical posterior distribution of object motion and feature separation parameters is approximated by weighted samples. The first algorithm addresses the problem when only two moving objects are present. A singular value decomposition (SVD)-based sample clustering algorithm is shown to be capable of separating samples related to different objects. A pair of SIS procedures is used to track the posterior distribution of the motion parameters. In the second algorithm, a balancing step is added into the SIS procedure to preserve samples of low weights so that all objects have enough samples to propagate empirical motion distributions. By using the proposed algorithms, the relative motions of all the moving objects with respect to the camera can be simultaneously estimated. Both algorithms have been tested on synthetic and real-image sequences. Improved results have been achieved.

Algorithms↗

In quest of an empirical potential for protein structure prediction.

Key to successful protein structure prediction is a potential that recognizes the native state from misfolded structures. Recent advances in empirical potentials based on known protein structures include improved reference states for assessing random interactions, sidechain-orientation-dependent pair potentials, potentials for describing secondary or supersecondary structural preferences and, most importantly, optimization protocols that sculpt the energy landscape to enhance the correlation between native-like features and the energy. Improved clustering algorithms that select native-like structures on the basis of cluster density also resulted in greater prediction accuracy. For template-based modeling, these advances allowed improvement in predicted structures relative to their initial template alignments over a wide range of target-template homology. This represents significant progress and suggests applications to proteome-scale structure prediction.

Algorithms↗

A quantitative comparison of functional MRI cluster analysis.

The aim of this work is to compare the efficiency and power of several cluster analysis techniques on fully artificial (mathematical) and synthesized (hybrid) functional magnetic resonance imaging (fMRI) data sets. The clustering algorithms used are hierarchical, crisp (neural gas, self-organizing maps, hard competitive learning, k-means, maximin-distance, CLARA) and fuzzy (c-means, fuzzy competitive learning). To compare these methods we use two performance measures, namely the correlation coefficient and the weighted Jaccard coefficient (wJC). Both performance coefficients (PCs) clearly show that the neural gas and the k-means algorithm perform significantly better than all the other methods using our setup. For the hierarchical methods the ward linkage algorithm performs best under our simulation design. In conclusion, the neural gas method seems to be the best choice for fMRI cluster analysis, given its correct classification of activated pixels (true positives (TPs)) whilst minimizing the misclassification of inactivated pixels (false positives (FPs)), and in the stability of the results achieved.

Algorithms↗

Mining the NCI anticancer drug discovery databases: genetic function approximation for the QSAR study of anticancer ellipticine analogues.

The U.S. National Cancer Institute (NCI) conducts a drug discovery program in which approximately 10,000 compounds are screened every year in vitro against a panel of 60 human cancer cell lines from different organs of origin. Since 1990, approximately 63,000 compounds have been tested, and their patterns of activity profiled. Recently, we analyzed the antitumor activity patterns of 112 ellipticine analogues using a hierarchical clustering algorithm. Dramatic coherence between molecular structures and activity patterns was observed qualitatively from the cluster tree. In the present study, we further investigate the quantitative structure-activity relationships (QSAR) of these compounds, in particular with respect to the influence of p53-status and the CNS cell selectivity of the activity patterns. Independent variables (i.e., chemical structural descriptors of the ellipticine analogues) were calculated from the Cerius2 molecular modeling package. Important structural descriptors, including partial atomic charges on the ellipticine ring-forming atoms, were identified by the recently developed genetic function approximation (GFA) method. For our data set, the GFA method gave better correlation and cross-validation results (R2 and CVR2 were usually approximately 0.3 higher) than did classical stepwise linear regression. A procedure for improving the performance of GFA is proposed, and the relative advantages and disadvantages of using GFA for QSAR studies are discussed.

Algorithms↗

Impact of plasmids and genetic change on the numerical classification of staphylococci.

Newly isolated bacterial strains often contain extrachromosomal DNA as plasmid DNA. These accessory components of the DNA gene pool confer additional phenotypic properties on their host but, despite this, little attention has been paid to the impact of plasmid-mediated characters on bacterial classification. In the present study, the effect of antibiotic resistance plasmids on the classification of representative staphylococci was determined using numerical phenetic techniques. Over sixty percent of the eighty-one test strains contained one or more plasmids which varied in molecular weight from 1.4 to 36 Mdal. Antibiotic resistance phenotypes were eliminated from strains of S. aureus, S. chromogenes, S. cohnii, S. hyicus and S. xylosus, and from a laboratory isolate, to give sixteen derivative strains. Fourteen had lost one or more plasmids and two had deleted plasmids. In addition three further derivative strains were isolated which showed no plasmid loss but exhibited gross phenotypic changes. The test and derivative strains were the subject of numerical phenetic analyses based on seventy-eight unit characters. Data were examined using the simple matching, Jaccard and pattern coefficients and clustering achieved using the unweighted pair group method with arithmetic averages algorithm. Cluster composition was not markedly affected by the statistics used or by test error, estimated at 1.02%. Numerically circumscribed clusters and subclusters were equated with the established species S. aureus, S. chromogenes, S. cohnii, S. hyicus, S. lentus, S. intermedius, S. sciuri and S. xylosus. The sixteen derivative strains with either lost or delected plasmids were recovered in the same cluster or subcluster as their corresponding parent indicating that the removal of plasmid-expressed characters had little effect on the structure of the numerical classification. In contrast, two of the three strains of S. xylosus with genomically-derived phenotypic variation formed a cluster that separated from their parent strain at the 70% similarity level in the SSM, UPGMA analysis.

Animals↗