PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

New findings on alternative criteria for PTSD in preschool children.

OBJECTIVE: An alternative set of criteria for posttraumatic stress disorder (PTSD) for preschool children was analyzed for validity. METHOD: Sixty-two traumatized children and 63 healthy controls, aged 20 months through 6 years, were assessed. The traumatic experiences included motor vehicle collisions, accidental injuries, abuse, and witnessing violence. The number of symptoms required for clusters C and D and the utility of proposed symptoms were systematically analyzed. RESULTS: No cases met the DSM-IV algorithm for PTSD. Cluster B was endorsed 67.9% of the time. The proportion of cases meeting the cluster C threshold was 2% when three symptoms were required, 11% when two symptoms were required, and 39% when one symptom was required. The rate of cluster D was 45% when two symptoms were required and 73% when one symptom was required. Four novel symptoms did not substantially add to the diagnostic validity of the criteria. The optimal algorithm (one cluster B symptom, one cluster C symptom, and two cluster D symptoms) diagnosed PTSD at a rate of 26%. Measures of comorbid symptoms concurrently provided convergent validation to support this revised algorithm. CONCLUSION: Revisions to the DSM-IV PTSD criteria continue to be supported so that highly symptomatic young children can be diagnosed.

Algorithms↗

K-means clustering: a half-century synthesis.

This paper synthesizes the results, methodology, and research conducted concerning the K-means clustering method over the last fifty years. The K-means method is first introduced, various formulations of the minimum variance loss function and alternative loss functions within the same class are outlined, and different methods of choosing the number of clusters and initialization, variable preprocessing, and data reduction schemes are discussed. Theoretic statistical results are provided and various extensions of K-means using different metrics or modifications of the original algorithm are given, leading to a unifying treatment of K-means and some of its extensions. Finally, several future studies are outlined that could enhance the understanding of numerous subtleties affecting the performance of the K-means method.

Algorithms↗

Hybridization of evolutionary algorithms and local search by means of a clustering method.

This paper presents a hybrid evolutionary algorithm (EA) to solve nonlinear-regression problems. Although EAs have proven their ability to explore large search spaces, they are comparatively inefficient in fine tuning the solution. This drawback is usually avoided by means of local optimization algorithms that are applied to the individuals of the population. The algorithms that use local optimization procedures are usually called hybrid algorithms. On the other hand, it is well known that the clustering process enables the creation of groups (clusters) with mutually close points that hopefully correspond to relevant regions of attraction. Local-search procedures can then be started once in every such region. This paper proposes the combination of an EA, a clustering process, and a local-search procedure to the evolutionary design of product-units neural networks. In the methodology presented, only a few individuals are subject to local optimization. Moreover, the local optimization algorithm is only applied at specific stages of the evolutionary process. Our results show a favorable performance when the regression method proposed is compared to other standard methods.

Algorithms↗

Quantitative analysis of cell columns in the cerebral cortex.

We present a quantified imaging method that describes the cell column in mammalian cortex. The minicolumn is an ideal template with which to examine cortical organization because it is a basic unit of function, complete in itself, which interacts with adjacent and distance columns to form more complex levels of organization. The subtle details of columnar anatomy should reflect physiological changes that have occurred in evolution as well as those that might be caused by pathologies in the brain. In this semiautomatic method, images of Nissl-stained tissue are digitized or scanned into a computer imaging system. The software detects the presence of cell columns and describes details of their morphology and of the surrounding space. Columns are detected automatically on the basis of cell-poor and cell-rich areas using a Gaussian distribution. A line is fit to the cell centers by least squares analysis. The line becomes the center of the column from which the precise location of every cell can be measured. On this basis several algorithms describe the distribution of cells from the center line and in relation to the available surrounding space. Other algorithms use cluster analyses to determine the spatial orientation of every column.

Algorithms↗

Assessing antibiotic resistance in fecal Escherichia coli in young calves using cluster analysis techniques.

This study uses cluster analysis techniques to describe the antibiotic susceptibility patterns seen in calf fecal Escherichia coli (E. coli). Cohorts of 30 dairy calves at six farms were sampled at 2-week intervals during the pre-weaning period. At each sampling occasion five fecal E. coli isolates per calf were analyzed for antibiotic susceptibility to 12 antibiotics using the disk diffusion method. All isolates had a profile consisting of the aggregate measured inhibition zone size for each of the evaluated antibiotics. Several cluster analytic algorithms were assessed to partition the E. coli isolates. For our data, Ward's minimum variance method met the objectives of the study. Relative to the number of possible combinations of resistance clusters, a parsimonious set of 14 patterns was developed. This set of E. coli isolates exhibited a limited set of resistance patterns to the different antibiotics indicating that certain resistance genes may be linked.

Animals↗

SaRAD: a Simple and Robust Abbreviation Dictionary.

MOTIVATION: Due to recent interest in the use of textual material to augment traditional experiments it has become necessary to automatically cluster, classify and filter natural language information. RESULTS: The Simple and Robust Abbreviation Dictionary (SaRAD) provides an easy to implement, high performance tool for the construction of a biomedical symbol dictionary. The algorithms, applied to the MEDLINE document set, result in a high quality dictionary and toolset to disambiguate abbreviation symbols automatically.

Abbreviations as Topic↗

Antigen microarray profiling of autoantibodies in rheumatoid arthritis.

OBJECTIVE: Because rheumatoid arthritis (RA) is a heterogeneous autoimmune disease in terms of disease manifestations, clinical outcomes, and therapeutic responses, we developed and applied a novel antigen microarray technology to identify distinct serum antibody profiles in patients with RA. METHODS: Synovial proteome microarrays, containing 225 peptides and proteins that represent candidate and control antigens, were developed. These arrays were used to profile autoantibodies in randomly selected sera from 2 different cohorts of patients: the Stanford Arthritis Center inception cohort, comprising 18 patients with established RA and 38 controls, and the Arthritis, Rheumatism, and Aging Medical Information System cohort, comprising 58 patients with a clinical diagnosis of RA of <6 months duration. Data were analyzed using the significance analysis of microarrays algorithm, the prediction analysis of microarrays algorithm, and Cluster software. RESULTS: Antigen microarrays demonstrated that autoreactive B cell responses targeting citrullinated epitopes were present in a subset of patients with early RA with features predictive of the development of severe RA. In contrast, autoimmune targeting of the native epitopes contained on synovial arrays, including several human cartilage gp39 peptides and type II collagen, were associated with features predictive of less severe RA. CONCLUSION: Proteomic analysis of autoantibody reactivities provides diagnostic information and allows stratification of patients with early RA into clinically relevant disease subsets.

Algorithms↗

MCMC methods for putative pollution source problems in environmental epidemiology.

This paper demonstrates the use of the Gibbs Sampler and other Markov Chain Monte Carlo (MCMC) methods in two applications in environmental epidemiology. The first example concerns the application of a Metropolis-Hastings/Gibbs sampler to a Cox process with a direction-dependent cluster variance parameter. The second example consists of the estimation of the posterior (spatial) distribution of a putative location.

Algorithms↗

Bivariate frailty model for the analysis of multivariate survival time.

Because of limitations of the univariate frailty model in analysis of multivariate survival data, a bivariate frailty model is introduced for the analysis of bivariate survival data. This provides tremendous flexibility especially in allowing negative associations between subjects within the same cluster. The approach involves incorporating into the model two possibly correlated frailties for each cluster. The bivariate lognormal distribution is used as the frailty distribution. The model is then generalized to multivariate survival data with two distinguished groups and also to alternating process data. A modified EM algorithm is developed with no requirement of specification of the baseline hazards. The estimators are generalized maximum likelihood estimators with subject-specific interpretation. The model is applied to a mental health study on evaluation of health policy effects for inpatient psychiatric care.

Algorithms↗

A multiple imputation approach to linear regression with clustered censored data.

We extend Wei and Tanner's (1991) multiple imputation approach in semi-parametric linear regression for univariate censored data to clustered censored data. The main idea is to iterate the following two steps: 1) using the data augmentation to impute for censored failure times; 2) fitting a linear model with imputed complete data, which takes into consideration of clustering among failure times. In particular, we propose using the generalized estimating equations (GEE) or a linear mixed-effects model to implement the second step. Through simulation studies our proposal compares favorably to the independence approach (Lee et al., 1993), which ignores the within-cluster correlation in estimating the regression coefficient. Our proposal is easy to implement by using existing softwares.

Algorithms↗

Toward an empirically derived typology of obese persons.

The MMPI, medical, anthropomorphic, and laboratory evaluations were completed by 260 obese patients (211 females, 49 males) at a New York hospital. Biological and psychological variables were separately subjected to principal components analyses. Fifteen biological and five psychological components were extracted. A three-cluster solution was selected from a K-means clustering on biological components, replicated via Ward's method, and validated via a discriminant analysis on psychological components. Cluster 1, 'android obesity', contained 75 percent of the males and was characterized by 'masculine phenotypy', 'poor conditioning' and 'adverse serum lipids', and less 'feminine' responding on the MMPI. Cluster 2, 'gynoid obesity', was low on components measuring physical stress and masculine phenotypy, was 95 percent female, moderately obese compared to clusters 1 and 3, and had a relatively healthy profile. Cluster 3 had elevations on overall fatness and physiological and psychological stress, and low scores on a 'healthy blood synthesis' component. This cluster, labeled 'morbidly obese', was the most obese and had profiles suggesting adverse effects of obesity.

Adult↗

Clustering proteins from interaction networks for the prediction of cellular functions.

BACKGROUND: Developing reliable and efficient strategies allowing to infer a function to yet uncharacterized proteins based on interaction networks is of crucial interest in the current context of high-throughput data generation. In this paper, we develop a new algorithm for clustering vertices of a protein-protein interaction network using a density function, providing disjoint classes. RESULTS: Applied to the yeast interaction network, the classes obtained appear to be biological significant. The partitions are then used to make functional predictions for uncharacterized yeast proteins, using an annotation procedure that takes into account the binary interactions between proteins inside the classes. We show that this procedure is able to enhance the performances with respect to previous approaches. Finally, we propose a new annotation for 37 previously uncharacterized yeast proteins. CONCLUSION: We believe that our results represent a significant improvement for the inference of cellular functions, that can be applied to other organism as well as to other type of interaction graph, such as genetic interactions.

Cluster Analysis↗

Calculating similarities between biological activities in the MDL Drug Data Report database.

There are a number of licensed databases that assign biological activities to druglike compounds. The MDL Drug Data Report (MDDR), compiled from the patent literature, is a popular example. It contains several hundred distinct activities, some of which are therapeutic areas (e.g., Antihypertensive) and some of which are related to specific enzymes or receptors (e.g., ACE inhibitor). There are several data mining applications where it would be useful to calculate a similarity between any two activities. Two distinct activity labels can have a significant similarity for a number of reasons: two activities can be nearly synonymous (e.g., CCK B antagonist vs Gastrin antagonist), one activity may be a subset of another (e.g., Dopamine (D2) agonist vs Dopamine agonist), or an activity can be the mechanism by which another activity works (e.g., ACE inhibitor vs Antihypertensive), etc. In an ideal world, similarities for two activities could be calculated simply by comparing the compounds they have in common, but in hand-curated databases such as the MDDR the assignment of activities to compounds are inevitably inconsistent and incomplete. We propose a number of methods of calculating activity-activity similarities that hopefully compensate for errors in hand-curation. Two of these, TIMI and trend vector, show promise. Soft clustering of the activities using a union of similarity methods shows a reasonable association of therapeutic areas with their mechanisms.

Algorithms↗

Two-dimensional gray-scale clustering for texture analysis.

OBJECTIVES: To develop a new quantitative method for the visual discrimination of image texture. METHODS: Two kinds of image phantoms were prepared, one for evaluating the effects of change in size and gray values of individual pixels (primitives) on perceived coarseness and the other for evaluating changes in groups of pixels (clusters) on perceived heterogeneity. The phantom images were displayed on a CRT and presented to 11 observers who assessed heterogeneity and coarseness on a 10-point scale between -5 and +5. On the basis of the observers' results, a new texture analysis method termed two-dimensional gray-scale clustering analysis was developed and applied to measure quantitatively the texture of the phantoms. The results obtained were then compared with those of the visual evaluation. RESULTS: The size of the primitives and the clusters greatly affected the visual evaluation of heterogeneity and coarseness. Changes in the gray value had only a slight effect. The intra-observer variation for heterogeneity was significantly larger than that for coarseness. Two-dimensional gray-scale clustering analysis could differentiate heterogeneity from coarseness. A high correlation was obtained between the visual evaluation and the quantitative data. CONCLUSION: Quantitative two-dimensional gray-scale clustering analysis appears to be a useful means of texture analysis.

Algorithms↗

Bagging to improve the accuracy of a clustering procedure.

MOTIVATION: The microarray technology is increasingly being applied in biological and medical research to address a wide range of problems such as the classification of tumors. An important statistical question associated with tumor classification is the identification of new tumor classes using gene expression profiles. Essential aspects of this clustering problem include identifying accurate partitions of the tumor samples into clusters and assessing the confidence of cluster assignments for individual samples. RESULTS: Two new resampling methods, inspired from bagging in prediction, are proposed to improve and assess the accuracy of a given clustering procedure. In these ensemble methods, a partitioning clustering procedure is applied to bootstrap learning sets and the resulting multiple partitions are combined by voting or the creation of a new dissimilarity matrix. As in prediction, the motivation behind bagging is to reduce variability in the partitioning results via averaging. The performances of the new and existing methods were compared using simulated data and gene expression data from two recently published cancer microarray studies. The bagged clustering procedures were in general at least as accurate and often substantially more accurate than a single application of the partitioning clustering procedure. A valuable by-product of bagged clustering are the cluster votes which can be used to assess the confidence of cluster assignments for individual observations. SUPPLEMENTARY INFORMATION: For supplementary information on datasets, analyses, and software, consult http://www.stat.berkeley.edu/~sandrine and http://www.bioconductor.org.

Algorithms↗

[A new method for EST clustering].

We developed an EST (expressed sequence tag) clustering method, ESTClustering, to generate high-quality unique expressed sequence based on large-scale EST sequencing. The method uses consensus sequences to sequence analyze with megablast and assemble each cluster with phrap in clustering process. The clustering strategy can efficiently identify gene family and alternate splicing forms of expressed sequences. It can also reduce the adverse effects caused by sequence errors. The ESTClustering method tends to provide more expressed gene forms comparing with the UniGene clustering method of the National Center for Biotechnology Information. Analysis of the 112,256 ESTs of Arabidopsis with ESTClustering produced 23,581 EST clusters. Among these Arabidopsis EST clusters, 13,597 have corresponding genome coding sequences and this number is close to the number of genes predicted with Arabidopsis ESTs. Using this clustering method, a total of 147,191 rice ESTs were clustered into 33,896 groups.

Algorithms↗

Spatial partitioning using multivariate cluster analysis and a contiguity algorithm.

Spatial analysis of epidemiological data can be a useful tool for identifying patterns of disease occurrence and can provide substantial support for prevention and control strategies. To obtain the greatest spatial resolution, it is important to use the smallest available areal units with homogeneous population. However, small areas usually have a small population, introducing spurious variability in the chosen indicators of disease occurrence. This paper describes an approach for combining small geographical units to stabilize mortality rates by pooling information across areas according to specified risk profiles. The procedure is based on a principal component analysis, followed by a cluster analysis of social-economic indicators to classify the risk profile of each small area. The classification is used in an algorithm to join neighbouring areas with similar profiles until an estimated population size is achieved. We applied this method to two Administrative Regions of the city of Rio de Janeiro, Brazil, using the census tracts as the basic areal unit. Census tracts were classified according to four socioeconomic categories distributed spatially as a mosaic, where tracts of differing categories neighbour each other. The aggregation algorithm produced a new partition of the region studied, with the created areal units preserving the internal socioeconomic homogeneity.

Adult↗

Spectral clustering of protein sequences.

An important problem in genomics is automatically clustering homologous proteins when only sequence information is available. Most methods for clustering proteins are local, and are based on simply thresholding a measure related to sequence distance. We first show how locality limits the performance of such methods by analysing the distribution of distances between protein sequences. We then present a global method based on spectral clustering and provide theoretical justification of why it will have a remarkable improvement over local methods. We extensively tested our method and compared its performance with other local methods on several subsets of the SCOP (Structural Classification of Proteins) database, a gold standard for protein structure classification. We consistently observed that, the number of clusters that we obtain for a given set of proteins is close to the number of superfamilies in that set; there are fewer singletons; and the method correctly groups most remote homologs. In our experiments, the quality of the clusters as quantified by a measure that combines sensitivity and specificity was consistently better [on average, improvements were 84% over hierarchical clustering, 34% over Connected Component Analysis (CCA) (similar to GeneRAGE) and 72% over another global method, TribeMCL].

Algorithms↗