PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Performance of multi-layer feedforward neural networks to predict liver transplantation outcome.

A novel multisolutional clustering and quantization (MCQ) algorithm has been developed that provides a flexible way to preprocess data. It was tested whether it would impact the neural network's performance favorably and whether the employment of the proposed algorithm would enable neural networks to handle missing data. This was assessed by comparing the performance of neural networks using a well-documented data set to predict outcome following liver transplantation. This new approach to data preprocessing leads to a statistically significant improvement in network performance when compared to simple linear scaling. The obtained results also showed that coding missing data as zeroes in combination with the MCQ algorithm, leads to a significant improvement in neural network performance on a data set containing missing values in 59.4% of cases when compared to replacement of missing values with either series means or medians.

Algorithms↗

Methods for robust clustering of epileptic EEG spikes.

We investigate algorithms for clustering of epileptic electroencephalogram (EEG) spikes. Such a method is useful prior to averaging and inverse computations since the spikes of a patient often belong to a few distinct classes. Data sets often contain outliers, which makes algorithms with robust performance desirable. We compare the fuzzy C-means (FCM) algorithm and a graph-theoretic algorithm. We give criteria for determination of the correct level of outlier contamination. The performance is then studied by aid of simulations, which show good results for a range of circumstances, for both algorithms. The graph-theoretic method gave better results than FCM for simulated signals. Also, when evaluating the methods on seven real-life data sets, the graph-theoretic method was the better method, in terms of closeness to the manual assessment by a neurophysiologist. However, there was some discrepancy between manual and automatic clustering and we suggest as an alternative method a human choice among a limited set of automatically obtained clusterings. Furthermore, we evaluate geometrically weighted feature extraction and conclude that it is useful as a supplementary dimension for clustering.

Algorithms↗

A modified K-means algorithm for circular invariant clustering.

Several important pattern recognition applications are based on feature vector extraction and vector clustering. Directional patterns are commonly represented by rotation-variant vectors Fd formed from features uniformly extracted in M directions. It is often desirable that pattern recognition algorithms are invariant under pattern rotation. This paper introduces a distance measure and a K-means-based algorithm, namely, Circular K-means (CK-means) to cluster vectors containing directional information, such as Fd, in a circular-shift invariant manner. A circular shift of Fd corresponds to pattern rotation, thus, the algorithm is rotation invariant. An efficient Fourier domain representation of the proposed measure is presented to reduce computational complexity. A split and merge approach (SMCK-means), suited to the proposed CK-means technique, is proposed to reduce the possibility of converging at local minima and to estimate the correct number of clusters. Experiments performed for textural images illustrate the superior performance of the proposed algorithm for clustering directional vectors Fd, compared to the alternative approach that uses the original K-means and rotation-invariant feature vectors transformed from Fd.

Algorithms↗

Multiple sequence alignment with hierarchical clustering.

An algorithm is presented for the multiple alignment of sequences, either proteins or nucleic acids, that is both accurate and easy to use on microcomputers. The approach is based on the conventional dynamic-programming method of pairwise alignment. Initially, a hierarchical clustering of the sequences is performed using the matrix of the pairwise alignment scores. The closest sequences are aligned creating groups of aligned sequences. Then close groups are aligned until all sequences are aligned in one group. The pairwise alignments included in the multiple alignment form a new matrix that is used to produce a hierarchical clustering. If it is different from the first one, iteration of the process can be performed. The method is illustrated by an example: a global alignment of 39 sequences of cytochrome c.

Algorithms↗

Many-body interaction analysis: algorithm development and application to large molecular clusters.

A completely automated algorithm for performing many-body interaction energy analysis of clusters (MBAC) [M. J. Elrodt and R. J. Saykally, Chem. Rev. 94, 1975 (1994); S. S. Xantheas, J. Chem. Phys. 104, 8821 (1996)] at restricted Hartree-Fock (RHF)/MA Plesset 2nd order perturbation theory (MP2)/density functional theory (DFT) level of theory is reported. Use of superior guess density matrices (DM's) for smaller fragments generated from DM of the parent system and elimination of energetically insignificant higher-body combinations, leads to a more efficient performance (speed-up up to 2) compared to the conventional procedure. MBAC approach has been tested out on several large-sized weakly bound molecular clusters such as (H(2)O)(n), n=8, 12, 16, 20 and hydrated clusters of amides and aldehydes. The MBAC results indicate that the amides interact more strongly with water than aldehydes in these clusters. It also reconfirms minimization of the basis set superposition error for large cluster on using superior quality basis set. In case of larger weakly bound clusters, the contributions higher than four body are found to be repulsive in nature and smaller in magnitude. The reason for this may be attributed to the increased random orientations of the interacting molecules separated from each other by large distances.

Journal Article↗

d2_cluster: a validated method for clustering EST and full-length cDNAsequences.

Several efforts are under way to condense single-read expressed sequence tags (ESTs) and full-length transcript data on a large scale by means of clustering or assembly. One goal of these projects is the construction of gene indices where transcripts are partitioned into index classes (or clusters) such that they are put into the same index class if and only if they represent the same gene. Accurate gene indexing facilitates gene expression studies and inexpensive and early partial gene sequence discovery through the assembly of ESTs that are derived from genes that have yet to be positionally cloned or obtained directly through genomic sequencing. We describe d2_cluster, an agglomerative algorithm for rapidly and accurately partitioning transcript databases into index classes by clustering sequences according to minimal linkage or "transitive closure" rules. We then evaluate the relative efficiency of d2_cluster with respect to other clustering tools. UniGene is chosen for comparison because of its high quality and wide acceptance. It is shown that although d2_cluster and UniGene produce results that are between 83% and 90% identical, the joining rate of d2_cluster is between 8% and 20% greater than UniGene. Finally, we present the first published rigorous evaluation of under and over clustering (in other words, of type I and type II errors) of a sequence clustering algorithm, although the existence of highly identical gene paralogs means that care must be taken in the interpretation of the type II error. Upper bounds for these d2_cluster error rates are estimated at 0.4% and 0.8%, respectively. In other words, the sensitivity and selectivity of d2_cluster are estimated to be >99.6% and 99.2%.

Algorithms↗

Evolutionary-based grouping of haplotypes in association analysis.

Haplotypes incorporate more information about the underlying polymorphisms than do genotypes for individual SNPs, and are considered as a more informative format of data in association analysis. To model haplotypes requires high degrees of freedom, which could decrease power and limit a model's capacity to incorporate other complex effects, such as gene-gene interactions. Even within haplotype blocks, high degrees of freedom are still a concern unless one chooses to discard rare haplotypes. To increase the efficiency and power of haplotype analysis, we adapt the evolutionary concepts of cladistic analyses and propose a grouping algorithm to cluster rare haplotypes to the corresponding ancestral haplotypes. The algorithm determines the cluster bases by preserving common haplotypes using a criterion built on the Shannon information content. Each haplotype is then assigned to its appropriate clusters probabilistically according to the cladistic relationship. Through this algorithm, we perform association analysis based on groups of haplotypes. Simulation results indicate power increases for performing tests on the haplotype clusters when compared to tests using original haplotypes or the truncated haplotype distribution.

Algorithms↗

Strategies for increasing the efficiency of a genetic algorithm for the structural optimization of nanoalloy clusters.

An improved genetic algorithm (GA) is described that has been developed to increase the efficiency of finding the global minimum energy isomers for nanoalloy clusters. The GA is optimized for the example Pt12Pd12, with specific investigation of: the effect of biasing the initial population by seeding; the effect of removing specified clusters from the population ("predation"); and the effect of varying the type of mutation operator applied. These changes are found to significantly enhance the efficiency of the GA, which is subsequently demonstrated by the application of the best strategy to a new cluster, namely Pt19Pd19.

Journal Article↗

Analysis of homologous gene clusters in Caenorhabditis elegans reveals striking regional cluster domains.

An algorithm for detecting local clusters of homologous genes was applied to the genome of Caenorhabditis elegans. Clusters of two or more homologous genes are abundant, totaling 1391 clusters containing 4607 genes, over one-fifth of all genes in C. elegans. Cluster genes are distributed unevenly in the genome, with the large majority located on autosomal chromosome arms, regions characterized by higher genetic recombination and more repeat sequences than autosomal centers and the X chromosome. Cluster genes are transcribed at much lower levels than average and very few have gross phenotypes as assayed by RNAi-mediated reduction of function. The molecular identity of cluster genes is unusual, with a preponderance of nematode-specific gene families that encode putative secreted and transmembrane proteins, and enrichment for genes implicated in xenobiotic detoxification and innate immunity. Gene clustering in Drosophila melanogaster is also substantial and the molecular identity of clustered genes follows a similar pattern. I hypothesize that autosomal chromosome arms in C. elegans undergo frequent local gene duplication and that these duplications support gene diversification and rapid evolution in response to environmental challenges. Although specific gene clusters have been documented in C. elegans, their abundance, genomic distribution, and unusual molecular identities were previously unrecognized.

Amino Acid Sequence↗

Evaluation of an algorithm of tagging SNPs selection by linkage disequilibrium.

BACKGROUND: Single nucleotide polymorphisms (SNPs) are the most abundant kind of genetic polymorphism in the human genome. They are important in both genetic research and genetic testing in a clinical setting, such as in the area of pharmacogenetics. In order to improve efficiency, tagging SNPs (tagSNPs) are selected in genes of interest to represent other co-related SNPs in linkage disequilibrium (LD) with the tagSNPs. Various algorithms have been proposed to identify a subset of single nucleotide polymorphisms as tagSNPs. Most algorithms of tagSNPs selection are haplotype-based, in which the spatial relationship between SNPs is considered. Currently, a more efficient cluster-based algorithm is proposed which clusters SNPs solely by a LD parameter, such as r(2). Here, we evaluated the sample distribution of r(2) and its effect on the cluster-based tagSNPs selection. DESIGN AND METHODS: The genotype data of 198 individual within a 500-kb region on 5q31 was used to evaluate the sample distribution of r(2) and its effect on the cluster-based tagSNPs selection. RESULTS: It was found that the degree of variation of LD depends on the LD structure of genes. CONCLUSION: As a cluster-based tagSNPs selection algorithm does not take into account the spatial position of SNPs, a more stringent r(2) threshold is required to achieve more reliable tagSNPs selection.

Algorithms↗

[Hierarchical image segmentation based on watershed filtering and fuzzy cluster].

Watershed algorithms is an automatic segmentation scheme to generate closed outlines, which might give rise to over-segmentation, i.e., numerous small segmented closed regions that blur the target contours or shapes. In this paper, the author proposed a scheme to merge the small regions and generate a hierarchical segmentation representation. We first define the dissimilarity measure between the neighboring regions, and based on this dissimilarity measurement, a fuzzy matrix calculation followed by hierarchical segmentation is performed. Experiments with various types of images are given.

Fuzzy Logic↗

Clustering cDNA sequences.

A set of programs has been written to quantify the similarities between large numbers of cDNA sequences. This information is used to cluster similar sequences together. The main program can cluster thousands of cDNA sequences per day using a novel, computationally inexpensive algorithm. The clustering information is kept in a small index file so that disk storage requirements are negligible. Using this index file, subsidiary programs create various views and statistical summaries of the entire cDNA sequence collection.

Algorithms↗

Clustering of domains of functionally related enzymes in the interaction database PRECISE by the generation of primary sequence patterns.

The PRECISE database was developed by our laboratory to allow for the systematic study of the ligand interactions common to a set of functionally related enzymes, where an interaction site is defined broadly as any residue(s) that interact with a ligand. During the construction of PRECISE, enzyme chains are extracted from the protein data bank (PDB) and clustered according to functional homology as defined by the enzyme commission (EC) nomenclature system. A sequence representative is chosen from each cluster based on the criterion set forth by the non-redundant PDB set, and pair-wise alignments of each cluster member to the representative are performed. Atom-based residue-ligand interactions are calculated for each cluster member, and the summation of ligand interactions for all cluster members at each aligned position is determined. Although we were able to successfully align most clusters using a simple dynamic programming algorithm, several cluster created exhibited poor pair-wise alignments of each cluster member to its sequence representative. We hypothesized that the observed alignment problems were, in most cases, due to the incorrect separation and alignment of different domains in multi-domain proteins, a mistake that frequently causes error proliferation in functional annotation. Here we present the results of generating primary sequence patterns for each poorly aligned cluster in PRECISE to assess the extent to which multi-domain proteins that are incorrectly aligned contributes to poor pair-wise alignments of each cluster member to its representative. This requires the use of an iterative locally optimal pair-wise alignment algorithm to build a hierarchical similarity-based sequence pattern for a set of functionally related enzymes. Our results show that poor alignments in PRECISE are caused most frequently by the misalignment of multi-domain proteins, and that the generation of primary sequence patterns for the assignment of sequence family membership yields better alignments for the functionally related enzyme clusters in PRECISE than our original alignment algorithm.

Amino Acid Sequence↗

Discovery of morphological subgroups that correlate with severity of symptoms in interstitial cystitis: a proposed biopsy classification system.

PURPOSE: We identified morphologically distinct subgroups in interstitial cystitis using cluster analysis and investigated the associations between cluster membership and urinary symptoms. MATERIALS AND METHODS: Of 637 patients enrolled in the Interstitial Cystitis Data Base Study 203 (32%) provided bladder biopsies at baseline screening, representing the focus of this analysis. A cluster analysis algorithm implemented in SAS PROC CLUSTER using standardized distances to measure the dissimilarity of each pair of patients with respect to select histopathological features was used to construct subgroups of these patients. Multivariate regression models for baseline nighttime and 24-hour voiding frequency, urinary urgency and pain were developed, incorporating indicator variables for cluster membership as predictors. Longitudinal urinary symptom profiles during 3 years of followup were also compared among the morphology clusters. RESULTS: Three morphology clusters were identified, corresponding to unique pathological groupings. In cluster C2 7 patients showed multiple pathological features of parenchymal damage, including several inflammatory features. In cluster C1 17 patients was characterized by complete denudation of the urothelium and variable edema. In cluster C0 in 179 patients none of the pathological features were present above the specified thresholds for C2. Cluster membership was significantly associated with baseline nighttime and 24-hour frequency (p <0.001, and with urinary urgency (p = 0.03). These significant increases in baseline symptom severity among clusters from C0 to C1 to C2 persisted throughout the 3 years of followup. CONCLUSIONS: These results suggest an important role for histopathological features in the predictive modeling of interstitial cystitis symptoms.

Adult↗

CLICK and EXPANDER: a system for clustering and visualizing gene expression data.

MOTIVATION: Microarrays have become a central tool in biological research. Their applications range from functional annotation to tissue classification and genetic network inference. A key step in the analysis of gene expression data is the identification of groups of genes that manifest similar expression patterns. This translates to the algorithmic problem of clustering genes based on their expression patterns. RESULTS: We present a novel clustering algorithm, called CLICK, and its applications to gene expression analysis. The algorithm utilizes graph-theoretic and statistical techniques to identify tight groups (kernels) of highly similar elements, which are likely to belong to the same true cluster. Several heuristic procedures are then used to expand the kernels into the full clusters. We report on the application of CLICK to a variety of gene expression data sets. In all those applications it outperformed extant algorithms according to several common figures of merit. We also point out that CLICK can be successfully used for the identification of common regulatory motifs in the upstream regions of co-regulated genes. Furthermore, we demonstrate how CLICK can be used to accurately classify tissue samples into disease types, based on their expression profiles. Finally, we present a new java-based graphical tool, called EXPANDER, for gene expression analysis and visualization, which incorporates CLICK and several other popular clustering algorithms. AVAILABILITY: http://www.cs.tau.ac.il/~rshamir/expander/expander.html

Algorithms↗

Multicentre, cluster-randomized clinical trial of algorithms for critical-care enteral and parenteral therapy (ACCEPT).

BACKGROUND: The provision of nutritional support for patients in intensive care units (ICUs) varies widely both within and between institutions. We tested the hypothesis that evidence-based algorithms to improve nutritional support in the ICU would improve patient outcomes. METHODS: A cluster-randomized controlled trial was performed in the ICUs of 11 community and 3 teaching hospitals between October 1997 and September 1998. Hospital ICUs were stratified by hospital type and randomized to the intervention or control arm. Patients at least 16 years of age with an expected ICU stay of at least 48 hours were enrolled in the study (n = 499). Evidence-based recommendations were introduced in the 7 intervention hospitals by means of in-service education sessions, reminders (local dietitian, posters) and academic detailing that stressed early institution of nutritional support, preferably enteral. RESULTS: Two hospitals crossed over and were excluded from the primary analysis. Compared with the patients in the control hospitals (n = 214), the patients in the intervention hospitals (n = 248) received significantly more days of enteral nutrition (6.7 v. 5.4 per 10 patient-days; p = 0.042), had a significantly shorter mean stay in hospital (25 v. 35 days; p = 0.003) and showed a trend toward reduced mortality (27% v. 37%; p = 0.058). The mean stay in the ICU did not differ between the control and intervention groups (10.9 v. 11.8 days; p = 0.7). INTERPRETATION: Implementation of evidence-based recommendations improved the provision of nutritional support and was associated with improved clinical outcomes.

APACHE↗

Algorithm for the detection of fine clustered calcifications on film mammograms.

An algorithmic process for the detection and marking of clustered calcifications in digitized film-screen mammograms has been applied to mammograms from 50 clinical cases sampled at two digitization levels, in both the craniocaudal and mediolateral views. In all but one case the detector accurately located suggestive clusters found by radiologists in normal screening. In five cases additional clusters were also found by the detector. The detector has a negligible false-positive rate for the detection of clustered calcifications, although it is sensitive to clusters of emulsion defects displayed as artifactual calcification densities in the original film. The detector is flexible in structure and is easily adapted to various calcification/cluster criteria. The detector shows considerable promise when applied to clinical examples but will require refinement before formal testing.

Algorithms↗

[An adaptive criterion for cluster number estimation and the optimal algorithm for image segmentation].

In the algorithms for image segmentation, the number of clusters (NOC), which impacts on the segmentation results, should be first solved, and its correct estimation both theoretically and in application is of much importance. The authors propose an adaptive total energy criterion (ATEC) based on Markov random fields (MRF). The correct NOC of different images can be obtained by minimizing the ATEC and the parameters in the criterion are estimated by expectation maximization algorithm and maximum pseudo-likelihood method. The experiments show that the NOC can be automatically detected by adjusting the parameters, and the segmentation with the estimated NOC can be obtained by the maximum a posteriori at the same time.

Algorithms↗