PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Multispectral magnetic resonance images segmentation using fuzzy Hopfield neural network.

This paper demonstrates a fuzzy Hopfield neural network for segmenting multispectral MR brain images. The proposed approach is a new unsupervised 2-D Hopfield neural network based upon the fuzzy clustering technique. Its implementation consists of the combination of 2-D Hopfield neural network and fuzzy c-means clustering algorithm in order to make parallel implementation for segmenting multispectral MR brain images feasible. For generating feasible results, a fuzzy c-means clustering strategy is included in the Hopfield neural network to eliminate the need for finding weighting factors in the energy function which is formulated and based on a basic concept commonly used in pattern classification, called the 'within-class scatter matrix' principle. The suggested fuzzy c-means clustering strategy has also been proven to be convergent and to allow the network to learn more effectively than the conventional Hopfield neural network. The experimental results show that a near optimal solution can be obtained using the fuzzy Hopfield neural network based on the within-class scatter matrix.

Algorithms↗

A phylogenomic analysis of the Ascomycota.

An automated procedure was developed to extract orthologous sequences from fungal genomes and incorporate them into phylogenomic analyses in a timely and efficient manner. This approach involves parsing an all versus all BLASTP search of 17 proteomes and creating a similarity matrix from e-values, which is then used to cluster proteins into related groups by means of a Markov Clustering algorithm. After performing this analysis at different stringency levels, 854 single copy protein clusters, which were ubiquitously distributed in all 17 proteomes, were identified. These clusters were culled to include only those clusters where all proteins had best hits to and received hits from a protein within the same cluster. The final data set included gapless alignments for 781 clusters of orthologous sequences that were concatenated into one super alignment containing 195,664 amino acid characters. Neighbor-joining distance and maximum likelihood analyses resulted in identical topologies and all except one node received 100% bootstrap support. The node supporting Stagonospora nodorum's position received 83% support or higher; it was also the only taxon differentially resolved in the maximum parsimony analyses. All analyses resolved the two derived subphyla Pezizomycotina and Saccharomycotina, and Schizosaccharomyces pombe as an early diverging lineage of the Ascomycota. Importantly, these analyses resolved the Leotiomycetes as the sister group to the Sordariomycetes, a region of the Ascomycota phylogeny that has remained problematic in molecular phylogenetic studies of more limited character sampling. Additional phylogenetic analyses which included orthologous sequences from an unannotated ascomycotan genome (e.g., Coccidioides immitis) and subsets of orthologs with different characteristics supported this topology. Phylogenetic analyses of the 595 orthologs which included C. immitis resulted in an identical topology to the previous 781 ortholog analysis and correctly placed C. immitis in the Eurotiomycetes. This demonstrated the correct identification of orthologs and the ability to incorporate unannotated genomic data into a common phylogenetic analysis.

Ascomycota↗

Assignment of Staphylococcus isolates to groups by spa typing, SmaI macrorestriction analysis, and multilocus sequence typing.

The implementation of the new clustering algorithm Based Upon Repeat Pattern (BURP) into the Ridom StaphType software tool enables clustering based on spa typing data for Staphylococcus aureus. We compared clustering results obtained by spa typing/BURP to those obtained by currently well-established methods, i.e., SmaI macrorestriction analysis and multilocus sequence typing/eBURST. A total of 99 clinical S. aureus strains, including MRSA and representing major clonal lineages associated with important kinds of infections which have been prevalent in Germany and Central Europe during the last 10 years, were used for comparison. SmaI macrorestriction analysis revealed the highest discriminatory power, and clustering results for all three methods resulted in concordance values ranging from 96.8% between the two sequence-based methods to 93.4% between spa typing/BURP and SmaI macrorestriction/cluster analysis. The results of this study indicate that spa typing, together with BURP clustering, is a useful tool in S. aureus epidemiology, especially because of ease of use and the advantages of unambiguous sequence analysis as well as reproducibility and exchange of typing data.

Bacterial Typing Techniques↗

Groups of histopathologic abnormalities in brains of very low birthweight infants.

The neuropathologic changes in brains of very premature infants are well recognized but relatively few studies have attempted to identify if specific neuropathologic features cluster together. These data could assist in determining pathogenetic mechanisms of immature brain injury. The goal of this study is to identify which, if any, combinations of histologic features occur together. We identified the presence or absence of 19 histologic features in the brains of 67 infants from a multicenter study of 1,665 prematurely born infants whose birthweight was 500-1,500 grams. We used clustering algorithms and factor analysis to group pathologic features that occurred together. Our results indicate that certain histopathologic features do cluster. For example, telencephalic white matter astrocytosis occurs in 2 groups: 1) associated with amphophilic globules, and, 2) in an uncorrelated group, associated with focal macrophage deposits and coagulative necroses. Parenchymal hemorrhage was not found to be associated with any telencephalic leukoencephalopathy, regardless of whether characterized by rarefaction, astrocytosis, focal coagulative necroses, or foci of macrophages in the white matter. Intraventricular hemorrhage and germinal matrix hemorrhage were not seen together more often than by chance expectation. Intraventricular hemorrhage was only marginally associated with parenchymal hemorrhage. Our data indicate that specific histopathologic features tend to preferentially cluster with each other in groups. This clustering may represent the manifestation of a common mechanism for each. These data should be valuable indicators for future research attempting to establish pathogenesis.

Algorithms↗

Review of MR image segmentation techniques using pattern recognition.

This paper has reviewed, with somewhat variable coverage, the nine MR image segmentation techniques itemized in Table II. A wide array of approaches have been discussed; each has its merits and drawbacks. We have also given pointers to other approaches not discussed in depth in this review. The methods reviewed fall roughly into four model groups: c-means, maximum likelihood, neural networks, and k-nearest neighbor rules. Both supervised and unsupervised schemes require human intervention to obtain clinically useful results in MR segmentation. Unsupervised techniques require somewhat less interaction on a per patient/image basis. Maximum likelihood techniques have had some success, but are very susceptible to the choice of training region, which may need to be chosen slice by slice for even one patient. Generally, techniques that must assume an underlying statistical distribution of the data (such as LML and UML) do not appear promising, since tissue regions of interest do not usually obey the distributional tendencies of probability density functions. The most promising supervised techniques reviewed seem to be FF/NN methods that allow hidden layers to be configured as examples are presented to the system. An example of a self-configuring network, FF/CC, was also discussed. The relatively simple k-nearest neighbor rule algorithms (hard and fuzzy) have also shown promise in the supervised category. Unsupervised techniques based upon fuzzy c-means clustering algorithms have also shown great promise in MR image segmentation. Several unsupervised connectionist techniques have recently been experimented with on MR images of the brain and have provided promising initial results. A pixel-intensity-based edge detection algorithm has recently been used to provide promising segmentations of the brain. This is also an unsupervised technique, older versions of which have been susceptible to oversegmenting the image because of the lack of clear boundaries between tissue types or finding uninteresting boundaries between slightly different types of the same tissue. To conclude, we offer some remarks about improving MR segmentation techniques. The better unsupervised techniques are too slow. Improving speed via parallelization and optimization will improve their competitiveness with, e.g., the k-nn rule, which is the fastest technique covered in this review. Another area for development is dynamic cluster validity. Unsupervised methods need better ways to specify and adjust c, the number of tissue classes found by the algorithm. Initialization is a third important area of research. Many of the schemes listed in Table II are sensitive to good initialization, both in terms of the parameters of the design, as well as operator selection of training data.(ABSTRACT TRUNCATED AT 400 WORDS)

Algorithms↗

Evaluating distance functions for clustering tandem repeats.

Tandem repeats are an important class of DNA repeats and much research has focused on their efficient identification, their use in DNA typing and fingerprinting, and their causative role in trinucleotide repeat diseases such as Huntington Disease, myotonic dystrophy, and Fragile-X mental retardation. We are interested in clustering tandem repeats into groups or families based on sequence similarity so that their biological importance may be further explored. To cluster tandem repeats we need a notion of pairwise distance which we obtain by alignment. In this paper we evaluate five distance functions used to produce those alignments--Consensus, Euclidean, Jensen-Shannon Divergence, Entropy-Surface, and Entropy-weighted. It is important to analyze and compare these functions because the choice of distance metric forms the core of any clustering algorithm. We employ a novel method to compare alignments and thereby compare the distance functions themselves. We rank the distance functions based on the cluster validation techniques--Average Cluster Density and Average Silhouette Width. Finally, we propose a multi-phase clustering method which produces good-quality clusters. In this study, we analyze clusters of tandem repeats from five sequences: Human Chromosomes 3, 5, 10 and X and C. elegans Chromosome III.

Algorithms↗

Framework for kernel regularization with application to protein clustering.

We develop and apply a previously undescribed framework that is designed to extract information in the form of a positive definite kernel matrix from possibly crude, noisy, incomplete, inconsistent dissimilarity information between pairs of objects, obtainable in a variety of contexts. Any positive definite kernel defines a consistent set of distances, and the fitted kernel provides a set of coordinates in Euclidean space that attempts to respect the information available while controlling for complexity of the kernel. The resulting set of coordinates is highly appropriate for visualization and as input to classification and clustering algorithms. The framework is formulated in terms of a class of optimization problems that can be solved efficiently by using modern convex cone programming software. The power of the method is illustrated in the context of protein clustering based on primary sequence data. An application to the globin family of proteins resulted in a readily visualizable 3D sequence space of globins, where several subfamilies and subgroupings consistent with the literature were easily identifiable.

Algorithms↗

Stokesian Dynamics Simulations of Ferromagnetic Colloidal Dispersions in a Simple Shear Flow.

We have investigated the behavior of clusters of ferromagnetic particles in a colloidal dispersion subjected to a simple shear flow. To do so, the Stokesian dynamics method has been used under the assumption that the effect of Brownian motion is negligible. For the case of no shear flow, the aggregate structures obtained by the Stokesian dynamics simulations agree well with Monte Carlo results qualitatively. We can, therefore, conclude that the Stokesian dynamics simulations can capture thick chainlike clusters without introducing a specific clustering algorithm, which is indispensable for Monte Carlo simulations. The behavior of the thick chainlike clusters in a simple shear flow is summarized as follows. The thick chainlike clusters decline in the shear flow direction as time advances. Since longer clusters experience larger shear forces, it is difficult for them to survive in such a situation. The thick chainlike clusters, therefore, dissociate into some short clusters. Such clusters are relatively stable in a shear flow, so that they do not decrease significantly any more. The viscosities have a strong relationship with the internal structures of the aggregates. The instantaneous viscosities, therefore, fluctuate significantly for the case of the thick chainlike clusters. Copyright 1998 Academic Press.

Journal Article↗

Hybrid systems for virtual screening: interest of fuzzy clustering applied to olfaction.

Kohonen neural networks, also known as Self Organizing Map (SOM), offer a useful 2D representation of the compound distribution inside a large chemical database. This distribution results from the compound organization in a molecular diversity hyperspace derived from a large set of molecular descriptors. Fuzzy techniques based on the "concept of partial truth" reveal to be also a valuable tool for the direct exploitation of chemical databases or SOM. In such cases a fuzzy clustering algorithm is used. In this paper, a complete hybrid system, combining SOM and fuzzy clustering, is applied. As example, a series of olfactory compounds was selected. The complexity of such information is that a same compound may exhibit different odors. It is shown how fuzzy logic helps to have a better understanding of the organization of the compounds. These hybrid systems, using simultaneously SOM and fuzzy clustering, are foreseen as powerful tools for "virtual pre-screening".

Fuzzy Logic↗

Venn Mapping: clustering of heterologous microarray data based on the number of co-occurring differentially expressed genes.

MOTIVATION: To evaluate microarray data, clustering is widely used to group biological samples or genes. However, problems arise when comparing heterologous databases. As the clustering algorithm searches for similarities between experiments, it will most likely first separate the data sets, masking relationships that exist between samples from different databases. RESULTS: We developed a program, Venn Mapper, to calculate the statistical significance of the number of co-occurring differentially expressed genes in any of the two experiments. For proof of principle, we analysed a heterologous data set of 170 microarrays including breast and prostate cancer microarray analyses. Significant overlap was found in an unsupervised analysis between metastasized prostate cancer and metastasized breast cancer and BRCA mutated breast cancer. A comparison between single microarray data and the averaged breast and prostate data sets was also evaluated. This analysis suggests that genes expressed higher in stromal cells are also implicated in metastatic prostate cancer and BRCA mutated breast cancer. The Venn Mapper program identifies overlaps between samples from heterologous data sets and directly extracts the genes responsible for the overlap. From this information novel biological hypotheses may be addressed. AVAILABILITY: Venn Mapper is freely available on http://www.erasmusmc.nl/gatcplatform. SUPPLEMENTARY INFORMATION: http://www.erasmusmc.nl/gatcplatform/vennmapper.html.

Algorithms↗

A latent variable model for chemogenomic profiling.

MOTIVATION: In haploinsufficiency profiling data, pleiotropic genes are often misclassified by clustering algorithms that impose the constraint that a gene or experiment belong to only one cluster. We have developed a general probabilistic model that clusters genes and experiments without requiring that a given gene or drug only appear in one cluster. The model also incorporates the functional annotation of known genes to guide the clustering procedure. RESULTS: We applied our model to the clustering of 79 chemogenomic experiments in yeast. Known pleiotropic genes PDR5 and MAL11 are more accurately represented by the model than by a clustering procedure that requires genes to belong to a single cluster. Drugs such as miconazole and fenpropimorph that have different targets but similar off-target genes are clustered more accurately by the model-based framework. We show that this model is useful for summarizing the relationship among treatments and genes affected by those treatments in a compendium of microarray profiles. AVAILABILITY: Supplementary information and computer code at http://genomics.lbl.gov/llda.

Computer Simulation↗

Molecular subtyping of breast cancer from traditional tumor marker profiles using parallel clustering methods.

PURPOSE: Recent small-sized genomic studies on the identification of breast cancer bioprofiles have led to profoundly dishomogenous results. Thus, we sought to identify distinct tumor profiles with possible clinical relevance based on clusters of immunohistochemical molecular markers measured on a large, single institution, case series. EXPERIMENTAL DESIGN: Tumor biological profiles were explored on 633 archival tissue samples analyzed by immunohistochemistry. Five validated markers were considered, i.e., estrogen receptors (ER), progesterone receptors (PR), Ki-67/MIB1 as a proliferation marker, HER2/NEU, and p53 in their original scale of measurement. The results obtained were analyzed by three different clustering algorithms. Four different indices were then used to select the different profiles (number of clusters). RESULTS: The best classification was obtained creating four clusters. Notably, three clusters were identified according to low, intermediate, and high ER/PR levels. A further subdivision in two biologically distinct subtypes was determined by the presence/absence of HER2/NEU and of p53. As expected, the cluster with high ER/PR levels was characterized by a much better prognosis and response to hormone therapy compared to that with the lowest ER/PR values. Notably, the cluster characterized by high HER2/NEU levels showed intermediate prognosis, but a rather poor response to hormone therapy. CONCLUSIONS: Our results show the possibility of profiling breast cancers by means of traditional markers, and have novel clinical implications on the definition of the prognosis of cancer patients. These findings support the existence of a tumor subtype that responds poorly to hormone therapy, characterized by HER2/NEU overexpression.

Adult↗

Use of cluster analysis technique for computerized recognition and prognosis of myocardial vulnerability.

Variability of factors exerting influence on arrhythmia origination leads to the appearance of polymodal distribution in sample space. Therefore, for the approximation of feature distribution, distribution mixture is used. To estimate the number of mixture components and to determine its parameters, cluster algorithm is used. The basic task of the algorithm is to identify the accumulation of vectors in the parallelepiped of their distribution. The accumulations of points are determined by testing statistical hypothesis of uniformity. On the basis of accumulations, clusters are formed and the parameters of normal mixtures of classes are estimated. Analysis of error matrix for recognition of mixture, enable to establish the parameters of the decision rule. The above algorithm was applied for recognition and prognosis of vulnerability of reentry and focal source in experiments on the right rabbit's atrium by using the electrophysiological parameters. We studied 30 cases of reentry, 36 cases of focal sources and 165-arrhythmia-free cases. As a result, we established the 7-class normal mixture which enabled a more effective (96.4%) recognition of the vulnerability types and 89.9% prognosis by features: increase in latency (theta), width of the interval of latency distribution, ratio theta/R, where R-refractory period.

Algorithms↗

A unified representation of multiprotein complex data for modeling interaction networks.

The protein interaction network presents one perspective for understanding cellular processes. Recent experiments employing high-throughput mass spectrometric characterizations have resulted in large data sets of physiologically relevant multiprotein complexes. We present a unified representation of such data sets based on an underlying bipartite graph model that is an advance over existing models of the network. Our unified representation allows for weighting of connections between proteins shared in more than one complex, as well as addressing the higher level organization that occurs when the network is viewed as consisting of protein complexes that share components. This representation also allows for the application of the rigorous MinMaxCut graph clustering algorithm for the determination of relevant protein modules in the networks. Statistically significant annotations of clusters in the protein-protein and complex-complex networks using terms from the Gene Ontology indicate that this method will be useful for posing hypotheses about uncharacterized components of protein complexes or uncharacterized relationships between protein complexes.

Algorithms↗

Seizure detection: evaluation of the Reveal algorithm.

OBJECTIVE: The aim of this study is to evaluate an improved seizure detection algorithm and to compare with two other algorithms and human experts. METHODS: 672 seizures from 426 epilepsy patients were examined with the (new) Reveal algorithm which utilizes 3 methods, novel in their application to seizure detection: Matching Pursuit, small neural network-rules and a new connected-object hierarchical clustering algorithm. RESULTS: Reveal had a sensitivity of 76% with a false positive rate of 0.11/h. Two other algorithms (Sensa and CNet) were tested and had sensitivities of 35.4 and 48.2% and false positive rates of 0.11/h and 0.75/h, respectively. CONCLUSIONS: This study validates the Reveal algorithm, and shows it to compare favorably with other methods. SIGNIFICANCE: Improved seizure detection can improve patient care in both the epilepsy monitoring unit and the intensive care unit.

Adolescent↗

Modeling and analysis of gene expression time-series based on co-expression.

In this paper a novel approach is introduced for modeling and clustering gene expression time-series. The radial basis function neural networks have been used to produce a generalized and smooth characterization of the expression time-series. A co-expression coefficient is defined to evaluate the similarities of the models based on their temporal shapes and the distribution of the time points. The profiles are grouped using a fuzzy clustering algorithm incorporated with the proposed co-expression coefficient metric. The results on artificial and real data are presented to illustrate the advantages of the metric and method in grouping temporal profiles. The proposed metric has also been compared with the commonly used correlation coefficient under the same procedures and the results show that the proposed method produces better biologically relevant clusters.

Algorithms↗

Incorporating genotyping uncertainty in haplotype inference for single-nucleotide polymorphisms.

The accuracy of the vast amount of genotypic information generated by high-throughput genotyping technologies is crucial in haplotype analyses and linkage-disequilibrium mapping for complex diseases. To date, most automated programs lack quality measures for the allele calls; therefore, human interventions, which are both labor intensive and error prone, have to be performed. Here, we propose a novel genotype clustering algorithm, GeneScore, based on a bivariate t-mixture model, which assigns a set of probabilities for each data point belonging to the candidate genotype clusters. Furthermore, we describe an expectation-maximization (EM) algorithm for haplotype phasing, GenoSpectrum (GS)-EM, which can use probabilistic multilocus genotype matrices (called "GenoSpectrum") as inputs. Combining these two model-based algorithms, we can perform haplotype inference directly on raw readouts from a genotyping machine, such as the TaqMan assay. By using both simulated and real data sets, we demonstrate the advantages of our probabilistic approach over the current genotype scoring methods, in terms of both the accuracy of haplotype inference and the statistical power of haplotype-based association analyses.

Algorithms↗

Identification and distribution of protein families in 120 completed genomes using Gene3D.

Using a new protocol, PFscape, we undertake a systematic identification of protein families and domain architectures in 120 complete genomes. PFscape clusters sequences into protein families using a Markov clustering algorithm (Enright et al., Nucleic Acids Res 2002;30:1575-1584) followed by complete linkage clustering according to sequence identity. Within each protein family, domains are recognized using a library of hidden Markov models comprising CATH structural and Pfam functional domains. Domain architectures are then determined using DomainFinder (Pearl et al., Protein Sci 2002;11:233-244) and the protein family and domain architecture data are amalgamated in the Gene3D database (Buchan et al., Genome Res 2002;12:503-514). Using Gene3D, we have investigated protein sequence space, the extent of structural annotation, and the distribution of different domain architectures in completed genomes from all kingdoms of life. As with earlier studies by other researchers, the distribution of domain families shows power-law behavior such that the largest 2,000 domain families can be mapped to approximately 70% of nonsingleton genome sequences; the remaining sequences are assigned to much smaller families. While approximately 50% of domain annotations within a genome are assigned to 219 universal domain families, a much smaller proportion (< 10%) of protein sequences are assigned to universal protein families. This supports the mosaic theory of evolution whereby domain duplication followed by domain shuffling gives rise to novel domain architectures that can expand the protein functional repertoire of an organism. Functional data (e.g. COG/KEGG/GO) integrated within Gene3D result in a comprehensive resource that is currently being used in structure genomics initiatives and can be accessed via http://www.biochem.ucl.ac.uk/bsm/cath/Gene3D/.

Amino Acid Sequence↗