PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

EvIdent: a functional magnetic resonance image analysis system.

EvIdent (EVent IDENTification) is a user-friendly, algorithm-rich, exploratory data analysis software for quickly detecting, investigating, and visualizing novel events in a set of images as they evolve in time and/or frequency. For instance, in a series of functional magnetic resonance neuroimages, novelty may manifest itself as neural activations in a time course. The core of the system is an enhanced variant of the fuzzy c-means clustering algorithm. Fuzzy clustering obviates the need for models of the underlying requisite biological function, models that are often statistically suspect.

Algorithms↗

Molecular profiles as predictive marker for the effect of overall treatment time of radiotherapy in supraglottic larynx squamous cell carcinomas.

BACKGROUND AND PURPOSE: Reduction of the overall treatment time of radiotherapy increases the probability of local tumour control, but it does not benefit all patients. Identification of molecular marker profiles may aid in the selection of patients likely to benefit from accelerated radiotherapy. PATIENTS AND METHODS: Two hundred and nine patients with SCC of the supraglottic larynx received primary radiotherapy in the randomised DAHANCA trials to 66-68 Gy, 2 Gy/fx but with different overall treatment times of 9.5 week, 6.5 week and 5.5 week. Formalin-fixed paraffin embedded tumour slides were assessed by immunohistochemistry for expression of EGFr, E-cadherin, KI-67 and Bcl-2 and the TP53 mutation profile was determined using PCR-amplification, DHPLC and sequencing. The profiles were established using a hierarchical clustering algorithm with a Bayesian information criterion for cluster number optimisation. RESULTS: Full data-set were available for 158 patients and four almost equally sized clusters were identified. One of these clusters differed significantly with respect to local control compared to the other clusters: the cluster (n=36) characterised by wild type TP53, low expression of E-cadherin and Bcl-2, moderate KI-67 and EGFr, was not influenced by a reduction in the overall treatment time (P=0.6) whereas the other clusters showed an increase in local control when the overall treatment time of radiotherapy was reduced. This was also partially seen with disease specific survival as the endpoint. CONCLUSIONS: Molecular marker profiling may aid in the selection of patients that will benefit of a reduction in overall treatment time of radiotherapy in SCC of the supraglottic larynx.

Algorithms↗

Statistically based postprocessing of phylogenetic analysis by clustering.

MOTIVATION: Phylogenetic analyses often produce thousands of candidate trees. Biologists resolve the conflict by computing the consensus of these trees. Single-tree consensus as postprocessing methods can be unsatisfactory due to their inherent limitations. RESULTS: In this paper we present an alternative approach by using clustering algorithms on the set of candidate trees. We propose bicriterion problems, in particular using the concept of information loss, and new consensus trees called characteristic trees that minimize the information loss. Our empirical study using four biological datasets shows that our approach provides a significant improvement in the information content, while adding only a small amount of complexity. Furthermore, the consensus trees we obtain for each of our large clusters are more resolved than the single-tree consensus trees. We also provide some initial progress on theoretical questions that arise in this context.

Algorithms↗

A fractal approach to the segmentation of microcalcifications in digital mammograms.

This paper presents a computerized method for the automated segmentation of individual microcalcifications in a region of interest (ROI) known to contain a cluster in digital mammograms. Mammographic parenchyma caj be accurately modeled with the fractal approach, but not areas with microcalcifications. The digitized image is divided into 16 x 16-pixel overlapping windows and those accurately modeled by the fractal model are eliminated. The next steps include local thresholding of the ROIs using an iterative method, the elimination of some of the artifacts and identification of the clustered microcalcifications using a clustering algorithm. The evaluation was performed on 81 simulated clusters superimposed on normal mammographic backgrounds and on a representative database of 408 real mammograms. Microcalcification locations were identified by two radiologists independently. These locations were compared to those found by the computer algorithm. An average of 59% of the simulated microcalcifications and 69% of the microcalcifications common to both radiologists were detected. The algorithm described provides a fully automated method for the segmentation of individual microcalcifications in an area of the mammogram known to contain a cluster.

Biopsy↗

Genes that co-cluster with estrogen receptor alpha in microarray analysis of breast biopsies.

The estrogen receptor plays a critical role in the pathogenesis and clinical behavior of breast cancer. To better understand the molecular basis of estrogen-dependent forms of this disease we studied gene expression profiles from 53 primary breast cancer biopsies. Gene expression data for more than 7000 genes were generated from each tumor sample with oligo microarrays. A standard correlation-clustering algorithm identified 18 genes that co-clustered with estrogen receptor alpha. Eleven of these genes had previously been associated with estrogen regulation or breast tumorigenesis including trefoil factor 1 and estrogen regulated LIV-1. Additional study of these 18 genes may further delineate the role of estrogen receptor in breast cancer, generate new predictive biomarkers for response to endocrine therapies and identify novel therapeutic targets.

Animals↗

A taxonomic study of Gardnerella vaginalis (Haemophilus vaginalis) Gardner and Dukes 1955.

Fifty-five strains received as Haemophilus vaginalis or as catalase-negative coryneform bacteria from the vagina together with 61 marker cultures were subjected to numerical phenetic analyses using 149 unit characters. The data were examined using the simple matching (SSM), Jaccard (SJ) and pattern (DP) coefficients and clustering was achieved using the average linkage algorithm. Cluster composition was not markedly affected by the coefficient used or by test error, estimated at 6 . 5%. The H. vaginalis strains formed a tight cluster which was only distantly related to representatives of the genera arthrobacter, Cellulomonas, Corynebacterium sensu stricto, Erysipelothrix, Haemophilus, Kurthia, Lactobacillus, Listeria and Propionibacterium but shared a high overall affinity to unclassified catalase-negative coryneforms which formed a discrete taxon, cluster 9. The H. vaginalis strains could be distinguished from the related strains in cluster 9 by several unrelated phenotypic characters. Using the S1 endonuclease assay, DNA-DNA hybridizations were performed with representative strains from the numerical as well as with reference strains of Bifidobacterium and Actinomyces. Haemophilus vaginalis was found to be a genotypically legitimate group and its DNA showed little homology with DNA from the marker strains tested. The DNA base composition of H. vaginalis was 42 to 44 mol % guanine plus cytosine. A new genus should be created to incorporate strains known as H. vaginalis or Corynebacterium vaginale. The name Gardnerella vaginalis proposed by Greenwood & Pickett (1979) is supported.

Base Composition↗

Numerical classification of Mycobacterium farcinogenes, Mycobacterium senegalense and related taxa.

Sixteen strains designated Mycobacterium farcinogenes, fifteen Mycobacterium senegalense, and ten Nocardia farcinica were, together with strains of Mycobacterium and Nocardia, subjected to numerical phenetic analyses using 96 unit characters. The data were examined using the simple matching (SSM), Jaccard (SJ) and pattern (DP) coefficients and clustering achieved using the unweighted average linkage algorithm. Cluster composition was not markedly affected by the coefficient used or by test error, estimated at 2.5%. The N. farcinica strains formed a distinct and homogeneous cluster in an aggregate taxon corresponding to the genus Nocardia. The M. farcinogenes and M. senegalense strains were recovered in well-defined and homogeneous phena within the genus Mycobacterium. Mycobacterium senegalense consistently showed a high overall similarity with clusters equated with M. chelonei and M. fortuitum, while the M. farcinogenes cluster was not closely associated with any of the mycobacterial clusters. The results are discussed in the light of other developments in the taxonomy of the bovine farcy bacteria.

Drug Resistance, Microbial↗

Regularized color clustering in medical image database.

A regularized color clustering algorithm is proposed to solve the color clustering problem in medical image database. By incorporating both measures of cluster separability and cluster compactness, regularized color clustering allows the automatic extraction of significant color groups with varying populations. Experimental results in different color spaces show that the regularized color clustering gives superior results in extracting significant distinct/abnormal color clusters without significant increases in cluster compactness. Furthermore, results of color clustering in different color spaces show that the LUV color space is more suitable for color clustering. Methods for selecting the regularization constants have also been suggested.

Color↗

Rare genetic variant risks in patients with sepsis-associated acute respiratory distress syndrome.

BACKGROUND: Acute respiratory distress syndrome (ARDS) is a complex, heterogeneous, and deadly condition often resulting from pulmonary lesions due to sepsis, among other causes. There is a lack of targeted therapies to specifically treat the patients. Common genetic factors in the population (frequency&#x2009;>&#x2009;1%) have been associated with ARDS susceptibility, but systematic genetic screens of the role of rare genetic variants are lacking. We used the network of known molecular interactions to identify ARDS risks from clusters of biologically related genes containing qualifying variants (QVs) with frequency&#x2009;<&#x2009;1% likely affecting function. METHODS: We conducted whole-exome sequencing in sepsis patients from the GEN-SEP cohort (n&#x2009;=&#x2009;822, of which 272 developed ARDS). A network-based heterogeneity clustering algorithm was used to discover significant gene clusters (p&#x2009;<&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5). Gene-set enrichment analysis and logistic regression models aggregating QVs were used for cross-verification to confirm consistency and deepen understanding of the effect sizes of gene clusters. RESULTS: We identified 19 significant clusters (plowest&#x2009;=&#x2009;3.29&#x2009;&#xd7;&#x2009;10&#x2013;10), each containing an average of 102 genes (11.6% mean similarity). QVs in nine gene clusters were associated with sepsis-associated ARDS (plowest&#x2009;=&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5) but were not associated with 28-day survival. Clusters were enriched in several biological pathways, notably the Toll-like receptor cascades. CONCLUSIONS: These results support a marked genetic heterogeneity underlying ARDS susceptibility and the presence of rare risk variants involving multiple biological processes that are associated with sepsis outcomes. Particularly, they underscore the importance of rare variants in genes of the Toll-like receptor cascades in the risk for sepsis-associated ARDS.

Humans↗

A meta-clustering analysis indicates distinct pattern alteration between two series of gene expression profiles for induced ischemic tolerance in rats.

We have developed a visualization methodology, called a "cluster overlap distribution map" (CODM), for comparing the clustering results of time series gene expression profiles generated under two different conditions. Although various clustering algorithms for gene expression data have been proposed, there are few effective methods to compare clustering results for different conditions. With CODM, the utilization of three-dimensional space and color allows intuitive visualization of changes in cluster set composition, changes in the expression patterns of genes between the two conditions, and relationship with other known gene information, such as transcription factors. We applied CODM to time series gene expression profiles obtained from rat four-vessel occlusion models combined with systemic hypotension and time-matched sham control animals (with sham operation), identifying distinct pattern alteration between the two. Comparisons of dynamic changes of time series gene expression levels under different conditions are important in various fields of gene expression profiling analysis, including toxicogenomics and pharmacogenomics. CODM will be valuable for various types of analyses within these fields, because it integrates and simultaneously visualizes various types of information across clustering results.

Algorithms↗

Spike detection II: automatic, perception-based detection and clustering.

OBJECTIVES: We developed perception-based spike detection and clustering algorithms. METHODS: The detection algorithm employs a novel, multiple monotonic neural network (MMNN). It is tested on two short-duration EEG databases containing 2400 spikes from 50 epilepsy patients and 10 control subjects. Previous studies are compared for database difficulty and reliability and algorithm accuracy. Automatic grouping of spikes via hierarchical clustering (using topology and morphology) is visually compared with hand marked grouping on a single record. RESULTS: The MMNN algorithm is found to operate close to the ability of a human expert while alleviating problems related to overtraining. The hierarchical and hand marked spike groupings are found to be strikingly similar. CONCLUSIONS: An automatic detection algorithm need not be as accurate as a human expert to be clinically useful. A user interface that allows the neurologist to quickly delete artifacts and determine whether there are multiple spike generators is sufficient.

Adolescent↗

Tolerating some redundancy significantly speeds up clustering of large protein databases.

MOTIVATION: Sequence clustering replaces groups of similar sequences in a database with single representatives. Clustering large protein databases like the NCBI Non-Redundant database (NR) using even the best currently available clustering algorithms is very time-consuming and only practical at relatively high sequence identity thresholds. Our previous program, CD-HI, clustered NR at 90% identity in approximately 1 h and at 75% identity in approximately 1 day on a 1 GHz Linux PC (Li et al., Bioinformatics, 17, 282, 2001); however even faster clustering speed is needed because the size of protein databases are rapidly growing and many applications desire a lower attainable thresholds. RESULTS: For our previous algorithm (CD-HI), we have employed short-word filters to speed up the clustering. In this paper, we show that tolerating some redundancy makes for more efficient use of these short-word filters and increases the program's speed 100 times. Our new program implements this technique and clusters NR at 70% identity within 2 h, and at 50% identity in approximately 5 days. Although some redundancy is present after clustering, our new program's results only differ from our previous program's by less than 0.4%.

Algorithms↗

Simulation algorithms for the random-cluster model.

We compare the performance of Monte Carlo algorithms for the simulation of the random-cluster representation of the q-state Potts model for continuous values of q. In particular we consider a local bond update method, a statistical reweighting method of percolation configurations, and a cluster algorithm, all of which generate Boltzmann statistics. The dynamic exponent z of the cluster algorithm appears to be quite small, and to assume the values of the Swendsen-Wang algorithm for q = 2 and 3. The cluster algorithm appears to be much more efficient than our versions of the other two methods for the simulation of the random-cluster model. The higher efficiency of the cluster method with respect to the local method is primarily due to the fact that the computer time usage of the local method increases more rapidly with system size; the difference between the dynamic exponents is less important.

Journal Article↗

Creating knowledgebases to text-mine PUBMED articles using clustering techniques.

Knowledgebase-mediated text-mining approaches work best when processing the natural language of domain-specific text. To enhance the utility of our successfully tested program-NeuroText, and to extend its methodologies to other domains, we have designed clustering algorithms, which is the principal step in automatically creating a knowledgebase. Our algorithms are designed to improve the quality of clustering by parsing the test corpus to include semantic and syntactic parsing

Algorithms↗

SEQOPTICS: a protein sequence clustering system.

BACKGROUND: Protein sequence clustering has been widely used as a part of the analysis of protein structure and function. In most cases single linkage or graph-based clustering algorithms have been applied. OPTICS (Ordering Points To Identify the Clustering Structure) is an attractive approach due to its emphasis on visualization of results and support for interactive work, e.g., in choosing parameters. However, OPTICS has not been used, as far as we know, for protein sequence clustering. RESULTS: In this paper, a system of clustering proteins, SEQOPTICS (SEQuence clustering with OPTICS) is demonstrated. The system is implemented with Smith-Waterman as protein distance measurement and OPTICS at its core to perform protein sequence clustering. SEQOPTICS is tested with four data sets from different data sources. Visualization of the sequence clustering structure is demonstrated as well. CONCLUSION: The system was evaluated by comparison with other existing methods. Analysis of the results demonstrates that SEQOPTICS performs better based on some evaluation criteria including Jaccard coefficient, Precision, and Recall. It is a promising protein sequence clustering method with future possible improvement on parallel computing and other protein distance measurements.

Algorithms↗

SamCluster: an integrated scheme for automatic discovery of sample classes using gene expression profile.

MOTIVATION: Feature (gene) selection can dramatically improve the accuracy of gene expression profile based sample class prediction. Many statistical methods for feature (gene) selection such as stepwise optimization and Monte Carlo simulation have been developed for tissue sample classification. In contrast to class prediction, few statistical and computational methods for feature selection have been applied to clustering algorithms for pattern discovery. RESULTS: An integrated scheme and corresponding program SamCluster for automatic discovery of sample classes based on gene expression profile is presented in this report. The scheme incorporates the feature selection algorithms based on the calculation of CV (coefficient of variation) and t-test into hierarchical clustering and proceeds as follows. At first, the genes with their CV greater than the pre-specified threshold are selected for cluster analysis, which results in two putative sample classes. Then, significantly differentially expressed genes in the two putative sample classes with p-values < or = 0.01, 0.05, or 0.1 from t-test are selected for further cluster analysis. The above processes were iterated until the two stable sample classes were found. Finally, the consensus sample classes are constructed from the putative classes that are derived from the different CV thresholds, and the best putative sample classes that have the minimum distance between the consensus classes and the putative classes are identified. To evaluate the performance of the feature selection for cluster analysis, the proposed scheme was applied to four expression datasets COLON, LEUKEMIA72, LEUKEMIA38, and OVARIAN. The results show that there are only 5, 1, 0, and 0 samples that have been misclassified, respectively. We conclude that the proposed scheme, SamCluster, is an efficient method for discovery of sample classes using gene expression profile. AVAILABILITY: The related program SamCluster is available upon request or from the web page http://www.sph.uth.tmc.edu:8052/hgc/Downloads.asp.

Algorithms↗

A note on the ICS algorithm with corrections and theoretical analysis.

In [1], Ozdemir and Akarun proposed an intercluster separation (ICS) fuzzy clustering algorithm. The ICS algorithm is useful in combined quantization and dithering. However, there are two errors in the update equations for the ICS algorithm. This correspondence first points out these errors and gives their corrections. Since the parameters m, c, and gamma are important factors in the performance of ICS, we also conduct a theoretical analysis of these ICS parameters. In order to analyze the parameters in ICS, we devise a theorem for the calculation of the Hessian matrix from the ICS objective function. We establish the fixed-point property of ICS based on the decomposition of the Hessian matrix and then analyze the effect of the parameters. Finally, we propose a numerical approach in choosing the appropriate parameters m and gamma for ICS. These experimental results give a better numerical perspective on the effect of parameters in ICS and have conclusions consistent with our theoretical analysis.

Algorithms↗

On clustering fMRI time series.

Analysis of fMRI time series is often performed by extracting one or more parameters for the individual voxels. Methods based, e.g., on various statistical tests are then used to yield parameters corresponding to probability of activation or activation strength. However, these methods do not indicate whether sets of voxels are activated in a similar way or in different ways. Typically, delays between two activated signals are not identified. In this article, we use clustering methods to detect similarities in activation between voxels. We employ a novel metric that measures the similarity between the activation stimulus and the fMRI signal. We present two different clustering algorithms and use them to identify regions of similar activations in an fMRI experiment involving a visual stimulus.

Algorithms↗