PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Joint entropy maximization in kernel-based topographic maps.

A new learning algorithm for kernel-based topographic map formation is introduced. The kernel parameters are adjusted individually so as to maximize the joint entropy of the kernel outputs. This is done by maximizing the differential entropies of the individual kernel outputs, given that the map's output redundancy, due to the kernel overlap, needs to be minimized. The latter is achieved by minimizing the mutual information between the kernel outputs. As a kernel, the (radial) incomplete gamma distribution is taken since, for a gaussian input density, the differential entropy of the kernel output will be maximal. Since the theoretically optimal joint entropy performance can be derived for the case of nonoverlapping gaussian mixture densities, a new clustering algorithm is suggested that uses this optimum as its "null" distribution. Finally, it is shown that the learning algorithm is similar to one that performs stochastic gradient descent on the Kullback-Leibler divergence for a heteroskedastic gaussian mixture density model.

Algorithms↗

The phylogenetic diversity of eukaryotic transcription.

Eukaryotic transcription is a highly regulated process involving interactions between large numbers of proteins. To analyse the phylogenetic distribution of the components of this process, six crown eukaryote group genomes were queried with a reference set of transcription-associated (TA) proteins. On average, one in 10 proteins encoded by these genomes were found to be homologous to sequences in the reference set. Analysis of families identified using an accurate sequence clustering algorithm and containing both TA proteins and eukaryotic sequences showed that in two-thirds of the families the homologues originate from a single kingdom. Furthermore, in only 15% of the fungal-specific clusters are the homologues present in both budding and fission yeast, as compared with the metazoan-specific clusters where 53% of the homologues originate from two or more species. Families whose members comprise general transcription factor or RNA polymerase subunits exhibit a low degree of taxon specificity, suggesting that the transcription initiation complex is highly conserved. This contrasts with transcriptional regulator families, that are primarily taxon-specific, indicating proteins controlling gene activation exhibit considerable sequence diversity across the eukaryotic domain.

Animals↗

5S ribosomal ribonucleic acid sequences in Bacteroides and Fusobacterium: evolutionary relationships within these genera and among eubacteria in general.

The 5S ribosomal ribonucleic acid (rRNA) sequences were determined for Bacteroides fragilis, Bacteroides thetaiotaomicron, Bacteroides capillosus, Bacteroides veroralis, Porphyromonas gingivalis, Anaerorhabdus furcosus, Fusobacterium nucleatum, Fusobacterium mortiferum, and Fusobacterium varium. A dendrogram constructed by a clustering algorithm from these sequences, which were aligned with all other hitherto known eubacterial 5S rRNA sequences, showed differences as well as similarities with respect to results derived from 16S rRNA analyses. In the 5S rRNA dendrogram, Bacteroides clustered together with Cytophaga and Fusobacterium, as in 16S rRNA analyses. Intraphylum relationships deduced from 5S rRNAs suggested that Bacteroides is specifically related to Cytophaga rather than to Fusobacterium, as was suggested by 16S rRNA analyses. Previous taxonomic considerations concerning the genus Bacteroides, based on biochemical and physiological data, were confirmed by the 5S rRNA sequence analysis.

Algorithms↗

Meta-analyzing left hemisphere language areas: phonology, semantics, and sentence processing.

The advent of functional neuroimaging has allowed tremendous advances in our understanding of brain-language relationships, in addition to generating substantial empirical data on this subject in the form of thousands of activation peak coordinates reported in a decade of language studies. We performed a large-scale meta-analysis of this literature, aimed at defining the composition of the phonological, semantic, and sentence processing networks in the frontal, temporal, and inferior parietal regions of the left cerebral hemisphere. For each of these language components, activation peaks issued from relevant component-specific contrasts were submitted to a spatial clustering algorithm, which gathered activation peaks on the basis of their relative distance in the MNI space. From a sample of 730 activation peaks extracted from 129 scientific reports selected among 260, we isolated 30 activation clusters, defining the functional fields constituting three distributed networks of frontal and temporal areas and revealing the functional organization of the left hemisphere for language. The functional role of each activation cluster is discussed based on the nature of the tasks in which it was involved. This meta-analysis sheds light on several contemporary issues, notably on the fine-scale functional architecture of the inferior frontal gyrus for phonological and semantic processing, the evidence for an elementary audio-motor loop involved in both comprehension and production of syllables including the primary auditory areas and the motor mouth area, evidence of areas of overlap between phonological and semantic processing, in particular at the location of the selective human voice area that was the seat of partial overlap of the three language components, the evidence of a cortical area in the pars opercularis of the inferior frontal gyrus dedicated to syntactic processing and in the posterior part of the superior temporal gyrus a region selectively activated by sentence and text processing, and the hypothesis that different working memory perception-actions loops are identifiable for the different language components. These results argue for large-scale architecture networks rather than modular organization of language in the left hemisphere.

Brain Mapping↗

Metric and multidimensional scaling: efficient tools for clustering molecular conformations.

The application of metric and multidimensional scaling to conformer ensembles was demonstrated in this work. An automated process was devised to cluster and assign group memberships and cluster representatives. The method allows rapid clustering, leading to intuitive results that can be visually inspected. Multidimensional scaling was found to be superior to metric scaling for clustering conformers. The performance of different hierarchical clustering algorithms was compared using multidimensional plots, and the group average method was found to perform best.

Amino Acid Sequence↗

Docking of flexible molecules using multiscale ligand representations.

Structural genomics will yield an immense number of protein three-dimensional structures in the near future. Automated theoretical methodologies are needed to exploit this information and are likely to play a pivotal role in drug discovery. Here, we present a fully automated, efficient docking methodology that does not require any a priori knowledge about the location of the binding site or function of the protein. The method relies on a multiscale concept where we deal with a hierarchy of models generated for the potential ligand. The models are created using the k-means clustering algorithm. The method was tested on seven protein-ligand complexes. In the largest complex, human immunodeficiency virus reverse transcriptase/nevirapin, the root mean square deviation value when comparing our results to the crystal structure was 0.29 A. We demonstrate on an additional 25 protein-ligand complexes that the methodology may be applicable to high throughput docking. This work reveals three striking results. First, a ligand can be docked using a very small number of feature points. Second, when using a multiscale concept, the number of conformers that require to be generated can be significantly reduced. Third, fully flexible ligands can be treated as a small set of rigid k-means clusters.

Algorithms↗

Different gene expression patterns in invasive lobular and ductal carcinomas of the breast.

Invasive ductal carcinoma (IDC) and invasive lobular carcinoma (ILC) are the two major histological types of breast cancer worldwide. Whereas IDC incidence has remained stable, ILC is the most rapidly increasing breast cancer phenotype in the United States and Western Europe. It is not clear whether IDC and ILC represent molecularly distinct entities and what genes might be involved in the development of these two phenotypes. We conducted comprehensive gene expression profiling studies to address these questions. Total RNA from 21 ILCs, 38 IDCs, two lymph node metastases, and three normal tissues were amplified and hybridized to approximately 42,000 clone cDNA microarrays. Data were analyzed using hierarchical clustering algorithms and statistical analyses that identify differentially expressed genes (significance analysis of microarrays) and minimal subsets of genes (prediction analysis for microarrays) that succinctly distinguish ILCs and IDCs. Eleven of 21 (52%) of the ILCs ("typical" ILCs) clustered together and displayed different gene expression profiles from IDCs, whereas the other ILCs ("ductal-like" ILCs) were distributed between different IDC subtypes. Many of the differentially expressed genes between ILCs and IDCs code for proteins involved in cell adhesion/motility, lipid/fatty acid transport and metabolism, immune/defense response, and electron transport. Many genes that distinguish typical and ductal-like ILCs are involved in regulation of cell growth and immune response. Our data strongly suggest that over half the ILCs differ from IDCs not only in histological and clinical features but also in global transcription programs. The remaining ILCs closely resemble IDCs in their transcription patterns. Further studies are needed to explore the differences between ILC molecular subtypes and to determine whether they require different therapeutic strategies.

Breast Neoplasms↗

Pulsatile thyrotropin release in patients with untreated pituitary disease.

Pulsatile and nocturnal TSH secretion was investigated in 16 healthy controls (group A) and 19 patients with untreated pituitary disease [7 were euthyroid without suprasellar extension (group B), 6 were euthyroid with suprasellar extension (group C) of pituitary lesions, and 6 were hypothyroid with or without suprasellar extension (group D)]. Pulse analysis was performed using Desade and Cluster algorithms. No changes were observed among groups A-D in mean 24-h TSH pulse amplitude [values given as mean +/- SD; Desade, 0.4 +/- 0.2 vs. 0.7 +/- 0.4 vs. 0.6 +/- 0.4 vs. 0.5 +/- 0.2 mU/L (P = NS); Cluster, 0.4 +/- 0.2 vs. 0.7 +/- 0.4 vs. 0.5 +/- 0.3 vs. 0.4 +/- 0.2 mU/L (P = NS)] or in the mean 24-h TSH pulse frequency (approximately 10 pulses/24 h). The mean 24-h TSH concentration was highly correlated to the mean 24-h TSH pulse amplitude in controls (r = 0.93; P < 0.001) and patients (r = 0.63; P < 0.01), but not to the mean 24-h TSH pulse frequency. The nocturnal TSH surge was similar in controls and euthyroid patients without suprasellar extension (group A, 1.0 +/- 0.6; group B, 1.3 +/- 1.3 mU/L; P = NS), but was decreased in euthyroid patients with suprasellar extension (group C, 0.3 +/- 1.0 mU/L; P < 0.05) and hypothyroid patients (group D, 0.4 +/- 0.4 mU/L; P < 0.05). The decreased nocturnal TSH surge was associated with a loss of the usual nocturnal increase in TSH pulse amplitude, whereas the usual nocturnal increase in TSH pulse frequency was maintained. In conclusion, 1) mean 24-h TSH pulse amplitude and frequency are unchanged in untreated patients with pituitary disease; and 2) patients with central hypothyroidism as well as euthyroid patients with suprasellar extension of pituitary lesions had a decreased nocturnal TSH surge associated with a loss of the usual nocturnal increase in TSH amplitude, but not TSH pulse frequency.

Adult↗

Identifying superficial, muscle-invasive, and metastasizing transitional cell carcinoma of the bladder: use of cDNA array analysis of gene expression profiles.

PURPOSE: Expression profiling by DNA microarray technology permits the identification of genes underlying clinical heterogeneity of bladder cancer and which might contribute to disease progression, thereby improving assessment of treatment and prediction of patient outcome. EXPERIMENTAL DESIGN: Invasive (20) and superficial (22) human bladder tumors from 34 patients with known outcome regarding disease recurrence and progression were analyzed by filter-based cDNA arrays (Atlas Human Cancer 1.2; BD Biosciences Clontech) containing 1185 genes. For 9 genes, array data were confirmed using real-time reverse transcription-PCR. Additionally, Atlas array data were validated using Affymetrix GeneChip oligonucleotide arrays with 22,283 human gene fragments and expressed sequence tags sequences in a subset of three superficial and six invasive bladder tumors. RESULTS: A two-way clustering algorithm using different subsets of gene expression data, including a subset of 41 genes validated by the oligonucleotide array (Affymetrix), classified tumor samples according to clinical outcome as superficial, invasive, or metastasizing. Furthermore, (a) a clonal origin of superficial tumors, (b) highly similar gene expression patterns in different areas of invasive tumors, and (c) an invasive-like pattern was observed in bladder mucosas derived from patients with locally advanced disease. Several gene clusters that characterized invasive or superficial tumors were identified. In superficial bladder tumors, increased mRNA levels of genes encoding transcription factors, molecules involved in protein synthesis and metabolism, and some proteins involved into cell cycle progression and differentiation were observed, whereas transcripts for immune, extracellular matrix, adhesion, peritumoral stroma and muscle tissue components, proliferation, and cell cycle controllers were up-regulated in invasive tumors. CONCLUSIONS: Gene expression profiling of human bladder cancers provides insight into the biology of bladder cancer progression and identifies patients with distinct clinical phenotypes.

Algorithms↗

Cluster Monte Carlo algorithm for the quantum rotor model.

We propose a highly efficient "worm"-like cluster Monte Carlo algorithm for the quantum rotor model in the link-current representation. We explicitly prove detailed balance for the algorithm even in the presence of disorder. For the pure quantum rotor model with mu=0, the algorithm yields high- precision estimates for the critical point K(c)=0.333 05(5) and the correlation length exponent nu=0.670(3). For the disordered case, mu=1 / 2+/-1 / 2, we find nu=1.15(10).

Journal Article↗

Gene expression profiles of normal proliferating and differentiating human intestinal epithelial cells: a comparison with the Caco-2 cell model.

cDNA microarray technology enables detailed analysis of gene expression throughout complex processes such as differentiation. The aim of this study was to analyze the gene expression profile of normal human intestinal epithelial cells using cell models that recapitulate the crypt-villus axis of intestinal differentiation in comparison with the widely used Caco-2 cell model. cDNA microarrays (19,200 human genes) and a clustering algorithm were used to identify patterns of gene expression in the crypt-like proliferative HIEC and tsFHI cells, and villus epithelial cells as well as Caco-2/15 cells at two distinct stages of differentiation. Unsupervised hierarchical clustering analysis of global gene expression among the cell lines identified two branches: one for the HIEC cells versus a second comprised of two sub-groups: (a) the proliferative Caco-2 cells and (b) the differentiated Caco-2 cells and closely related villus epithelial cells. At the gene level, supervised hierarchical clustering with 272 differentially expressed genes revealed distinct expression patterns specific to each cell phenotype. We identified several upregulated genes that could lead to the identification of new regulatory pathways involved in cell differentiation and carcinogenesis. The combined use of microarray analysis and human intestinal cell models thus provides a powerful tool for establishing detailed gene expression profiles of proliferative to terminally differentiated intestinal cells. Furthermore, the molecular differences between the normal human intestinal cell models and Caco-2 cells clearly point out the strengths and limitations of this widely used experimental model for studying intestinal cell proliferation and differentiation.

Caco-2 Cells↗

Investigations of dipole localization accuracy in MEG using the bootstrap.

We describe the use of the nonparametric bootstrap to investigate the accuracy of current dipole localization from magnetoencephalography (MEG) studies of event-related neural activity. The bootstrap is well suited to the analysis of event-related MEG data since the experiments are repeated tens or even hundreds of times and averaged to achieve acceptable signal-to-noise ratios (SNRs). The set of repetitions or epochs can be viewed as a set of independent realizations of the brain's response to the experiment. Bootstrap resamples can be generated by sampling with replacement from these epochs and averaging. In this study, we applied the bootstrap resampling technique to MEG data from somatotopic experimental and simulated data. Four fingers of the right and left hand of a healthy subject were electrically stimulated, and about 400 trials per stimulation were recorded and averaged in order to measure the somatotopic mapping of the fingers in the S1 area of the brain. Based on single-trial recordings for each finger we performed 5000 bootstrap resamples. We reconstructed dipoles from these resampled averages using the Recursively Applied and Projected (RAP)-MUSIC source localization algorithm. We also performed a simulation for two dipolar sources with overlapping time courses embedded in realistic background brain activity generated using the prestimulus segments of the somatotopic data. To find correspondences between multiple sources in each bootstrap, sample dipoles with similar time series and forward fields were assumed to represent the same source. These dipoles were then clustered by a Gaussian Mixture Model (GMM) clustering algorithm using their combined normalized time series and topographies as feature vectors. The mean and standard deviation of the dipole position and the dipole time series in each cluster were computed to provide estimates of the accuracy of the reconstructed source locations and time series.

Brain Mapping↗

A genetic algorithm for maximum-likelihood phylogeny inference using nucleotide sequence data.

Phylogeny reconstruction is a difficult computational problem, because the number of possible solutions increases with the number of included taxa. For example, for only 14 taxa, there are more than seven trillion possible unrooted phylogenetic trees. For this reason, phylogenetic inference methods commonly use clustering algorithms (e.g., the neighbor-joining method) or heuristic search strategies to minimize the amount of time spent evaluating nonoptimal trees. Even heuristic searches can be painfully slow, especially when computationally intensive optimality criteria such as maximum likelihood are used. I describe here a different approach to heuristic searching (using a genetic algorithm) that can tremendously reduce the time required for maximum-likelihood phylogenetic inference, especially for data sets involving large numbers of taxa. Genetic algorithms are simulations of natural selection in which individuals are encoded solutions to the problem of interest. Here, labeled phylogenetic trees are the individuals, and differential reproduction is effected by allowing the number of offspring produced by each individual to be proportional to that individual's rank likelihood score. Natural selection increases the average likelihood in the evolving population of phylogenetic trees, and the genetic algorithm is allowed to proceed until the likelihood of the best individual ceases to improve over time. An example is presented involving rbcL sequence data for 55 taxa of green plants. The genetic algorithm described here required only 6% of the computational effort required by a conventional heuristic search using tree bisection/reconnection (TBR) branch swapping to obtain the same maximum-likelihood topology.

Algorithms↗

Protein fold comparison by the alignment of topological strings.

Using the definitions of protein folds encoded in a text string, a dynamic programming algorithm was devised to compare these and identify their largest common substructure and calculate the distance (in terms of the number of edit operations) that this lay from each structure. This provided a metric on which the folds were clustered into a 'phylogenetic' tree. This construction differs from previous automatic structure clustering algorithms as it has explicit representation of the structures at 'ancestral' branching nodes, even when these have no corresponding known structure. The resulting tree was compared with that compiled by an 'expert' in the field and while there was broad agreement, differences were found that resulted from differing degrees of emphasis being placed on the types of operations that can be used to transform structures. Some concluding speculations on the relationship of such trees to the evolutionary history and folding of the proteins are advanced.

Computational Biology↗

The ClusNet algorithm and time series prediction.

This paper describes a novel neural network architecture named ClusNet. This network is designed to study the trade-offs between the simplicity of instance-based methods and the accuracy of the more computational intensive learning methods. The features that make this network different from existing learning algorithms are outlined. A simple proof of convergence of the ClusNet algorithm is given. Experimental results showing the convergence of the algorithm on a specific problem is also presented. In this paper, ClusNet is applied to predict the temporal continuation of the Mackey-Glass chaotic time series. A comparison between the results obtained with ClusNet and other neural network algorithms is made. For example, ClusNet requires one-tenth the computing resources of the instance-based local linear method for this application while achieving comparable accuracy in this task. The sensitivity of ClusNet prediction accuracies on specific clustering algorithms is examined for an application. The simplicity and fast convergence of ClusNet makes it ideal as a rapid prototyping tool for applications where on-line learning is required.

Algorithms↗

Selective averaging of evoked potentials using trajectory-based clustering.

A clustering method has been developed to group evoked potentials that display similar prestimulus dynamic behavior. The procedure involves using the method of time delay embedding to construct a trajectory in state space from a time series. Certain features that characterize the geometry of the trajectory have been defined. The trajectory-based clustering algorithm has been applied to visual evoked potentials to determine relationships between prestimulus EEG and evoked potential shape.

Cluster Analysis↗

Design of new selective inhibitors of cyclooxygenase-2 by dynamic assembly of molecular building blocks.

A method of dynamically assembling molecular building blocks - DycoBlock - has been proposed and tested by Liu et al. This method is based on multiple-copy stochastic dynamics simulation in the presence of a receptor molecule. In this method, a novel algorithm was used to dynamically assemble the molecular building blocks to form candidate compounds. Currently, some new improvements have been incorporated into DycoBlock to make it more efficient. In the new version of DycoBlock, the binding energy and solvent accessible surface area (SASA) can be used to screen the resulting compounds. A simple clustering algorithm based on molecular similarity was developed and used to classify the remaining compounds. The revised DycoBlock was tested by breaking SC-558 - a selective inhibitor of cyclooxygenase-2 (COX-2) - into building blocks and reassembling them in the active site of the enzyme. The accuracy of recovery grew to 58.8% while it was only 16.7% in the previous version. Then, thirty-three kinds of molecular building blocks were used in the design of novel inhibitors and the investigation of diversity. As a result, a total of 1441 compounds was generated with high diversity. After the first screening procedure, there remained 864 reasonable compounds. The results from clustering indicate that the structural motifs in the diarylheterocycle class of COX-2-selective inhibitors have been generated using the revised DycoBlock, and their binding modes were investigated.

Algorithms↗

Into the heart of darkness: large-scale clustering of human non-coding DNA.

MOTIVATION: It is currently believed that the human genome contains about twice as much non-coding functional regions as it does protein-coding genes, yet our understanding of these regions is very limited. RESULTS: We examine the intersection between syntenically conserved sequences in the human, mouse and rat genomes, and sequence similarities within the human genome itself, in search of families of non-protein-coding elements. For this purpose we develop a graph theoretic clustering algorithm, akin to the highly successful methods used in elucidating protein sequence family relationships. The algorithm is applied to a highly filtered set of about 700 000 human-rodent evolutionarily conserved regions, not resembling any known coding sequence, which encompasses 3.7% of the human genome. From these, we obtain roughly 12 000 non-singleton clusters, dense in significant sequence similarities. Further analysis of genomic location, evidence of transcription and RNA secondary structure reveals many clusters to be significantly homogeneous in one or more characteristics. This subset of the highly conserved non-protein-coding elements in the human genome thus contains rich family-like structures, which merit in-depth analysis. AVAILABILITY: Supplementary material to this work is available at http://www.soe.ucsc.edu/~jill/dark.html

Animals↗