PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

ASK-GraphView: A large scale graph visualization system.

We describe ASK-GraphView, a node-link-based graph visualization system that allows clustering and interactive navigation of large graphs, ranging in size up to 16 million edges. The system uses a scalable architecture and a series of increasingly sophisticated clustering algorithms to construct a hierarchy on an arbitrary, weighted undirected input graph. By lowering the interactivity requirements we can scale to substantially bigger graphs. The user is allowed to navigate this hierarchy in a top down manner by interactively expanding individual clusters. ASK-GraphView also provides facilities for filtering and coloring, annotation and cluster labeling.

Journal Article↗

Phylogeny of bovine species based on AFLP fingerprinting.

The Bovini species comprise both domestic and wild cattle species. Published phylogenies of this tribe based on mitochondrial DNA contain anomalies, while nuclear sequences show only low variation. We have used amplified fragment length polymorphism (AFLP) fingerprinting in order to detect variation in loci distributed over the nuclear genome. Computer-assisted scoring of electrophoretic fingerprinting patterns yielded 361 markers, which provided sufficient redundancy to suppress stochastic effects of intraspecies polymorphisms and length homoplasies (comigration of non-homologous fragments). Tree reconstructions reveal three clusters: African buffalo with water buffalo, ox with zebu, and bison with wisent. Similarity values suggest a clustering of gaur and banteng, but bifurcating clustering algorithms did not assign consistent positions to these species and yak. We propose that because of shared polymorphisms and reticulations, tree topologies are only partially adequate to represent the phylogeny of the Bovini. Principal-coordinate analysis positions zebu between a gaur/banteng cluster and taurine cattle. This correlates with the region of origin of these species and suggests that genomic distances between the cattle species have been influenced by genetic exchange between neighbouring ancestral populations.

Animals↗

Red-bond exponents of the critical and the tricritical Ising model in three dimensions.

Using the Wolff and geometric cluster algorithms and finite-size scaling analysis, we investigate the critical Ising and the tricritical Blume-Capel models with nearest-neighbor interactions on the simple-cubic lattice. The sampling procedure involves the decomposition of the Ising configuration into geometric clusters, each of which consists of a set of nearest-neighboring spins of the same sign connected with bond probability p. These clusters include the well-known Kasteleyn-Fortuin clusters as a special case for p=1-exp(-2K) , where K is the Ising spin-spin coupling. Along the critical line K=Kc , the size distribution of geometric clusters is investigated as a function of p . We observe that, unlike in the case of two-dimensional tricriticality, the percolation threshold in both models lies at pc =1-exp(-2Kc) . Further, we determine the corresponding red-bond exponents as yr =0.757(2) and 0.501(5) for the critical Ising and the tricritical Blume-Capel models, respectively. On this basis, we conjecture yr =1/2 for the latter model.

Journal Article↗

Optimal protein structure alignments by multiple linkage clustering: application to distantly related proteins.

A fully automatic procedure for aligning two protein structures is presented. It uses as sole structural similarity measure the root mean square (r.m.s.) deviation of superimposed backbone atoms (N, C alpha, C and O) and is designed to yield optimal solutions with respect to this measure. In a first step, the procedure identifies protein segments with similar conformations in both proteins. In a second step, a novel multiple linkage clustering algorithm is used to identify segment combinations which yield optimal global structure alignments. Several structure alignments can usually be obtained for a given pair of proteins, which are exploited here to define automatically the common structural core of a protein family. Furthermore, an automatic analysis of the clustering trees is described which enables detection of rigid-body movements between structure elements. To illustrate the performance of our procedure, we apply it to families of distantly related proteins. One groups the three alpha + beta proteins ubiquitin, ferredoxin and the B1-domain of protein G. Their common structure motif consists of four beta-strands and the only alpha-helix, with one strand and the helix being displaced as a rigid body relative to the remaining three beta-strands. The other family consists of beta-proteins from the Greek key group, in particular actinoxanthin, the immunoglobulin variable domain and plastocyanin. Their consensus motif, composed of five beta-strands and a turn, is identified, mostly intact, in all Greek key proteins except the trypsins, and interestingly also in three other beta-protein families, the lipocalins, the neuraminidases and the lectins. This result provides new insights into the evolutionary relationships in the very diverse group of all beta-proteins.

Algorithms↗

Co-clustering and visualization of gene expression data and gene ontology terms for Saccharomyces cerevisiae using self-organizing maps.

We propose a novel co-clustering algorithm that is based on self-organizing maps (SOMs). The method is applied to group yeast (Saccharomyces cerevisiae) genes according to both expression profiles and Gene Ontology (GO) annotations. The combination of multiple databases is supposed to provide a better biological definition and separation of gene clusters. We compare different levels of genome-wide co-clustering by weighting the involved sources of information differently. Clustering quality is determined by both general and SOM-specific validation measures. Co-clustering relies on a sufficient correlation between the different datasets. We investigate in various experiments how much GO information is contained in the applied gene expression dataset and vice versa. The second major contribution is a visualization technique that applies the cluster structure of SOMs for a better biological interpretation of gene (expression) clusterings. Our GO term maps reveal functional neighborhoods between clusters forming biologically meaningful functional SOM regions. To cope with the high variety and specificity of GO terms, gene and cluster annotations are mapped to a reduced vocabulary of more general GO terms. In particular, this advances the ability of SOMs to act as gene function predictors.

Artificial Intelligence↗

Reuse of imputed data in microarray analysis increases imputation efficiency.

BACKGROUND: The imputation of missing values is necessary for the efficient use of DNA microarray data, because many clustering algorithms and some statistical analysis require a complete data set. A few imputation methods for DNA microarray data have been introduced, but the efficiency of the methods was low and the validity of imputed values in these methods had not been fully checked. RESULTS: We developed a new cluster-based imputation method called sequential K-nearest neighbor (SKNN) method. This imputes the missing values sequentially from the gene having least missing values, and uses the imputed values for the later imputation. Although it uses the imputed values, the efficiency of this new method is greatly improved in its accuracy and computational complexity over the conventional KNN-based method and other methods based on maximum likelihood estimation. The performance of SKNN was in particular higher than other imputation methods for the data with high missing rates and large number of experiments. Application of Expectation Maximization (EM) to the SKNN method improved the accuracy, but increased computational time proportional to the number of iterations. The Multiple Imputation (MI) method, which is well known but not applied previously to microarray data, showed a similarly high accuracy as the SKNN method, with slightly higher dependency on the types of data sets. CONCLUSIONS: Sequential reuse of imputed data in KNN-based imputation greatly increases the efficiency of imputation. The SKNN method should be practically useful to save the data of some microarray experiments which have high amounts of missing entries. The SKNN method generates reliable imputed values which can be used for further cluster-based analysis of microarray data.

Efficiency, Organizational↗

MRI fuzzy segmentation of brain tissue using neighborhood attraction with neural-network optimization.

Image segmentation is an indispensable process in the visualization of human tissues, particularly during clinical analysis of magnetic resonance (MR) images. Unfortunately, MR images always contain a significant amount of noise caused by operator performance, equipment, and the environment, which can lead to serious inaccuracies with segmentation. A robust segmentation technique based on an extension to the traditional fuzzy c-means (FCM) clustering algorithm is proposed in this paper. A neighborhood attraction, which is dependent on the relative location and features of neighboring pixels, is shown to improve the segmentation performance dramatically. The degree of attraction is optimized by a neural-network model. Simulated and real brain MR images with different noise levels are segmented to demonstrate the superiority of the proposed technique compared to other FCM-based methods. This segmentation method is a key component of an MR image-based classification system for brain tumors, currently being developed. Index Terms-Improved fuzzy c-means clustering (IFCM), magnetic resonance imaging (MRI), neighborhood attraction, segmentation.

Algorithms↗

5S rRNA sequences of representatives of the genera Chlorobium, Prosthecochloris, Thermomicrobium, Cytophaga, Flavobacterium, Flexibacter and Saprospira and a discussion of the evolution of eubacteria in general.

5S rRNA sequences were determined for the green sulphur bacteria Chlorobium limicola, Chlorobium phaeobacteroides and Prosthecochloris aestuarii, for Thermomicrobium roseum, which is a relative of the green non-sulphur bacteria, and for Cytophaga aquatilis, Cytophaga heparina, Cytophaga johnsonae, Flavobacterium breve, Flexibacter sp. and Saprospira grandis, organisms allotted to the phylum 'Bacteroides-Cytophaga-Flavobacterium' and relatives as determined by 16S rRNA analyses. By using a clustering algorithm a dendrogram was constructed from these sequences and from all other known eubacterial 5S RNA sequences. The dendrogram showed differences, as well as similarities, with respect to results obtained by 16S RNA analyses. The 5S RNA sequences of green sulphur bacteria were closely related to one another, and to a cluster containing 5S RNA sequences from Bacteroides and its relatives, including Cytophaga aquatilis. 5S RNA sequences of all other representatives of the 'Bacteroides-Cytophaga-Flavobacterium' phylum as distinguished by 16S RNA analysis failed to group with Bacteroides and related clusters. On the basis of 5S RNA sequences, Thermomicrobium roseum clustered with Chloroflexus aurantiacus, as was expected from 16S RNA analysis.

Base Sequence↗

Forecasting epilepsy from the heart rate signal.

Information contained in the R-R interval series, specific to the pre-ictal period, was sought by applying an unsupervised fuzzy clustering algorithm to the N-dimensional phase space of N consecutive interval durations or the absolute value of duration differences. Data sources were individual, complex partial seizures of temporal-lobe epileptics and generalised seizures of rats rendered epileptic with hyperbaric oxygen. Forecasting success was 86% and 82% (zero false positives in resistant rats), respectively, at times ranging from 10 min to 30 s prior to seizure onset Although certain forecasting clusters predominated in the patient group and different ones predominated in the animal group, forecasting on the whole was seizure-specific. The high prediction sensitivity of this method, which matches that of EEG-based methods, seems promising. It is believed that an on-line version of the algorithm, trained on each patient's peri-ictal ECG, could serve as a basis for a simple seizure alarm system.

Algorithms↗

Extracting gene networks for low-dose radiation using graph theoretical algorithms.

Genes with common functions often exhibit correlated expression levels, which can be used to identify sets of interacting genes from microarray data. Microarrays typically measure expression across genomic space, creating a massive matrix of co-expression that must be mined to extract only the most relevant gene interactions. We describe a graph theoretical approach to extracting co-expressed sets of genes, based on the computation of cliques. Unlike the results of traditional clustering algorithms, cliques are not disjoint and allow genes to be assigned to multiple sets of interacting partners, consistent with biological reality. A graph is created by thresholding the correlation matrix to include only the correlations most likely to signify functional relationships. Cliques computed from the graph correspond to sets of genes for which significant edges are present between all members of the set, representing potential members of common or interacting pathways. Clique membership can be used to infer function about poorly annotated genes, based on the known functions of better-annotated genes with which they share clique membership (i.e., "guilt-by-association"). We illustrate our method by applying it to microarray data collected from the spleens of mice exposed to low-dose ionizing radiation. Differential analysis is used to identify sets of genes whose interactions are impacted by radiation exposure. The correlation graph is also queried independently of clique to extract edges that are impacted by radiation. We present several examples of multiple gene interactions that are altered by radiation exposure and thus represent potential molecular pathways that mediate the radiation response.

Algorithms↗

Using information theory to discover side chain rotamer classes: analysis of the effects of local backbone structure.

An understanding of the regularities in the side chain conformations of proteins and how these are related to local backbone structures is important for protein modeling and design. Previous work using regular secondary structures and regular divisions of the backbone dihedral angle data has shown that these rotamers are sensitive to the protein's local backbone conformation. In this preliminary study, we demonstrate a method for combining a more general backbone structure model with an objective clustering algorithm to investigate the effects of backbone structures on side chain rotamer classes and distributions. For the local structure classification, we use the Structural Building Blocks (SBB) categories, which represent all types of secondary structure, including regular structures, capping structures, and loops. For classification of side chain data, we use Minimum Message Length (MML) clustering from information theory. We show an example of how MML clustering on data classified by backbone SBBs can reveal different distributions of rotamer classes among the SBBs. Using these preliminary results, some of the characteristics of a rotamer library created using MML clustering on SBB dependent rotamer data are demonstrated.

Computational Biology↗

Molecular classification of breast cancer patients by gene expression profiling.

For many tumors, pathological subclasses exist which have to be further defined by genetic markers to improve therapy and follow-up strategies. In this study, cDNA array analyses of breast cancers have been performed to classify tumors into categories based on expression patterns. Comparing purified normal ductal epithelial cells and corresponding tumour tissues, the expression of only a small fraction of genes was found to be significantly changed. A subset of genes repeatedly found to be differentially expressed in breast cancers was subsequently employed to perform a classification of 82 normal and malignant breast specimens by cluster analysis. This analysis identifies a subgroup of transcriptionally related tumours, designated class A, which can be further subdivided into A1 and A2. Correlation with classical clinicopathological parameters revealed that subgroup A1 was characterized by a high number of node-positive tumours (14 of 16). In this subgroup there was a disproportionate number of patients who had already developed distant metastases at the time of diagnosis (25% in this subgroup, compared with 5% among the rest of the samples). Taken together, the use of these differentially expressed marker genes in conjunction with sample clustering algorithms provides a novel molecular classification of breast cancer specimens, which facilitates the identification of patients with a higher risk of recurrence.

Breast Neoplasms↗

Considerations in applying clustering techniques to speaker-independent word recognition.

Recent work at Bell Laboratories has demonstrated the utility of applying sophisticated pattern recognition techniques to obtain a set of speaker-independent word templates for an isolated word recognition system [Levinson et al.,IEEE Trans. Acoust. Speech Signal Process. ASSP-27 (2), 134--141 (1979); Rabiner et al., IEEE Trans. Acoust. Speech Signal Process.(in press)]. In these studies, it was shown that a careful experimenter could guide the clustering algorithms to choose a small set of templates that were representative of a large number of replications for each word in the vocabulary. Subsequent word recognition tests verified that the templates chosen were indeed representative of a fairly large population of talkers. Given the success of this approach, the next important step is to investigate fully automatic techniques for clustering multiple versions of a single word into a set of speaker-independent word templates. Two such techniques are described in this paper. The first method uses distance data (between replications of a word) to segment the population into stable clusters. The word template is obtained as either the cluster minimax, or as an averaged version of all the elements in the cluster. The second method is a variation of the one described by Rabiner [IEEE Trans. Acoust. Speech Signal Process. ASSP-26 (3), 34--42 (1978)] in which averaging techniques are directly combined with the nearest neighbor rule to simultaneously define both the word template (i.e., the cluster center) and the elements in the cluster. Experimental data show the first method to be superior to the second method when three or more clusters per word are used in the recognition task.

Humans↗

Cluster analysis to improve food classification within commodity groups.

Mathematical clustering algorithms were used to classify foods within dairy, grain, and fat commodity groups on the basis of nutrients with limited availability in the food supply as well as those posing a possible health risk due to excess consumption. The procedure overcomes the problem that has made objective and accurate grouping, i.e., dealing simultaneously with 10 or more nutrients, difficult. The clustering routine classifies foods on the basis of similar nutrient content for any number of food attributes and assigns a degree of association to each food to indicate its compositional similarity to a prototype food for the cluster group. Foods within dairy, grain, and fat commodity groups were clustered on the basis of similar content of vitamin B-6, calcium, iron, magnesium, folacin, zinc, and added sugar, fat, cholesterol, and sodium. Whole milk and natural cheese clustered together on the basis of their moderate nutrient and relatively high fat and sodium content. Whole wheat breads, pumpernickel bread, and pancakes from mix constituted a grain subgroup with highest nutrient content, lowest cholesterol and sugar, lower fat, and higher sodium. Other subgroups based upon similarities in attributes were identified within food commodity categories. The result is an expansion of some food groups to incorporate concepts of both nutritional adequacy and moderation of food components of current nutritional concern.

Dairy Products↗

Molecular characterization of some Indian Basmati and other elite rice genotypes using fluorescent-AFLP.

Cultivated rice is a high-volume, low-value cereal crop providing staple food to more than 50% of the world populace. A small group of rice cultivars, traditionally produced on the Indo-Gangetic plains and popularly known as Basmati, have exquisite quality grain characteristics and are a prized commercial commodity. Efforts to improve the yield potential of Basmati have led to the development of several crossbred Basmati-like cultivars. In this study we have analysed the genetic diversity and interrelationships among 33 rice genotypes consisting of the traditional Basmati, improved Basmati-like genotypes developed in India and elsewhere, American long-grain rice and a few non-aromatic rice using a DNA marker-based approach - fluorescent-amplified fragment length polymorphism (f-AFLP). Using a set of nine primer-pairs we scored a total of 10,672 data points over all of the genotypes in the size range of 75-500 bp. The scored data points corresponded to a total of 501 AFLP markers (putative loci/genome landmarks) of which 327 markers (65%) were polymorphic. The f-AFLP marker data, which were analysed using different clustering algorithms and principal component analysis, indicate that: (1) considerable genetic variability exists in the analysed genotypes; (2) traditional Basmati cultivars could be distinctly separated from the crossbred Basmati-like genotypes as well as from the non-aromatic rice; (3) the crossbred Basmati-like cultivars from the subcontinent and elsewhere are genetically very distinct; (4) f-AFLP-based clustering, in general, conforms to the putative pedigree of the improved genotypes. Moreover, analysis to ascertain the scope of AFLP as a technique suggests that the polymorphism revealed by three selective primer-pair combinations is sufficient to obtain reliable estimates of genetic diversity for the type of material used in this study. However, its utility to identify group-specific DNA markers was discounted due to a low frequency of observed group-specific discrete markers.

Journal Article↗

National Cooperative Growth Study substudy. II: Do growth hormone levels from serial sampling add important diagnostic information?

The National Cooperative Growth Study includes growth data on more the 24,000 children in the United States and Canada who have been treated with growth hormone (GH). To determine whether dysregulation of GH release causes growth failure in children, we initiated the National Cooperative Growth Study substudy II to evaluate the diagnostic utility of serially sampled GH levels and to determine whether those patterns were responsible for the low growth rates in certain subsets of short children and whether children in any of the diagnostic categories would respond to GH therapy. A total of 3744 subjects whose mean height standardized for their chronological age was -2.8 SD and whose pretreatment growth rate was 4.2 cm/yr had complete 12-hour data sets-- serial samples obtained in a 12-hour overnight period. Pulsatile characteristics of GH release were assessed with the cluster algorithm. There was a virtually complete overlap of the GH pulsatile characteristics between control subjects and short children, but the insulin-like growth factor I (IGF-I) levels were markedly lower in the short children, suggesting impairment in the GH-IGF-I axis. THe growth response to administered GH showed only very weak correlations with the various cluster-derived parameters. Our results indicate that one must look beyond the release of GH to find an explanation for the short statures and low IGF-I levels in the subsets of children with idiopathic short stature.

Activity Cycles↗

RGFGA: an efficient representation and crossover for grouping genetic algorithms.

There is substantial research into genetic algorithms that are used to group large numbers of objects into mutually exclusive subsets based upon some fitness function. However, nearly all methods involve degeneracy to some degree. We introduce a new representation for grouping genetic algorithms, the restricted growth function genetic algorithm, that effectively removes all degeneracy, resulting in a more efficient search. A new crossover operator is also described that exploits a measure of similarity between chromosomes in a population. Using several synthetic datasets, we compare the performance of our representation and crossover with another well known state-of-the-art GA method, a strawman optimisation method and a well-established statistical clustering algorithm, with encouraging results.

Algorithms↗

Protein array autoantibody profiles for insights into systemic lupus erythematosus and incomplete lupus syndromes.

The objective of this study was to investigate the prevalence and clinical significance of a spectrum of autoantibodies in systemic lupus erythematosus and incomplete lupus syndromes using a proteome microarray bearing 70 autoantigens. Microarrays containing candidate autoantigens or control proteins were printed on 16-section slides. These arrays were used to profile 93 serum samples from patients with systemic lupus erythematosus (SLE (n = 33), incomplete LE (ILE; n = 23), first-degree relatives (FDRs) of SLE patients (n = 20) and non-autoimmune controls (NC; n = 17). Data were analysed using the significance analysis of microarray (SAM) and clustering algorithms. Correlations with disease features were determined. Serum from ILE and SLE patients contained high levels of IgG autoantibodies to 50 autoantigens and IgM autoantibodies to 12 autoantigens. Elevated levels of at least one IgG autoantibody were detected in 26% of SLE and 19% of ILE samples; elevated IgM autoantibodies were present in 13% of SLE and 17% of ILE samples. IgG autoantibodies segregated into seven clusters including two specific for DNA and RNA autoantigens that were correlated with the number of lupus criteria. Three IgG autoantibody clusters specific for collagens, DNA and histones, were correlated with renal involvement. Of the four IgM autoantibody clusters, two were correlated negatively with the number of lupus criteria; none were correlated with renal disease. The IgG : IgM autoantibody ratios generally showed a stepwise increase in the groups following disease burden from NC to SLE. Insights derived from the expanded autoantibody profiling made possible with the antigen array suggest differences in autoreactivity in ILE and SLE. Determining whether the IgM aurotreactivity that predominates in ILE represents an early stage prior to IgG switching or is persistent and relatively protective will require further longitudinal studies.

Adult↗