PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Pulsatile thyrotropin release in patients with untreated pituitary disease.

Pulsatile and nocturnal TSH secretion was investigated in 16 healthy controls (group A) and 19 patients with untreated pituitary disease [7 were euthyroid without suprasellar extension (group B), 6 were euthyroid with suprasellar extension (group C) of pituitary lesions, and 6 were hypothyroid with or without suprasellar extension (group D)]. Pulse analysis was performed using Desade and Cluster algorithms. No changes were observed among groups A-D in mean 24-h TSH pulse amplitude [values given as mean +/- SD; Desade, 0.4 +/- 0.2 vs. 0.7 +/- 0.4 vs. 0.6 +/- 0.4 vs. 0.5 +/- 0.2 mU/L (P = NS); Cluster, 0.4 +/- 0.2 vs. 0.7 +/- 0.4 vs. 0.5 +/- 0.3 vs. 0.4 +/- 0.2 mU/L (P = NS)] or in the mean 24-h TSH pulse frequency (approximately 10 pulses/24 h). The mean 24-h TSH concentration was highly correlated to the mean 24-h TSH pulse amplitude in controls (r = 0.93; P < 0.001) and patients (r = 0.63; P < 0.01), but not to the mean 24-h TSH pulse frequency. The nocturnal TSH surge was similar in controls and euthyroid patients without suprasellar extension (group A, 1.0 +/- 0.6; group B, 1.3 +/- 1.3 mU/L; P = NS), but was decreased in euthyroid patients with suprasellar extension (group C, 0.3 +/- 1.0 mU/L; P < 0.05) and hypothyroid patients (group D, 0.4 +/- 0.4 mU/L; P < 0.05). The decreased nocturnal TSH surge was associated with a loss of the usual nocturnal increase in TSH pulse amplitude, whereas the usual nocturnal increase in TSH pulse frequency was maintained. In conclusion, 1) mean 24-h TSH pulse amplitude and frequency are unchanged in untreated patients with pituitary disease; and 2) patients with central hypothyroidism as well as euthyroid patients with suprasellar extension of pituitary lesions had a decreased nocturnal TSH surge associated with a loss of the usual nocturnal increase in TSH amplitude, but not TSH pulse frequency.

Adult↗

Cluster Monte Carlo algorithm for the quantum rotor model.

We propose a highly efficient "worm"-like cluster Monte Carlo algorithm for the quantum rotor model in the link-current representation. We explicitly prove detailed balance for the algorithm even in the presence of disorder. For the pure quantum rotor model with mu=0, the algorithm yields high- precision estimates for the critical point K(c)=0.333 05(5) and the correlation length exponent nu=0.670(3). For the disordered case, mu=1 / 2+/-1 / 2, we find nu=1.15(10).

Journal Article↗

A genetic algorithm for maximum-likelihood phylogeny inference using nucleotide sequence data.

Phylogeny reconstruction is a difficult computational problem, because the number of possible solutions increases with the number of included taxa. For example, for only 14 taxa, there are more than seven trillion possible unrooted phylogenetic trees. For this reason, phylogenetic inference methods commonly use clustering algorithms (e.g., the neighbor-joining method) or heuristic search strategies to minimize the amount of time spent evaluating nonoptimal trees. Even heuristic searches can be painfully slow, especially when computationally intensive optimality criteria such as maximum likelihood are used. I describe here a different approach to heuristic searching (using a genetic algorithm) that can tremendously reduce the time required for maximum-likelihood phylogenetic inference, especially for data sets involving large numbers of taxa. Genetic algorithms are simulations of natural selection in which individuals are encoded solutions to the problem of interest. Here, labeled phylogenetic trees are the individuals, and differential reproduction is effected by allowing the number of offspring produced by each individual to be proportional to that individual's rank likelihood score. Natural selection increases the average likelihood in the evolving population of phylogenetic trees, and the genetic algorithm is allowed to proceed until the likelihood of the best individual ceases to improve over time. An example is presented involving rbcL sequence data for 55 taxa of green plants. The genetic algorithm described here required only 6% of the computational effort required by a conventional heuristic search using tree bisection/reconnection (TBR) branch swapping to obtain the same maximum-likelihood topology.

Algorithms↗

The ClusNet algorithm and time series prediction.

This paper describes a novel neural network architecture named ClusNet. This network is designed to study the trade-offs between the simplicity of instance-based methods and the accuracy of the more computational intensive learning methods. The features that make this network different from existing learning algorithms are outlined. A simple proof of convergence of the ClusNet algorithm is given. Experimental results showing the convergence of the algorithm on a specific problem is also presented. In this paper, ClusNet is applied to predict the temporal continuation of the Mackey-Glass chaotic time series. A comparison between the results obtained with ClusNet and other neural network algorithms is made. For example, ClusNet requires one-tenth the computing resources of the instance-based local linear method for this application while achieving comparable accuracy in this task. The sensitivity of ClusNet prediction accuracies on specific clustering algorithms is examined for an application. The simplicity and fast convergence of ClusNet makes it ideal as a rapid prototyping tool for applications where on-line learning is required.

Algorithms↗

Selective averaging of evoked potentials using trajectory-based clustering.

A clustering method has been developed to group evoked potentials that display similar prestimulus dynamic behavior. The procedure involves using the method of time delay embedding to construct a trajectory in state space from a time series. Certain features that characterize the geometry of the trajectory have been defined. The trajectory-based clustering algorithm has been applied to visual evoked potentials to determine relationships between prestimulus EEG and evoked potential shape.

Cluster Analysis↗

Design of new selective inhibitors of cyclooxygenase-2 by dynamic assembly of molecular building blocks.

A method of dynamically assembling molecular building blocks - DycoBlock - has been proposed and tested by Liu et al. This method is based on multiple-copy stochastic dynamics simulation in the presence of a receptor molecule. In this method, a novel algorithm was used to dynamically assemble the molecular building blocks to form candidate compounds. Currently, some new improvements have been incorporated into DycoBlock to make it more efficient. In the new version of DycoBlock, the binding energy and solvent accessible surface area (SASA) can be used to screen the resulting compounds. A simple clustering algorithm based on molecular similarity was developed and used to classify the remaining compounds. The revised DycoBlock was tested by breaking SC-558 - a selective inhibitor of cyclooxygenase-2 (COX-2) - into building blocks and reassembling them in the active site of the enzyme. The accuracy of recovery grew to 58.8% while it was only 16.7% in the previous version. Then, thirty-three kinds of molecular building blocks were used in the design of novel inhibitors and the investigation of diversity. As a result, a total of 1441 compounds was generated with high diversity. After the first screening procedure, there remained 864 reasonable compounds. The results from clustering indicate that the structural motifs in the diarylheterocycle class of COX-2-selective inhibitors have been generated using the revised DycoBlock, and their binding modes were investigated.

Algorithms↗

Phylogeny of bovine species based on AFLP fingerprinting.

The Bovini species comprise both domestic and wild cattle species. Published phylogenies of this tribe based on mitochondrial DNA contain anomalies, while nuclear sequences show only low variation. We have used amplified fragment length polymorphism (AFLP) fingerprinting in order to detect variation in loci distributed over the nuclear genome. Computer-assisted scoring of electrophoretic fingerprinting patterns yielded 361 markers, which provided sufficient redundancy to suppress stochastic effects of intraspecies polymorphisms and length homoplasies (comigration of non-homologous fragments). Tree reconstructions reveal three clusters: African buffalo with water buffalo, ox with zebu, and bison with wisent. Similarity values suggest a clustering of gaur and banteng, but bifurcating clustering algorithms did not assign consistent positions to these species and yak. We propose that because of shared polymorphisms and reticulations, tree topologies are only partially adequate to represent the phylogeny of the Bovini. Principal-coordinate analysis positions zebu between a gaur/banteng cluster and taurine cattle. This correlates with the region of origin of these species and suggests that genomic distances between the cattle species have been influenced by genetic exchange between neighbouring ancestral populations.

Animals↗

Optimal protein structure alignments by multiple linkage clustering: application to distantly related proteins.

A fully automatic procedure for aligning two protein structures is presented. It uses as sole structural similarity measure the root mean square (r.m.s.) deviation of superimposed backbone atoms (N, C alpha, C and O) and is designed to yield optimal solutions with respect to this measure. In a first step, the procedure identifies protein segments with similar conformations in both proteins. In a second step, a novel multiple linkage clustering algorithm is used to identify segment combinations which yield optimal global structure alignments. Several structure alignments can usually be obtained for a given pair of proteins, which are exploited here to define automatically the common structural core of a protein family. Furthermore, an automatic analysis of the clustering trees is described which enables detection of rigid-body movements between structure elements. To illustrate the performance of our procedure, we apply it to families of distantly related proteins. One groups the three alpha + beta proteins ubiquitin, ferredoxin and the B1-domain of protein G. Their common structure motif consists of four beta-strands and the only alpha-helix, with one strand and the helix being displaced as a rigid body relative to the remaining three beta-strands. The other family consists of beta-proteins from the Greek key group, in particular actinoxanthin, the immunoglobulin variable domain and plastocyanin. Their consensus motif, composed of five beta-strands and a turn, is identified, mostly intact, in all Greek key proteins except the trypsins, and interestingly also in three other beta-protein families, the lipocalins, the neuraminidases and the lectins. This result provides new insights into the evolutionary relationships in the very diverse group of all beta-proteins.

Algorithms↗

5S rRNA sequences of representatives of the genera Chlorobium, Prosthecochloris, Thermomicrobium, Cytophaga, Flavobacterium, Flexibacter and Saprospira and a discussion of the evolution of eubacteria in general.

5S rRNA sequences were determined for the green sulphur bacteria Chlorobium limicola, Chlorobium phaeobacteroides and Prosthecochloris aestuarii, for Thermomicrobium roseum, which is a relative of the green non-sulphur bacteria, and for Cytophaga aquatilis, Cytophaga heparina, Cytophaga johnsonae, Flavobacterium breve, Flexibacter sp. and Saprospira grandis, organisms allotted to the phylum 'Bacteroides-Cytophaga-Flavobacterium' and relatives as determined by 16S rRNA analyses. By using a clustering algorithm a dendrogram was constructed from these sequences and from all other known eubacterial 5S RNA sequences. The dendrogram showed differences, as well as similarities, with respect to results obtained by 16S RNA analyses. The 5S RNA sequences of green sulphur bacteria were closely related to one another, and to a cluster containing 5S RNA sequences from Bacteroides and its relatives, including Cytophaga aquatilis. 5S RNA sequences of all other representatives of the 'Bacteroides-Cytophaga-Flavobacterium' phylum as distinguished by 16S RNA analysis failed to group with Bacteroides and related clusters. On the basis of 5S RNA sequences, Thermomicrobium roseum clustered with Chloroflexus aurantiacus, as was expected from 16S RNA analysis.

Base Sequence↗

Using information theory to discover side chain rotamer classes: analysis of the effects of local backbone structure.

An understanding of the regularities in the side chain conformations of proteins and how these are related to local backbone structures is important for protein modeling and design. Previous work using regular secondary structures and regular divisions of the backbone dihedral angle data has shown that these rotamers are sensitive to the protein's local backbone conformation. In this preliminary study, we demonstrate a method for combining a more general backbone structure model with an objective clustering algorithm to investigate the effects of backbone structures on side chain rotamer classes and distributions. For the local structure classification, we use the Structural Building Blocks (SBB) categories, which represent all types of secondary structure, including regular structures, capping structures, and loops. For classification of side chain data, we use Minimum Message Length (MML) clustering from information theory. We show an example of how MML clustering on data classified by backbone SBBs can reveal different distributions of rotamer classes among the SBBs. Using these preliminary results, some of the characteristics of a rotamer library created using MML clustering on SBB dependent rotamer data are demonstrated.

Computational Biology↗

Molecular classification of breast cancer patients by gene expression profiling.

For many tumors, pathological subclasses exist which have to be further defined by genetic markers to improve therapy and follow-up strategies. In this study, cDNA array analyses of breast cancers have been performed to classify tumors into categories based on expression patterns. Comparing purified normal ductal epithelial cells and corresponding tumour tissues, the expression of only a small fraction of genes was found to be significantly changed. A subset of genes repeatedly found to be differentially expressed in breast cancers was subsequently employed to perform a classification of 82 normal and malignant breast specimens by cluster analysis. This analysis identifies a subgroup of transcriptionally related tumours, designated class A, which can be further subdivided into A1 and A2. Correlation with classical clinicopathological parameters revealed that subgroup A1 was characterized by a high number of node-positive tumours (14 of 16). In this subgroup there was a disproportionate number of patients who had already developed distant metastases at the time of diagnosis (25% in this subgroup, compared with 5% among the rest of the samples). Taken together, the use of these differentially expressed marker genes in conjunction with sample clustering algorithms provides a novel molecular classification of breast cancer specimens, which facilitates the identification of patients with a higher risk of recurrence.

Breast Neoplasms↗

Considerations in applying clustering techniques to speaker-independent word recognition.

Recent work at Bell Laboratories has demonstrated the utility of applying sophisticated pattern recognition techniques to obtain a set of speaker-independent word templates for an isolated word recognition system [Levinson et al.,IEEE Trans. Acoust. Speech Signal Process. ASSP-27 (2), 134--141 (1979); Rabiner et al., IEEE Trans. Acoust. Speech Signal Process.(in press)]. In these studies, it was shown that a careful experimenter could guide the clustering algorithms to choose a small set of templates that were representative of a large number of replications for each word in the vocabulary. Subsequent word recognition tests verified that the templates chosen were indeed representative of a fairly large population of talkers. Given the success of this approach, the next important step is to investigate fully automatic techniques for clustering multiple versions of a single word into a set of speaker-independent word templates. Two such techniques are described in this paper. The first method uses distance data (between replications of a word) to segment the population into stable clusters. The word template is obtained as either the cluster minimax, or as an averaged version of all the elements in the cluster. The second method is a variation of the one described by Rabiner [IEEE Trans. Acoust. Speech Signal Process. ASSP-26 (3), 34--42 (1978)] in which averaging techniques are directly combined with the nearest neighbor rule to simultaneously define both the word template (i.e., the cluster center) and the elements in the cluster. Experimental data show the first method to be superior to the second method when three or more clusters per word are used in the recognition task.

Humans↗

Cluster analysis to improve food classification within commodity groups.

Mathematical clustering algorithms were used to classify foods within dairy, grain, and fat commodity groups on the basis of nutrients with limited availability in the food supply as well as those posing a possible health risk due to excess consumption. The procedure overcomes the problem that has made objective and accurate grouping, i.e., dealing simultaneously with 10 or more nutrients, difficult. The clustering routine classifies foods on the basis of similar nutrient content for any number of food attributes and assigns a degree of association to each food to indicate its compositional similarity to a prototype food for the cluster group. Foods within dairy, grain, and fat commodity groups were clustered on the basis of similar content of vitamin B-6, calcium, iron, magnesium, folacin, zinc, and added sugar, fat, cholesterol, and sodium. Whole milk and natural cheese clustered together on the basis of their moderate nutrient and relatively high fat and sodium content. Whole wheat breads, pumpernickel bread, and pancakes from mix constituted a grain subgroup with highest nutrient content, lowest cholesterol and sugar, lower fat, and higher sodium. Other subgroups based upon similarities in attributes were identified within food commodity categories. The result is an expansion of some food groups to incorporate concepts of both nutritional adequacy and moderation of food components of current nutritional concern.

Dairy Products↗

Molecular characterization of some Indian Basmati and other elite rice genotypes using fluorescent-AFLP.

Cultivated rice is a high-volume, low-value cereal crop providing staple food to more than 50% of the world populace. A small group of rice cultivars, traditionally produced on the Indo-Gangetic plains and popularly known as Basmati, have exquisite quality grain characteristics and are a prized commercial commodity. Efforts to improve the yield potential of Basmati have led to the development of several crossbred Basmati-like cultivars. In this study we have analysed the genetic diversity and interrelationships among 33 rice genotypes consisting of the traditional Basmati, improved Basmati-like genotypes developed in India and elsewhere, American long-grain rice and a few non-aromatic rice using a DNA marker-based approach - fluorescent-amplified fragment length polymorphism (f-AFLP). Using a set of nine primer-pairs we scored a total of 10,672 data points over all of the genotypes in the size range of 75-500 bp. The scored data points corresponded to a total of 501 AFLP markers (putative loci/genome landmarks) of which 327 markers (65%) were polymorphic. The f-AFLP marker data, which were analysed using different clustering algorithms and principal component analysis, indicate that: (1) considerable genetic variability exists in the analysed genotypes; (2) traditional Basmati cultivars could be distinctly separated from the crossbred Basmati-like genotypes as well as from the non-aromatic rice; (3) the crossbred Basmati-like cultivars from the subcontinent and elsewhere are genetically very distinct; (4) f-AFLP-based clustering, in general, conforms to the putative pedigree of the improved genotypes. Moreover, analysis to ascertain the scope of AFLP as a technique suggests that the polymorphism revealed by three selective primer-pair combinations is sufficient to obtain reliable estimates of genetic diversity for the type of material used in this study. However, its utility to identify group-specific DNA markers was discounted due to a low frequency of observed group-specific discrete markers.

Journal Article↗

National Cooperative Growth Study substudy. II: Do growth hormone levels from serial sampling add important diagnostic information?

The National Cooperative Growth Study includes growth data on more the 24,000 children in the United States and Canada who have been treated with growth hormone (GH). To determine whether dysregulation of GH release causes growth failure in children, we initiated the National Cooperative Growth Study substudy II to evaluate the diagnostic utility of serially sampled GH levels and to determine whether those patterns were responsible for the low growth rates in certain subsets of short children and whether children in any of the diagnostic categories would respond to GH therapy. A total of 3744 subjects whose mean height standardized for their chronological age was -2.8 SD and whose pretreatment growth rate was 4.2 cm/yr had complete 12-hour data sets-- serial samples obtained in a 12-hour overnight period. Pulsatile characteristics of GH release were assessed with the cluster algorithm. There was a virtually complete overlap of the GH pulsatile characteristics between control subjects and short children, but the insulin-like growth factor I (IGF-I) levels were markedly lower in the short children, suggesting impairment in the GH-IGF-I axis. THe growth response to administered GH showed only very weak correlations with the various cluster-derived parameters. Our results indicate that one must look beyond the release of GH to find an explanation for the short statures and low IGF-I levels in the subsets of children with idiopathic short stature.

Activity Cycles↗

A fully automatic multimodality image registration algorithm.

OBJECTIVE: A fully automatic multimodality image registration algorithm is presented. The method is primarily designed for 3D registration of MR and PET images of the brain. However, it has also been successfully applied to CT-PET, MR-CT, and MR-SPECT registrations. MATERIALS AND METHODS: The head contour is detected on the MR image using a gradient threshold method. The head region in the MR image is then segmented into a set of connected components using the K-means clustering algorithm. When the two image sets are registered, the segmentation of the MR image indirectly generates a segmentation of the PET image. The best registration is taken to be the one that optimizes the segmentation induced on the PET image. In this article, the K-means minimum variance criterion is used as a cost function, and the optimization is performed using the method of coordinate descent. RESULTS: The algorithm was tested on 80 H2 15O PET and MR image pairs from 10 subjects. Qualitatively correct results were obtained in all cases. With use of external markers visible in both image modalities, the average registration error was estimated to be < 3 mm. CONCLUSION: The algorithm presented in this article requires no user interaction and can be applied to a wide range of registration problems. Quantitative and qualitative evaluations of the algorithm indicate a high degree of accuracy.

Algorithms↗

Unique gene expression profiles of human macrophages and dendritic cells to phylogenetically distinct parasites.

Monocyte-derived dendritic cells (DCs) and macrophages (Ms) generated in vitro from the same individual blood donors were exposed to 5 different pathogens, and gene expression profiles were assessed by microarray analysis. Responses to Mycobacterium tuberculosis and to phylogenetically distinct protozoan (Leishmania major, Leishmania donovani, Toxoplasma gondii) and helminth (Brugia malayi) parasites were examined, each of which produces chronic infections in humans yet vary considerably in the nature of the immune responses they trigger. In the absence of microbial stimulation, DCs and Ms constitutively expressed approximately 4000 genes, 96% of which were shared between the 2 cell types. In contrast, the genes altered transcriptionally in DCs and Ms following pathogen exposure were largely cell specific. Profiling of the gene expression data led to the identification of sets of tightly coregulated genes across all experimental conditions tested. A newly devised literature-based clustering algorithm enabled the identification of functionally and transcriptionally homogenous groups of genes. A comparison of the responses induced by the individual pathogens by means of this strategy revealed major differences in the functionally related gene profiles associated with each infectious agent. Although the intracellular pathogens induced responses clearly distinct from the extracellular B malayi, they each displayed a unique pattern of gene expression that would not necessarily be predicted on the basis of their phylogenetic relationship. The association of characteristic functional clusters with each infectious agent is consistent with the concept that antigen-presenting cells have prewired signaling patterns for use in the response to different pathogens.

Animals↗

Selection of surrogate marker genes in primary central nervous system lymphomas for radio-chemotherapy by DNA array analysis of gene expression profiles.

Primary central nervous system lymphomas (PCNSLs) are extra nodal B-cell non-Hodgkin's lymphomas with primary manifestation in the brain, and their incidence has been increasing among both immunocompetent and immunocompromised populations. Samples of oligodendroglioma (n=5), glioblastoma (n=7), PCNSL (n=6), and normal brain (n=3) were studied (total of 21 samples) using cDNA array technology. The hierarchical clustering algorithm was used to obtain a phylogenetic tree, and it revealed a striking feature: PCNSL was clearly separated. The genes encoding laminin receptor 2, thioredoxin peroxidase, and elongation factor-1 were selected as specific genes in PCNSL by principal component analysis (PCA). When Mann-Whitney tests were performed to identify genes responsible for the differences between responders and non-responders to the treatment schedule for PCNSL, 76 known genes were found to show significantly different expression patterns between the two groups at the P<0.01 level. The two groups were clearly separated by the re-clustering method using the selected genes related to response to chemo-radiotherapy. This is the first report describing the gene expression profiles of PCNSL. In conclusion, accumulation of data with respect to the expression profiles of PCNSL specimens, clinicopathological data, susceptibility to treatment, and outcome will provide information for identifying optimal therapeutic modalities for individual patients and novel therapeutic targets.

Adult↗