PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

The ClusNet algorithm and time series prediction.

This paper describes a novel neural network architecture named ClusNet. This network is designed to study the trade-offs between the simplicity of instance-based methods and the accuracy of the more computational intensive learning methods. The features that make this network different from existing learning algorithms are outlined. A simple proof of convergence of the ClusNet algorithm is given. Experimental results showing the convergence of the algorithm on a specific problem is also presented. In this paper, ClusNet is applied to predict the temporal continuation of the Mackey-Glass chaotic time series. A comparison between the results obtained with ClusNet and other neural network algorithms is made. For example, ClusNet requires one-tenth the computing resources of the instance-based local linear method for this application while achieving comparable accuracy in this task. The sensitivity of ClusNet prediction accuracies on specific clustering algorithms is examined for an application. The simplicity and fast convergence of ClusNet makes it ideal as a rapid prototyping tool for applications where on-line learning is required.

Algorithms↗

Selective averaging of evoked potentials using trajectory-based clustering.

A clustering method has been developed to group evoked potentials that display similar prestimulus dynamic behavior. The procedure involves using the method of time delay embedding to construct a trajectory in state space from a time series. Certain features that characterize the geometry of the trajectory have been defined. The trajectory-based clustering algorithm has been applied to visual evoked potentials to determine relationships between prestimulus EEG and evoked potential shape.

Cluster Analysis↗

Design of new selective inhibitors of cyclooxygenase-2 by dynamic assembly of molecular building blocks.

A method of dynamically assembling molecular building blocks - DycoBlock - has been proposed and tested by Liu et al. This method is based on multiple-copy stochastic dynamics simulation in the presence of a receptor molecule. In this method, a novel algorithm was used to dynamically assemble the molecular building blocks to form candidate compounds. Currently, some new improvements have been incorporated into DycoBlock to make it more efficient. In the new version of DycoBlock, the binding energy and solvent accessible surface area (SASA) can be used to screen the resulting compounds. A simple clustering algorithm based on molecular similarity was developed and used to classify the remaining compounds. The revised DycoBlock was tested by breaking SC-558 - a selective inhibitor of cyclooxygenase-2 (COX-2) - into building blocks and reassembling them in the active site of the enzyme. The accuracy of recovery grew to 58.8% while it was only 16.7% in the previous version. Then, thirty-three kinds of molecular building blocks were used in the design of novel inhibitors and the investigation of diversity. As a result, a total of 1441 compounds was generated with high diversity. After the first screening procedure, there remained 864 reasonable compounds. The results from clustering indicate that the structural motifs in the diarylheterocycle class of COX-2-selective inhibitors have been generated using the revised DycoBlock, and their binding modes were investigated.

Algorithms↗

Optimal protein structure alignments by multiple linkage clustering: application to distantly related proteins.

A fully automatic procedure for aligning two protein structures is presented. It uses as sole structural similarity measure the root mean square (r.m.s.) deviation of superimposed backbone atoms (N, C alpha, C and O) and is designed to yield optimal solutions with respect to this measure. In a first step, the procedure identifies protein segments with similar conformations in both proteins. In a second step, a novel multiple linkage clustering algorithm is used to identify segment combinations which yield optimal global structure alignments. Several structure alignments can usually be obtained for a given pair of proteins, which are exploited here to define automatically the common structural core of a protein family. Furthermore, an automatic analysis of the clustering trees is described which enables detection of rigid-body movements between structure elements. To illustrate the performance of our procedure, we apply it to families of distantly related proteins. One groups the three alpha + beta proteins ubiquitin, ferredoxin and the B1-domain of protein G. Their common structure motif consists of four beta-strands and the only alpha-helix, with one strand and the helix being displaced as a rigid body relative to the remaining three beta-strands. The other family consists of beta-proteins from the Greek key group, in particular actinoxanthin, the immunoglobulin variable domain and plastocyanin. Their consensus motif, composed of five beta-strands and a turn, is identified, mostly intact, in all Greek key proteins except the trypsins, and interestingly also in three other beta-protein families, the lipocalins, the neuraminidases and the lectins. This result provides new insights into the evolutionary relationships in the very diverse group of all beta-proteins.

Algorithms↗

5S rRNA sequences of representatives of the genera Chlorobium, Prosthecochloris, Thermomicrobium, Cytophaga, Flavobacterium, Flexibacter and Saprospira and a discussion of the evolution of eubacteria in general.

5S rRNA sequences were determined for the green sulphur bacteria Chlorobium limicola, Chlorobium phaeobacteroides and Prosthecochloris aestuarii, for Thermomicrobium roseum, which is a relative of the green non-sulphur bacteria, and for Cytophaga aquatilis, Cytophaga heparina, Cytophaga johnsonae, Flavobacterium breve, Flexibacter sp. and Saprospira grandis, organisms allotted to the phylum 'Bacteroides-Cytophaga-Flavobacterium' and relatives as determined by 16S rRNA analyses. By using a clustering algorithm a dendrogram was constructed from these sequences and from all other known eubacterial 5S RNA sequences. The dendrogram showed differences, as well as similarities, with respect to results obtained by 16S RNA analyses. The 5S RNA sequences of green sulphur bacteria were closely related to one another, and to a cluster containing 5S RNA sequences from Bacteroides and its relatives, including Cytophaga aquatilis. 5S RNA sequences of all other representatives of the 'Bacteroides-Cytophaga-Flavobacterium' phylum as distinguished by 16S RNA analysis failed to group with Bacteroides and related clusters. On the basis of 5S RNA sequences, Thermomicrobium roseum clustered with Chloroflexus aurantiacus, as was expected from 16S RNA analysis.

Base Sequence↗

Using information theory to discover side chain rotamer classes: analysis of the effects of local backbone structure.

An understanding of the regularities in the side chain conformations of proteins and how these are related to local backbone structures is important for protein modeling and design. Previous work using regular secondary structures and regular divisions of the backbone dihedral angle data has shown that these rotamers are sensitive to the protein's local backbone conformation. In this preliminary study, we demonstrate a method for combining a more general backbone structure model with an objective clustering algorithm to investigate the effects of backbone structures on side chain rotamer classes and distributions. For the local structure classification, we use the Structural Building Blocks (SBB) categories, which represent all types of secondary structure, including regular structures, capping structures, and loops. For classification of side chain data, we use Minimum Message Length (MML) clustering from information theory. We show an example of how MML clustering on data classified by backbone SBBs can reveal different distributions of rotamer classes among the SBBs. Using these preliminary results, some of the characteristics of a rotamer library created using MML clustering on SBB dependent rotamer data are demonstrated.

Computational Biology↗

Molecular classification of breast cancer patients by gene expression profiling.

For many tumors, pathological subclasses exist which have to be further defined by genetic markers to improve therapy and follow-up strategies. In this study, cDNA array analyses of breast cancers have been performed to classify tumors into categories based on expression patterns. Comparing purified normal ductal epithelial cells and corresponding tumour tissues, the expression of only a small fraction of genes was found to be significantly changed. A subset of genes repeatedly found to be differentially expressed in breast cancers was subsequently employed to perform a classification of 82 normal and malignant breast specimens by cluster analysis. This analysis identifies a subgroup of transcriptionally related tumours, designated class A, which can be further subdivided into A1 and A2. Correlation with classical clinicopathological parameters revealed that subgroup A1 was characterized by a high number of node-positive tumours (14 of 16). In this subgroup there was a disproportionate number of patients who had already developed distant metastases at the time of diagnosis (25% in this subgroup, compared with 5% among the rest of the samples). Taken together, the use of these differentially expressed marker genes in conjunction with sample clustering algorithms provides a novel molecular classification of breast cancer specimens, which facilitates the identification of patients with a higher risk of recurrence.

Breast Neoplasms↗

Considerations in applying clustering techniques to speaker-independent word recognition.

Recent work at Bell Laboratories has demonstrated the utility of applying sophisticated pattern recognition techniques to obtain a set of speaker-independent word templates for an isolated word recognition system [Levinson et al.,IEEE Trans. Acoust. Speech Signal Process. ASSP-27 (2), 134--141 (1979); Rabiner et al., IEEE Trans. Acoust. Speech Signal Process.(in press)]. In these studies, it was shown that a careful experimenter could guide the clustering algorithms to choose a small set of templates that were representative of a large number of replications for each word in the vocabulary. Subsequent word recognition tests verified that the templates chosen were indeed representative of a fairly large population of talkers. Given the success of this approach, the next important step is to investigate fully automatic techniques for clustering multiple versions of a single word into a set of speaker-independent word templates. Two such techniques are described in this paper. The first method uses distance data (between replications of a word) to segment the population into stable clusters. The word template is obtained as either the cluster minimax, or as an averaged version of all the elements in the cluster. The second method is a variation of the one described by Rabiner [IEEE Trans. Acoust. Speech Signal Process. ASSP-26 (3), 34--42 (1978)] in which averaging techniques are directly combined with the nearest neighbor rule to simultaneously define both the word template (i.e., the cluster center) and the elements in the cluster. Experimental data show the first method to be superior to the second method when three or more clusters per word are used in the recognition task.

Humans↗

Cluster analysis to improve food classification within commodity groups.

Mathematical clustering algorithms were used to classify foods within dairy, grain, and fat commodity groups on the basis of nutrients with limited availability in the food supply as well as those posing a possible health risk due to excess consumption. The procedure overcomes the problem that has made objective and accurate grouping, i.e., dealing simultaneously with 10 or more nutrients, difficult. The clustering routine classifies foods on the basis of similar nutrient content for any number of food attributes and assigns a degree of association to each food to indicate its compositional similarity to a prototype food for the cluster group. Foods within dairy, grain, and fat commodity groups were clustered on the basis of similar content of vitamin B-6, calcium, iron, magnesium, folacin, zinc, and added sugar, fat, cholesterol, and sodium. Whole milk and natural cheese clustered together on the basis of their moderate nutrient and relatively high fat and sodium content. Whole wheat breads, pumpernickel bread, and pancakes from mix constituted a grain subgroup with highest nutrient content, lowest cholesterol and sugar, lower fat, and higher sodium. Other subgroups based upon similarities in attributes were identified within food commodity categories. The result is an expansion of some food groups to incorporate concepts of both nutritional adequacy and moderation of food components of current nutritional concern.

Dairy Products↗

National Cooperative Growth Study substudy. II: Do growth hormone levels from serial sampling add important diagnostic information?

The National Cooperative Growth Study includes growth data on more the 24,000 children in the United States and Canada who have been treated with growth hormone (GH). To determine whether dysregulation of GH release causes growth failure in children, we initiated the National Cooperative Growth Study substudy II to evaluate the diagnostic utility of serially sampled GH levels and to determine whether those patterns were responsible for the low growth rates in certain subsets of short children and whether children in any of the diagnostic categories would respond to GH therapy. A total of 3744 subjects whose mean height standardized for their chronological age was -2.8 SD and whose pretreatment growth rate was 4.2 cm/yr had complete 12-hour data sets-- serial samples obtained in a 12-hour overnight period. Pulsatile characteristics of GH release were assessed with the cluster algorithm. There was a virtually complete overlap of the GH pulsatile characteristics between control subjects and short children, but the insulin-like growth factor I (IGF-I) levels were markedly lower in the short children, suggesting impairment in the GH-IGF-I axis. THe growth response to administered GH showed only very weak correlations with the various cluster-derived parameters. Our results indicate that one must look beyond the release of GH to find an explanation for the short statures and low IGF-I levels in the subsets of children with idiopathic short stature.

Activity Cycles↗

A fully automatic multimodality image registration algorithm.

OBJECTIVE: A fully automatic multimodality image registration algorithm is presented. The method is primarily designed for 3D registration of MR and PET images of the brain. However, it has also been successfully applied to CT-PET, MR-CT, and MR-SPECT registrations. MATERIALS AND METHODS: The head contour is detected on the MR image using a gradient threshold method. The head region in the MR image is then segmented into a set of connected components using the K-means clustering algorithm. When the two image sets are registered, the segmentation of the MR image indirectly generates a segmentation of the PET image. The best registration is taken to be the one that optimizes the segmentation induced on the PET image. In this article, the K-means minimum variance criterion is used as a cost function, and the optimization is performed using the method of coordinate descent. RESULTS: The algorithm was tested on 80 H2 15O PET and MR image pairs from 10 subjects. Qualitatively correct results were obtained in all cases. With use of external markers visible in both image modalities, the average registration error was estimated to be < 3 mm. CONCLUSION: The algorithm presented in this article requires no user interaction and can be applied to a wide range of registration problems. Quantitative and qualitative evaluations of the algorithm indicate a high degree of accuracy.

Algorithms↗

Identity structure, narrative accounts, and commitment to a volunteer role.

Degree of commitment was explored in relation to core self and role-identity. Thirty-one American emergency medical technicians (EMTs) described themselves in the EMT role (EMT now) and the way they anticipated they would be in the future (EMT future) by selecting items from an adjective checklist. Participants also described "real me," "ideal me," and "ought me." Ratings of commitment and extranormative activity were also obtained. Finally, participants described a positive and a negative episode they had experienced as an EMT in an open-ended question that was coded for task and relational content. Each participant's checklist data set was individually analyzed using HICLAS, a clustering algorithm for binary data (P. DeBoeck, S. Rosenberg, & I. Van Mechelen, 1993). Results indicate that the similarity between EMT now and real me best predicted activity and the similarity between EMT future and real me best predicted commitment (positive correlations in both cases). Older, more experienced EMTs tended to describe positive episodes in relational terms, whereas younger, less experienced EMTs described positive experiences in task-oriented terms.

Adult↗

An algorithm for comparing RNA secondary structures and searching for similar substructures.

To access the functional informations carried by RNA molecules at the level of their secondary structure interactions, we propose a comparison method based on a tree edit algorithm which takes into account the tree structure of RNA foldings. Any secondary structure is translated into a tree involving all its elementary substructures; then a shorter condensed tree is built in which any unbranched helix interspersed with bulges and interior loops is taken as a single node. This method includes several parameters: a comparison matrix between structural units, gap penalties, and the scoring between nodes of the condensed trees. Their effects have been analysed using as a model a rapidly divergent domain of the large ribosomal RNA, for which structural variation during evolution is well known. This method allows one to recognize precisely, in large target molecules, definite substructures that present with the query molecules only a limited set of closely related secondary structure features; it is still efficient if intervening features, which can correspond to insertion/deletion of entire stem regions, separate such structural elements. When coupled with a hierarchical clustering algorithm, this method is suitable for classifying RNA molecules according to their secondary structure homologies.

Algorithms↗

Saturated BLAST: an automated multiple intermediate sequence search used to detect distant homology.

MOTIVATION: Two proteins can have a similar 3-dimensional structure and biological function, but have sequences sufficiently different that traditional protein sequence comparison algorithms do not identify their relationship. The desire to identify such relations has led to the development of more sensitive sequence alignment strategies. One such strategy is the Intermediate Sequence Search (ISS), which connects two proteins through one or more intermediate sequences. In its brute-force implementation, ISS is a strategy that repetitively uses the results of the previous query as new search seeds, making it time-consuming and difficult to analyze. RESULTS: Saturated BLAST is a package that performs ISS in an efficient and automated manner. It was developed using Perl and Perl/Tk and implemented on the LINUX operating system. Starting with a protein sequence, Saturated BLAST runs a BLAST search and identifies representative sequences for the next generation of searches. The procedure is run until convergence or until some predefined criteria are met. Saturated BLAST has a friendly graphic user interface, a built-in BLAST result parser, several multiple alignment tools, clustering algorithms and various filters for the elimination of false positives, thereby providing an easy way to edit, visualize, analyze, monitor and control the search. Besides detecting remote homologies, Saturated BLAST can be used to maintain protein family databases and to search for new genes in genomic databases.

Algorithms↗

A new approach to analysis of synchronized sympathetic nerve activity.

Renal sympathetic nerve activity (RSNA) recorded from the multifiber preparation is a continuously fluctuating variable in terms of period and amplitude, reflecting a coordinated tonic level of output from the vasomotor center. Yet current methods of analysis cannot simultaneously measure both of these parameters. A new accurate technique for assessing changes in global sympathetic activity is required. We made a novel application of a computerized peak detection algorithm (Cluster program) to recordings of synchronized sympathetic nerve discharges. The procedure was applied to this new area to retrieve information about the characteristics of synchronized RSNA. Peaks in synchronized RSNA activity were detected from short-term (20 ms) integrated recordings in which voltage changes had been digitized at 200 Hz and stored on computer. The program scanned the data series for significant increases followed by significant decreases in a small cluster of voltage values. The program permits the input of the cluster sample sizes for the test peaks and pre- and postpeak nadirs and also the minimum height to be defined as a peak. Once each synchronized RSNA peak had been detected, its corresponding amplitude, width, and peak-to-peak interval were calculated. The program successfully characterized RSNA in a group of eight cats and yielded results comparable to other analysis techniques. The peak-to-peak interval period showed two modes of synchronized discharge, one related to the cardiac cycle and a faster 8- to 14-Hz frequency. The synchronized peak amplitude and width showed unimodal frequency distributions. The relationship between each of the three variables was examined; only the peak height and width were significantly related to each other.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Global gene expression analysis of gastric cancer by oligonucleotide microarrays.

To gain molecular understanding of carcinogenesis, progression, and diversity of gastric cancer, 22 primary human advanced gastric cancer tissues and 8 noncancerous gastric tissues were analyzed by high-density oligonucleotide microarray in this study. Based on expression analysis of approximately 6800 genes, a two-way clustering algorithm successfully distinguished cancer tissues from noncancerous tissues. Subsequently, genes that were differentially expressed in cancer and noncancerous tissues were identified; 162 and 129 genes were highly expressed (P < 0.05) >2.5-fold in cancer tissues and noncancerous tissues, respectively. In cancer tissues, genes related to cell cycle, growth factor, cell motility, cell adhesion, and matrix remodeling were highly expressed. In noncancerous tissues, genes related to gastrointestinal-specific function and immune response were highly expressed. Furthermore, we identified several genes associated with lymph node metastasis including Oct-2 or histological types including Liver-Intestine Cadherin. These results provide not only a new molecular basis for understanding biological properties of gastric cancer, but also useful resources for future development of therapeutic targets and diagnostic markers for gastric cancer.

Cluster Analysis↗

A typology of parasuicide.

Parasuicide is not a single syndrome. Subtypes at present recognized are based largely on clinically derived stereotypes. When considering a series of patients, the clinician is unable to handle more than a few attributes at a time. This paper describes the application of three very different clustering algorithms to a material of 350 treated parasuicide patients. Mathematically, three types emerge. Clinically, two of these are interpretable and make sense. The types established are: I (n = 107) a group not characterized by any of the variables we examined; this group is a puzzle, mainly because the reasons for the parasuicidal act are not clear. II (n = 132) a depressed, alienated group with high life-endangerment. III (n = III) a group whose act was highly operant: they felt alienated and were angry with others. These groups did not differ significantly on demographic variables. The usefulness of this typology, particularly for management, after-care and prevention, has now to be assessed.

Anger↗

Alterations in luteinizing hormone secretory activity in women with insulin-dependent diabetes mellitus and secondary amenorrhea.

To investigate hypothalamic and/or pituitary abnormalities in women with poorly controlled insulin-dependent diabetes mellitus (IDDM) and secondary amenorrhea, we measured serum LH every 10 min for 24 h and for 2 additional h after the administration of exogenous GnRH in 8 women with IDDM and amenorrhea and compared these to data from 15 eumenorrheic nondiabetic women. LH pulses were characterized by the pulse detection algorithm Cluster, and secretory episodes were evaluated using the multiple parameter deconvolution procedure Deconv. Cluster analysis revealed fewer LH pulses per 24 h (14.3 +/- 1.2 vs. 19.9 +/- 0.6; P < 0.001; mean +/- SEM), a greater peak width (63 +/- 4.9 vs. 44 +/- 2.2 min; P < 0.01), and greater peak area (136 +/- 17 vs. 89 +/- 13 IU/L.min; P < 0.01) in the diabetic women. Analysis with Deconv revealed fewer LH secretory episodes per 24 h in the diabetic women (14.4 +/- 0.9 vs. 20.4 +/- 0.5; P < 0.001) and no statistical difference in LH half-lives. The IDDM women responded to a 10-micrograms GnRH bolus with LH pulses of larger total (51 +/- 15.9 vs. 15 +/- 1.4 IU/L; P < 0.01) and incremental (29 +/- 7.6 vs. 9 +/- 1.2; P < 0.001) amplitude. In summary, we observed that amenorrheic diabetic women have fewer LH pulses/secretory episodes than normal women. However, they respond well to exogenous GnRH, suggesting that compromise of the GnRH pulse generator, rather than pituitary dysfunction, is responsible for their menstrual dysfunction.

Adult↗