PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cluster analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Statistical shape analysis: clustering, learning, and testing.

Using a differential-geometric treatment of planar shapes, we present tools for: 1) hierarchical clustering of imaged objects according to the shapes of their boundaries, 2) learning of probability models for clusters of shapes, and 3) testing of newly observed shapes under competing probability models. Clustering at any level of hierarchy is performed using a mimimum variance type criterion criterion and a Markov process. Statistical means of clusters provide shapes to be clustered at the next higher level, thus building a hierarchy of shapes. Using finite-dimensional approximations of spaces tangent to the shape space at sample means, we (implicitly) impose probability models on the shape space, and results are illustrated via random sampling and classification (hypothesis testing). Together, hierarchical clustering and hypothesis testing provide an efficient framework for shape retrieval. Examples are presented using shapes and images from ETH, Surrey, and AMCOM databases.

Algorithms↗

Chromosomal regions in prostatic carcinomas studied by comparative genomic hybridization, hierarchical cluster analysis and self-organizing feature maps.

Comparative genomic hybridization (CGH) is an established genetic method which enables a genome-wide survey of chromosomal imbalances. For each chromosome region, one obtains the information whether there is a loss or gain of genetic material, or whether there is no change at that place. Therefore, large amounts of data quickly accumulate which must be put into a logical order. Cluster analysis can be used to assign individual cases (samples) to different clusters of cases, which are similar and where each cluster may be related to a different tumour biology. Another approach consists in a clustering of chromosomal regions by rewriting the original data matrix, where the cases are written as rows and the chromosomal regions as columns, in a transposed form. In this paper we applied hierarchical cluster analysis as well as two implementations of self-organizing feature maps as classical and neuronal tools for cluster analysis of CGH data from prostatic carcinomas to such transposed data sets. Self-organizing maps are artificial neural networks with the capability to form clusters on the basis of an unsupervised learning rule. We studied a group of 48 cases of incidental carcinomas, a tumour category which has not been evaluated by CGH before. In addition we studied a group of 50 cases of pT2N0-tumours and a group of 20 pT3N0-carcinomas. The results show in all case groups three clusters of chromosomal regions, which are (i) normal or minimally affected by losses and gains, (ii) regions with many losses and few gains and (iii) regions with many gains and few losses. Moreover, for the pT2N0- and pT3N0-groups, it could be shown that the regions 6q, 8p and 13q lay all on the same cluster (associated with losses), and that the regions 9q and 20q belonged to the same cluster (associated with gains). For the incidental cancers such clear correlations could not be demonstrated.

Chromosome Mapping↗

[Dynamic pattern of gene expression in rat hearts responding to transient ischemia/reperfusion detected by cDNA microarray and cluster analysis].

OBJECTIVE: To analyze the pattern of gene expression programs during rat myocardial ischemia/reperfusion at multiple time points. METHODS: Rat model of myocardial ischemia/reperfusion was established by repeating the occlusion and relaxation of left coronary artery. cDNA microarray was used to analyze the pattern of gene expression programs in rat myocardium at 1, 3, 6, 12, 24 hours after reperfusion. SOM cluster analysis was used to identify different cluster of genes in which each cluster had similar expression pattern. RESULTS: Altogether 75, 779, 205, 155, and 166 genes were differentially expressed at 1, 3, 6, 12, 24 hours after the reperfusion respectively. Clusters analysis identified 12 clusters of genes in which each cluster had similar expression pattern. CONCLUSION: Analysis of gene expression pattern revealed sequential induction of subsets of genes that characterize each response.

Animals↗

Classification of primate spinothalamic and somatosensory thalamic neurons based on cluster analysis.

Data analyzed in this study were derived from the responses of 128 spinothalamic tract (STT) cells and 110 thalamic neurons recorded in 75 anesthetized monkeys. A k-means cluster analysis, a nonhierarchical clustering technique, was performed using the relative magnitudes of responses to a graded series of innocuous and noxious mechanical stimuli applied to the receptive field. For comparison, a parallel analysis was performed based on definitions of low-threshold (LT), wide dynamic range (WDR), and high-threshold (HT) cells used by our laboratory. For 128 STT cells, a classification scheme with three clusters was found statistically to be the best. This yielded groups of 22, 57, and 49 cells in clusters 1, 2, and 3, respectively. Cluster 1 cells were activated best by low-intensity mechanical stimuli, whereas cluster 3 cells were activated primarily by nociceptive stimuli. Cluster 2 cells had intermediate characteristics. When the classification scheme based on the cluster analysis was compared with the classification of the same neurons as LT, WDR, and HT cells, cluster 1 cells were divided into LT and WDR cells, whereas cluster 2 and 3 cells included WDR and HT cells. For 110 thalamic neurons, a classification scheme with five clusters was found statistically to be the best. Clusters 1-5 contained 25, 34, 17, 10, and 24 cells, respectively. Response characteristics of cells in each group indicated a gradual change in sensitivity to higher intensities of peripheral input from cluster 1 to 5. When this classification scheme was compared with the classification scheme previously used by our laboratory, cluster 1 cells belonged to the LT group, clusters 2 and 3 split into LT and WDR cells, and clusters 4 and 5 included WDR and HT cells. It is concluded that a classification scheme based on a cluster analysis of the responses of neurons to standardized stimuli may provide an objective and functionally meaningful way to categorize somatosensory neurons.

Animals↗

Spectral analysis of EEG in the late course of primary generalized myoclonic-astatic epilepsy. II. Cluster analysis of the power spectra.

Cluster analysis is applied to power spectra of the EEG of 38 patients with a primary generalized myoclonic astatic epilepsy (Gundel et al. 1981). The tendency of the data to form clusters within this group is indicated by a random experiment which has been performed with the data. The clustering algorithm divided the material in seven distinct groups which may be combined to three main types of power spectra. These types are power spectra with a 10 cps peak, a 4-7 cps peak and power spectra without remarkable rhythmization. The comparison of EEG types and clinical data shows a correlation of 4-7 cps rhythms with the occurrence of seizures. 4-7 cps rhythms are interpreted as a symptom of "centrencephalic" convulsibility.

Adolescent↗

[Clinical characteristics of seafood allergy and classification of 10 seafood allergens by cluster analysis].

The aim of this study was to investigate the clinical characteristics of children who showed sensitization to any type of seafood and to classify the 10 seafood allergens based on IgE reactivities by a cluster analysis. In children with bronchial asthma (BA) and/or atopic dermatitis (AD), we defined the 'seafood' group as 23 patients having any type of seafood specific IgE antibody (CAP system). We analyzed the clinical features, the serum total IgE and each allergen specific IgE level. In addition, ten seafood allergens were also classified by a cluster analysis. Three patients revealed immediate hypersensitivity to some seafood. The frequency of patients with AD and the total IgE in the seafood group were high and the patients in this group were tend to be sensitized to multiple allergens. Seafood allergens were classified into 4 groups, 1) salmon, sardine, horse mackerel and mackerel, 2) cod and tuna, 3) octopus and squid, and 4) crab and shrimp, by a cluster analysis. These findings corresponded to the biological classification and the classification by the reported common allergens among various types of seafood. Based on our findings, this classification is therefore considered to be useful when selecting allergens to screen for sensitization to seafood.

Allergens↗

Identification of "binge-prone" women: an experimentally and psychometrically validated cluster analysis in a college population.

This study investigated the escape model of binge eating through a cluster analysis using standardized measures. A sample of 126 undergraduate women underwent a manipulation of their level of cognition and were asked to "taste-test" several flavors of ice cream. Questionnaire data from these women were entered into a cluster analysis. Two groups emerged: women in the "binge-prone" group were significantly more depressed, had lower self-esteem, had more chaotic and extreme eating patterns, and were more self-conscious than those in the control group. In validation work, binge-prone women were shown to report elevated levels of bulimic symptomatology and, when in the presence of a food they enjoyed, to respond to increases in level of cognition by eating more. These results were consistent with some, but not all, of the components of the escape model.

Adult↗

[Cluster analysis of restriction fragment length polymorphism patterns of Mycobacterium tuberculosis isolated in Chiba prefecture].

Methods for cluster analysis of IS6110 based restriction fragment length polymorphism (RFLP) patterns of Mycobacterium tuberculosis isolates were studied for an epidemiological investigation in Chiba prefecture. To normalize patterns, external size markers were adopted instead of typical internal size markers used in the standard method. RFLP patterns were run on 1.4% agarose gels and external markers were applied to outside and middle lanes on each gel for precise comparison. The resulting RFLP patterns of 74 isolates were clustered by similarity. Similarity was calculated with the Dice coefficient using parameter settings at 0.8% tolerance and 0.5% optimization. Patterns of 19 isolates from 8 outbreaks showed high similarity within each outbreak. Cluster analysis, as described here, provides insights into epidemiological tracing of tuberculosis in Chiba prefecture.

Cluster Analysis↗

Structural symmetry of the extracellular domain of the cytokine/growth hormone/prolactin receptor family and interferon receptors revealed by hydrophobic cluster analysis.

Sequence comparison based on Hydrophobic Cluster Analysis procedures shows that the extracellular approximately 200 amino acids domains of cytokines receptors belonging to the Cytokine/Growth hormone/Prolactin receptor family and to the Interferon one are organized in two homologous subdomains. Further, comparison of the subdomains of 32 independent sequences and of a lot of already recognized homologous domains with data bases could lead to the hypothesis that these approximately 100 amino acids subdomains could possess the overall fold of the constant immunoglobulin domains and so could belong to the immunoglobulin superfamily.

Amino Acid Sequence↗

Hierarchical cluster analysis as an approach for systematic grouping of diet constituents on basis of fatty acid, energy and cholesterol content: application on consumable lamb products.

The role of dietary fat in the etiology of chronic diseases is both a qualitative and a quantitative issue. The dietary fat intake is largely influenced by behavioral and social influences on food choice. Ongoing scientific research has led to dietary recommendations with main concerns being the percentage of saturated, essential fatty acids and cholesterol with respect to total energy intake. However, the compositional complexity of food choice constituting the diet is a critical concept complicating the interpretation of epidemiologic, clinical and laboratory evidence to define the role of dietary fat in the etiology of diseases. This study was conducted on the observation of the need to better systematically classify consumable food based on complex composition and lamb meat is randomly selected as a non-specific subset for application of hierarchical cluster analysis method to obtain the dendogram using average linkage. Data on fat composition of consumable lamb prepared by different methods was obtained from USDA Nutrient Database for Standart Reference. Using agglomerative hierarchical cluster analysis lamb meat was grouped into two main clusters among which one divided into two families of which each was subdivided into two subfamilies based on fatty acids, cholesterol and energy composition. Present work may be considered as a leading study to systematically classify larger food sets. As high fat foods are rich in flavor and overall palatability, the outcome of this study may lead to behaviorally more acceptable but healthier dietary replacements. Besides future use of the results obtained may reveal the effect of complex compositional dietary influences on health and disease and may have superiority to studies questioning individual dietary items. Furthermore, hieararchial cluster analysis may be used to cluster food including other compositional data in food items like amino acids, vitamins, carbohydrates, as well.

Animals↗

Cluster analysis of levels of body fatness in children.

The purpose of this study was to examine, using cluster analysis, the levels of body fatness as defined in the current national programs for children and youth fitness. A total of 1,056 examinees were drawn randomly from the published data of the National Children and Youth Fitness Study II, with 525 boys and 531 girls, ages 6 to 9 years. Their triceps and medial calf skinfold measures were used for the cluster analysis, including both the kth nearest-neighbor and the Ward's minimum variance procedures. Although multimodal clusters were found at four age groups according to the kth nearest-neighbor procedure, the Ward procedure and plotting of these clusters did not support their existence. It was concluded that the levels of body fatness reported in the current national children and youth fitness programs were arbitrarily defined.

Child↗

Identification of homogeneous geographical areas of mortality for tumours from cluster analysis.

This paper attempts to demonstrate the utility of cluster analysis as a descriptive method of studying mortality in epidemiology. In order to verify which algorithms of clustering best fit the data structure, the method of cophenetic correlation was implemented. Furthermore the probabilistic algorithm proposed by Beale was used to assess the partition. The results show the presence of some striking clusters between Local Sanitary Units of the Emilia Romagna Region for four types of tumour in men.

Algorithms↗

[Cluster analysis study on the marshland of Schistosoma Japonicum using satellite TM image data in Peng Lake, Jiangxi province].

OBJECTIVE: To create a category land cover map of the marshland region of Schistosoma Japonicum using satellite TM data. METHODS: TM satellite images from Peng Lake of Jiangxi province were applied to the Cluster analysis. Then the resulting clusters were identified and reclassified by undertaking site visits. RESULTS: Eight land cover classes were generated, including Carexspp Zone that was the snail habitat place. CONCLUSION: Cluster analysis, which is a technique for the interpretation of remotely sensed imagery, could contribute to the study on the distribution of snail habitats and become a new tool for the epidemiological ecology study.

Animals↗

Investigating stress effect patterns in hospital staff nurses: results of a cluster analysis.

A comprehensive and reliable assessment of work stress, burnout, affective, and physical symptomatology was conducted with 260 hospital nurses. As previous attempts to categorize nursing stress and burnout by ward type have yielded inconsistent results, an alternative method for grouping nursing stress effects was sought. Cluster analysis was chosen as it offers a statistically sound means of delineating natural groupings within data. Sets of questionnaires measuring burnout, work stressors, and physical and emotional symptomatology were sent to all staff nurses at a large university hospital. Of 709 nurses employed there, a total of 260 nurses returned completed questionnaire packets. These nurses were separated into two equal groups using random sampling procedures. Cluster analysis of this data revealed groupings which were based on nursing stressors (particularly workload and conflict with physicians), social support, and patient loads. These cluster-analytic findings were replicated on both samples, and validated using data not used in the original cluster analysis. Results suggest that the effects of stress have more to do with the characteristics of the work environment and overall workload than with the degree of specialization on the unit. Results also suggest that intraprofessional conflict (i.e. with other nurses) is less psychologically damaging than is interprofessional conflict (i.e. conflict with physicians). Findings are discussed with respect to the burnout process and possible interventions.

Adult↗

Cluster analysis as selection and dereplication tool for the identification of new natural compounds from large sample sets.

Cluster analysis of gas-chromatographic (GC) data of ca. 500 bacterial isolates was used as an aid in detection and identification of new natural compounds. This approach reduces the number of GC/MS analysis (dereplication) and concomitantly improves the selection of samples with high probability to contain unknown natural products. Lipophilic bacterial extracts were derivatized and analyzed by GC under standardized conditions. A program was developed to convert chromatographic data into a two-dimensional matrix. Based on the results of hierarchical cluster analysis samples were selected for further investigation by GC/MS and NMR. This approach avoided unnecessary analysis of similar samples. By this method, the unusual oligoprenylsesquiterpenes 1 and 2 as well as new aromatic amides 7 and 8 were identified.

Bacteria↗

Molecular classification of borderline ovarian tumors using hierarchical cluster analysis of protein expression profiles.

Ovarian tumors range from benign to aggressive malignant tumors, including an intermediate class referred to as borderline carcinoma. The prognosis of the disease is strongly dependent on tumor classification, where patients with borderline tumors have much better prognosis than patients with carcinomas. We here describe the use of hierarchical clustering analysis of quantitative protein expression data for classification of this type of tumor. An accurate classification was not achieved using an unselected set of 1,584 protein spots for clustering analysis. Different approaches were used to select spots that were differentially expressed between tumors of different malignant potential and to use these sets of spots for classification. When sets of proteins were selected that differentiated benign and malignant tumors, borderline tumors clustered in the benign group. This is consistent with the biologic properties of these tumors. Our results indicate that hierarchical clustering analysis is a useful approach for analysis of protein profiles and show that this approach can be used for differential diagnosis of ovarian carcinomas and borderline tumors.

Adult↗

Cluster analysis and three-dimensional QSAR studies of HIV-1 integrase inhibitors.

Three-dimensional quantitative structure-activity relationship (3D QSAR) and cluster analysis were applied to a variety of HIV-1 integrase inhibitors. One structure was chosen from each of 11 classes of inhibitors to represent the whole class in descriptor-based cluster analysis. The 11 classes of inhibitors were classified into two groups. The molecular field analysis (MFA) models for these two clusters had r2 values of 0.90 and 0.95 and q2 values of 0.85 and 0.91 that were noticeably enhanced from those of conventional QSAR models. The five test compounds, which were proposed to have a common binding site near the metal in HIV-1 integrase based on docking studies by Sotriffer et al., were utilized to compare the predictive capability of MFA and conventional QSAR models. Among these five compounds, only L-chicoric acid belongs to cluster 1 and the other four belong to cluster 2. MFA models give better overall predictions and more importantly the activity of these test compounds is better predicted by the MFA model derived from the cluster each test compound belongs to. The necessity of dividing the inhibitors into two groups to obtain predictive QSAR models supports the likelihood of two separate binding sites.

Cluster Analysis↗