PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cluster analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Identifying eating patterns in male and female undergraduates using cluster analysis.

Although there is substantial research on eating patterns, little effort has been paid to developing a classification of eating behavior applicable to the general population, rather than to people seeking help for obesity or eating disorders. Using cluster analysis, this study identified six types of eating patterns among normal volunteers. One hundred and sixteen females and 70 males completed a questionnaire concerning weight history, food intake patterns, use of satiation cues, and attitudes toward weight gain. Subjects also completed the Restraint Scale (Herman & Polivy, 1975). Height and weight were measured. Factor analysis reduced the questionnaire to nine internally reliable and meaningful scales; these were then entered into a K-means cluster analysis of subjects. Of the six clusters, two represented mild forms of disordered eating, two could be considered to represent more regulated eating styles, and two were distinguished by differential sensitivity to internal satiation cues. Construct validity of clusters was explored against gender, degree of overweight and scores on the Restraint Scale. Discussion focuses on the values of a cluster analytic technique to identify multidimensional patterns of food intake.

Adult↗

Hierarchical cluster analysis for exposure assessment of workers in the Semiconductor Health Study.

The fabrication of integrated circuits in the semiconductor industry involves worker exposures to multiple chemical and physical agents. The potential for a high degree of correlation among exposure variables was of concern in the Semiconductor Health Study. Hierarchical cluster analysis was used to identify groups or "clusters" of correlated variables. Several variations of hierarchical cluster analysis were performed on 14 chemical and physical agents, using exposure data on 882 subjects from the historical cohort of the epidemiological studies. Similarity between agent pairs was determined by calculating two metrics of dissimilarity, and hierarchical trees were constructed using three clustering methods. Among subjects exposed to ethylene-based glycol ethers (EGE), xylene, or n-butyl acetate (nBA), 83% were exposed to EGE and xylene, 86% to EGE and nBA, and 94% to xylene and nBA, suggesting that exposures to EGE, xylene, and nBA were highly correlated. A high correlation was also found for subjects exposed to boron and phosphorus (80%). The trees also revealed cluster groups containing agents associated with work-group exposure categories developed for the epidemiologic analyses.

Cluster Analysis↗

[Cluster analysis on the temporal dynamics of arthropod community in a plum orchard].

In this paper, twelve investigations were conducted on the structural feature of arthropod community in a plum orchard, and a cluster analysis was made on the temporal dynamics of the community. The total community could be clustered into 5 clusters (D = 0.2000), i.e., that in March, in June, in July, in November, and in other months. The natural enemy and pest-neutral sub-communities could be also clustered into 5 clusters, respectively. For natural enemy sub-community, the clusters (D = 0.2000) were that in March, in July, in August, in September and October, and in other months, and for pest-neutral sub-community, they (D = 0.1000) were that in April, in July, in June, in November, and in other months. The results of cluster analysis partly reflected the seasonal differences of total community and sub-communities, while the temporal overlaps of cluster results reflected the complexity of community structure. Based on the optimization cut-apart, both the total community and the sub-communities were divided into 5 stages, i.e., 6 April as the first stage, 27 April to 8 June as the second stage, 27 June to 27 August as the third stage, 21 September to 19 October as the forth stage, and 22 November as the fifth stage.

Animals↗

Hip problems in older adults: classification by cluster analysis.

No validated classification system of hip disorders in primary care is available. This study explores whether it is possible to obtain such a classification with the method of cluster analyses. A total of 224 consecutive patients aged 50 years or older, consulting the general practitioner for pain in the hip region, and referred for X-ray investigation of the hip, underwent a standardized examination. Ward's cluster analysis with variables from history and physical examination of the hip region resulted in a classification with nine different clusters. These clusters were reproduced in 10 random subsamples and with an alternative cluster analysis. Significant relationships of various external variables (radiological and sonographic signs and variables of low-back and knee examination) with the distinctive clusters were found. Twenty of the approached experts recognized the symptoms in seven clusters as identifiable syndromes. However, further validation of the achieved classification system, especially with respect to the clinical importance, is needed before introducing it into clinical practice.

Cluster Analysis↗

VISCANA: visualized cluster analysis of protein-ligand interaction based on the ab initio fragment molecular orbital method for virtual ligand screening.

We have developed a visualized cluster analysis of protein-ligand interaction (VISCANA) that analyzes the pattern of the interaction of the receptor and ligand on the basis of quantum theory for virtual ligand screening. Kitaura et al. (Chem. Phys. Lett. 1999, 312, 319-324.) have proposed an ab initio fragment molecular orbital (FMO) method by which large molecules such as proteins can be easily treated with chemical accuracy. In the FMO method, a total energy of the molecule is evaluated by summation of fragment energies and interfragment interaction energies (IFIEs). In this paper, we have proposed a cluster analysis using the dissimilarity that is defined as the squared Euclidean distance between IFIEs of two ligands. Although the result of an ordered table by clustering is still a massive collection of numbers, we combine a clustering method with a graphical representation of the IFIEs by representing each data point with colors that quantitatively and qualitatively reflect the IFIEs. We applied VISCANA to a docking study of pharmacophores of the human estrogen receptor alpha ligand-binding domain (57 amino acid residues). By using VISCANA, we could classify even structurally different ligands into functionally similar clusters according to the interaction pattern of a ligand and amino acid residues of the receptor protein. In addition, VISCANA could estimate the correct docking conformation by analyzing patterns of the receptor-ligand interactions of some conformations through the docking calculation.

Binding Sites↗

[Cluster analysis of mirror type internal symmetry in amino acid sequences of heterotrimeric G-protein alpha-subunit superfamily].

For the identification of mirror type internal symmetry centers in amino acid sequences (AASs) the new method, named by the method of internal symmetry scanning, was developed. The method, contrary to earlier ones, can be used for analysis of large clusters of primary structures of related proteins. The internal symmetry centres, containing both one and two amino acid residues, can be identified rapidly and effectively by the method. Additionally, the new method allow to estimate quantitatively the homology of AASs, which are antiparallel to relation of the centres. The different modifications of the method can be used for revealing of both high conservative and unequal symmetrical structures in AASs of proteins. Usually the structures coincide with functionally important regions of protein molecules. The method was used for investigation of primary structures of members of heterotrimeric G-protein a a-subunit superfamily. The positive correlation between conservativity of primary structure and distribution of mirror type internal symmetry centres was shown.

Amino Acid Sequence↗

Cluster analysis in kinetic modelling of the brain: a noninvasive alternative to arterial sampling.

In emission tomography, quantification of brain tracer uptake, metabolism or binding requires knowledge of the cerebral input function. Traditionally, this is achieved with arterial blood sampling. We propose a noninvasive alternative via the use of a blood vessel time-activity curve (TAC) extracted directly from dynamic positron emission tomography (PET) scans by cluster analysis. Five healthy subjects were injected with the 5HT(2A)-receptor ligand [(18)F]-altanserin and blood samples were subsequently taken from the radial artery and cubital vein. Eight regions-of-interest (ROI) TACs were extracted from the PET data set. Hierarchical K-means cluster analysis was performed on the PET time series to extract a cerebral vasculature ROI. The number of clusters was varied from K = 1 to 10 for the second of the two-stage method. Determination of the correct number of clusters was performed by the 'within-variance' measure and by 3D visual inspection of the homogeneity of the determined clusters. The cluster-determined input curve was then used in Logan plot analysis and compared with the arterial and venous blood samples, and additionally with one of the currently used alternatives to arterial blood sampling, the Simplified Reference Tissue Model (SRTM) and Logan analysis with cerebellar TAC as an input. There was a good agreement (P < 0.05) between the values of Distribution Volume (DV) obtained from the K-means-clustered input function and those from the arterial blood samples. This work acts as a proof-of-principle that the use of cluster analysis on a PET data set could obviate the requirement for arterial cannulation when determining the input function for kinetic modelling of ligand binding, and that this may be a superior approach as compared to the other noninvasive alternatives.

Adult↗

[The computer image processing and cluster analysis on Fructus Cnidii from different habitats].

OBJECTIVE: To study the variation of shape of Fructus Cnidii from different habitats. METHODS: The methods of computer image processing and cluster analysis were adopted. RESULTS: The shape of Fructus Cnidii from different habitats varied significantly. CONCLUSION: According to the results of computer image processing and cluster analysis, Fructus Cnidii can be classified into three types.

China↗

Cluster analysis of childhood temperament data on adoptees.

Cluster analysis is used with behavioral data on 162 adoptees to assess temperament and to test the validity of the temperament typologies described by Thomas and Chess. Data are divided into two groups by sex. Results concur with the Thomas-Chess findings in identifying three main temperament groups: difficult, easy, and slow to warm up. Membership in the difficult group predicted later childhood behavior disorder in both sexes. Differing environmental factors associated with difficulty for males and for females are considered, and directions for further investigation are suggested.

Adolescent↗

A comparison of cluster analysis methods using DNA methylation data.

MOTIVATION: Aberrant DNA methylation is common in cancer. DNA methylation profiles differ between tumor types and subtypes and provide a powerful diagnostic tool for identifying clusters of samples and/or genes. DNA methylation data obtained with the quantitative, highly sensitive MethyLight technology is not normally distributed; it frequently contains an excess of zeros. Established tools to analyze this type of data do not exist. Here, we evaluate a variety of methods for cluster analysis to determine which is most reliable. RESULTS: We introduce a Bernoulli-lognormal mixture model for clustering DNA methylation data obtained using MethyLight. We model the outcomes using a two-part distribution having discrete and continuous components. It is compared with standard cluster analysis approaches for continuous data and for discrete data. In a simulation study, we find that the two-part model has the lowest classification error rate for mixture outcome data compared with other approaches. The methods are illustrated using DNA methylation data from a study of lung cancer cell lines. Compared with competing hierarchical clustering methods, the mixture model approaches have the lowest cross-validation error for detecting lung cancer subtype (non-small versus small cell). The Bernoulli-lognormal mixture assigns observations to subgroups with the lowest uncertainty. AVAILABILITY: Software is available upon request from the authors. SUPPLEMENTARY INFORMATION: http://www-rcf.usc.edu/~kims/SupplementaryInfo.html

Algorithms↗

Computerized three-dimensional localization of prostate cancer using contrast-enhanced power Doppler and clustering analysis.

OBJECTIVE: To evaluate the potential benefit of semiautomated localization of prostate cancer using clustering analysis on three-dimensional (3-D) contrast-enhanced power Doppler images. METHODS: Thirty patients with biopsy-proven prostate cancer and scheduled for radical prostatectomy underwent a 3-D contrast-enhanced power Doppler scan prior to surgery. A 3-D ellipsoid model was manually fitted around the prostate. The model automatically divided the prostate into 12 zones. After calculation of a so-called clustering map, the clustering values of each zone were calculated. They were compared with whole-mount section histopathology. Region-of-interest (ROI) analysis was performed with bootstrapping to evaluate overall performance. RESULTS: The ROI analysis yielded area under the curve (AUC) values of 0.65 with a corresponding standard error of 0.03. CONCLUSION: Semiautomatic localization based on clustering analysis of blood flow aids in localization of prostate tumors. A clustering map is an easy-to-interpret extension to standard power Doppler images.

Aged↗

Analyzing microarray data using cluster analysis.

As pharmacogenetics researchers gather more detailed and complex data on gene polymorphisms that effect drug metabolizing enzymes, drug target receptors and drug transporters, they will need access to advanced statistical tools to mine that data. These tools include approaches from classical biostatistics, such as logistic regression or linear discriminant analysis, and supervised learning methods from computer science, such as support vector machines and artificial neural networks. In this review, we present an overview of another class of models, cluster analysis, which will likely be less familiar to pharmacogenetics researchers. Cluster analysis is used to analyze data that is not a priori known to contain any specific subgroups. The goal is to use the data itself to identify meaningful or informative subgroups. Specifically, we will focus on demonstrating the use of distance-based methods of hierarchical clustering to analyze gene expression data.

Cluster Analysis↗

Application of gravitational clustering analysis to liquid gastric emptying.

A gravitational clustering analysis was applied to principal component weighting factors and t1/2 values from 100 liquid gastric emptying studies. Using principal components, groups of patients with clinical features in common were identified, whereas the analysis of the t1/2 values failed to differentiate them.

Duodenal Ulcer↗

Cluster analysis of clinical data to identify subtypes within a study population following treatment with a new pentapeptide antidepressant.

Cluster analysis was used to evaluate the data from a placebo-controlled, double-blind clinical trial with a new pentapeptide antidepressant (INN 00835) in major depression. The objective of this paper is to examine the effect of separating the study population into homogeneous subgroups (clusters) with relatively similar response to treatment within subgroups, and significantly different response between subgroups. The list of variables for cluster analysis was selected only from the efficacy parameters investigated in the study. Three to six clusters were modelled to obtain the optimal number of clusters, based on a proportional contribution of subjects per cluster, and the maximum statistical difference between clusters. After separation, the variability of response among drug-treated subjects by cluster was attributed to plasma drug concentration. Platelet serotonin uptake, which is a putative biochemical marker of effective treatment of depression, also reproduced the same effect of separation as the initially established cluster variables.

Journal Article↗

Prognosis of hemolytic anemia in G6PD- subjects. Multifactorial cluster analysis of biochemical characteristics of red cell age groups.

Individual susceptibility of 10 G6PD- hemizygotes to oxidative hemolytic agents was tested on the basis of multifactorial cluster analysis of biochemical indices of erythrocyte populations; the indices related to G6PD activity and glucose metabolism were analyzed under physiological and oxidative stress conditions in very young, exactly adult and very old red cell suspensions. Biochemical images of G6PD- erythrocytes were obtained and compared with the donor (7 subjects) biochemical image on a IBM-PC computer according to a special "taxon" program. As a result, a stable subdivision of 10 Gd- biochemical images into 5 taxons was formed; each taxon included G6PD subjects with a certain form of clinical appearance of G6PD deficiency. Multifactorial cluster analysis of biochemical data on the erythrocyte population allows a clinical prognosis for G6PD- subjects.

Anemia, Hemolytic↗

Gene expression in rat leydig cells during development from the progenitor to adult stage: a cluster analysis.

The postnatal development of Leydig cells can be divided into three distinct stages: initially they exist as fibroblast-like progenitor Leydig cells (PLCs) appearing in the testis by Days 14-21; subsequently, by Day 35, they become immature Leydig cells (ILCs) acquiring steroidogenic organelle structure and enzyme activities but metabolizing most of the testosterone they produce; finally, as adult Leydig cells (ALCs) by Day 90, they actively produce testosterone. The factors controlling proliferation and differentiation of Leydig cells remain largely unknown, and the aim of the present study was to identify changes in gene expression during development through cDNA array analysis of PLCs, ILCs, and ALCs. By cluster analysis, it was determined that the transitions from PLC to ILC to ALC were associated with downregulation of mRNAs corresponding to 107 genes. The downregulated genes included cell-cycle regulators, e.g., cyclin D1 (Ccnd1); growth factors, e.g., basic fibroblast growth factor (Fgf2); growth-factor-related receptors, e.g., platelet-derived growth factor alpha receptor (Pdgfra); oncogenes, e.g., kit oncogene (Kit); and transcription factors, e.g., early growth response 1 (Egr1). Conversely, expression levels of 264 genes were increased by at least twofold. Most of these were related to differentiated function and included steroidogenic enzymes, e.g., 11beta-hydroxysteroid dehydrogenase 2 (Hsd11b2); neurotransmitter receptors, e.g., acetylcholine receptor nicotinic alpha 4 (Chrna4); stress response factors, e.g., glutathione transferase 8 (Gsta4); and protein turnover enzymes, e.g., tissue inhibitor of metalloproteinase 2 (Timp2). The detection of Hsd11b2 mRNA in the array was the first indication that this gene is expressed in Leydig cells, and parallel increases in Hsd11b2 mRNA and enzyme activity were recorded. Thus, gene profiling demonstrates that postnatal development is associated with changes in the expression levels of several different clusters of genes consistent with the processes of Leydig cell growth and differentiation.

11-beta-Hydroxysteroid Dehydrogenase Type 2↗

Population recovery capabilities of 35 cluster analysis methods.

Comparative evaluation of population recovery capabilities of 35 cluster analysis methods defined by different combinations of 5 profile similarity measures and 7 agglomeration rules was undertaken using artificial data that represented duplicate mixture samples from 4 latent populations. The latent population mean profiles differed primarily in elevation or in pattern parameters. Latent population sampling variances were controlled to provide two different levels of realistic overlap. The within-population distributions were multivariate normal with diagonal covariance structure. Across all conditions examined, complete linkage and Ward's minimum variance methods, used with Euclidian or city block interprofile distance measures, performed best. Single linkage, median, and centroid methods were substantially inferior for clustering individuals in accordance with true population memberships.

Analysis of Variance↗

Hydrophobic-cluster analysis of plant protein sequences. A domain homology between storage and lipid-transfer proteins.

Hydrophobic-cluster analysis was used to characterize a conserved domain located near the C-terminal amino acid sequence of wheat (Triticum aestivum) storage proteins. This domain was transformed into a linear template for a global search for similarities in over 5200 protein sequences. In addition to proteins that had already been found to exhibit homology to wheat storage proteins, a previously unreported homology was found with non-specific lipid-transfer proteins from castor bean (Ricinus communis) and from spinach (Spinacia oleracea) leaf. Hydrophobic-cluster analysis of various members of the present protein group clearly shows a typical domain structure where (i) variable and conserved domains are located along the sequence at precise positions, (ii) the conserved domains probably reflect a common ancestor, and (iii) the unique properties of a given protein (chain cut into subunits, repetitive domains, trypsin-inhibitor active site) are associated with the variable domains.

Amino Acid Sequence↗