PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cluster analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Robust cluster analysis of microarray gene expression data with the number of clusters determined biologically.

MOTIVATION: The success of each method of cluster analysis depends on how well its underlying model describes the patterns of expression. Outlier-resistant and distribution-insensitive clustering of genes are robust against violations of model assumptions. RESULTS: A measure of dissimilarity that combines advantages of the Euclidean distance and the correlation coefficient is introduced. The measure can be made robust using a rank order correlation coefficient. A robust graphical method of summarizing the results of cluster analysis and a biological method of determining the number of clusters are also presented. These methods are applied to a public data set, showing that rank-based methods perform better than log-based methods. AVAILABILITY: Software is available from http://www.davidbickel.com.

Algorithms↗

Cluster analysis of dynamic parameters of gene expression.

Cluster analysis has proven to be a valuable statistical method for analyzing whole genome expression data. Although clustering methods have great utility, they do represent a lower level statistical analysis that is not directly tied to a specific model. To extend such methods and to allow for more sophisticated lines of inference, we use cluster analysis in conjunction with a specific model of gene expression dynamics. This model provides phenomenological dynamic parameters on both linear and non-linear responses of the system. This analysis determines the parameters of two different transition matrices (linear and nonlinear) that describe the influence of one gene expression level on another. Using yeast cell cycle microarray data as test set, we calculated the transition matrices and used these dynamic parameters as a metric for cluster analysis. Hierarchical cluster analysis of this transition matrix reveals how a set of genes influence the expression of other genes activated during different cell cycle phases. Most strikingly, genes in different stages of cell cycle preferentially activate or inactivate genes in other stages of cell cycle, and this relationship can be readily visualized in a two-way clustering image. The observation is prior to any knowledge of the chronological characteristics of the cell cycle process. This method shows the utility of using model parameters as a metric in cluster analysis.

Cell Cycle↗

[Comparisons of brain weight and neocortical thickness in 21-day-old fetuses and 1-day-old rats with the histophysiological indices of the ovaries and adrenal glands of their mothers (a correlational and cluster analysis)].

The correlation and cluster analysis revealed that the brain weight of the rat 21-day-old foetus obviously depended on 3 beta-hydroxysteroiddehydrogenase activity in adrenal cortex. No such dependence was found in 1-day-old foetus. Thickness of the brain cortex in 21-day-old foetus and in 1-day-old rat correlated with parameters of the enzyme activity in the corpus luteum, corpus atretic, and theca folliculi of mother's ovaries.

3-Hydroxysteroid Dehydrogenases↗

TreeSOM: Cluster analysis in the self-organizing map.

Clustering problems arise in various domains of science and engineering. A large number of methods have been developed to date. The Kohonen self-organizing map (SOM) is a popular tool that maps a high-dimensional space onto a small number of dimensions by placing similar elements close together, forming clusters. Cluster analysis is often left to the user. In this paper we present the method TreeSOM and a set of tools to perform unsupervised SOM cluster analysis, determine cluster confidence and visualize the result as a tree facilitating comparison with existing hierarchical classifiers. We also introduce a distance measure for cluster trees that allows one to select a SOM with the most confident clusters.

Age Factors↗

Identifying diverse HIV risk groups among American Indian young adults: the utility of cluster analysis.

We demonstrate the utility of cluster analysis for identifying diverse HIV risk groups found in a community-based sample. Within a group of 706 American Indian young adults, we used cluster analysis to identify four profiles of HIV risk/protection. The High Efficacy/Low Risk cluster had high levels of knowledge/education, self-efficacy, and outcome expectations about HIV protection, with low levels of risk behaviors. Low Efficacy/Low Risk had low levels of HIV knowledge/education, self-efficacy, and outcome expectations, but high levels of perceived risk for HIV with low levels of HIV risk behaviors. Low Efficacy/Moderate Risk was similar to the previous group, but its members had moderately higher levels of several risk behaviors and higher condom use. Low Efficacy/High Risk had high rates of several high-risk behaviors such as exchanging sex for money or injection drug use. Validation analyses highlighted differences that can be useful for the development of preventive interventions.

Adult↗

Fuzzy cluster analysis of molecular dynamics trajectories.

We propose fuzzy clustering as a method to analyze molecular dynamics (MD) trajectories, especially of proteins and polypeptides. A fuzzy cluster analysis locates classes of similar three-dimensional conformations explored during a molecular dynamics simulation. The method can be readily applied to results from both equilibrium and nonequilibrium simulations, with clustering on either global or local structural parameters. The potential of this technique is illustrated by results from fuzzy cluster analyses of trajectories from MD simulations of various fragments of human parathyroid hormone (PTH). For large molecules, it is more efficient to analyze the clustering of root-mean-square distances between conformations comprising the trajectory. We found that the results of the clustering analysis were unambiguous, in terms of the optimal number of clusters of conformations, for the majority of the trajectories examined. The conformation closest to the cluster center can be chosen as being representative of the class of structures making up the cluster, and can be further analyzed, for example, in terms of its secondary structure. The CPU time used by the cluster analysis was negligible compared to the MD simulation time.

Amino Acid Sequence↗

Finding needles in haystacks: Reranking DOT results by using shape complementarity, cluster analysis, and biological information.

We present an evaluation of our results for the first Critical Assessment of PRedicted Interaction (CAPRI). The methods used include the molecular docking program DOT, shape analysis tool FADE, cluster analysis and filtering based on biological data. Good results were obtained for most of the seven CAPRI targets, and for two systems, submissions having the highest number of correctly predicted contacts were produced.

Algorithms↗

Cluster analysis for automatic image segmentation in dynamic scintigraphies.

An original and entirely automatic algorithm is proposed to select regions of interest (ROIs) on dynamic scintigrams. This algorithm is based on factor analysis and on cluster analysis. It consists of first extracting the orthogonal factor images of the series using factor analysis of correspondence. These factor images are then automatically segmented in ROIs using a hierarchical ascendant classification procedure. The distance used for the classification is the 'minimum added intra-class variance' distance. This algorithm has been implemented on a fast computer dedicated to nuclear medicine (Nodecrest Micas V system). The time of calculation on 1000 pixels from 40 images is less than 5 min when three factor images are used. This algorithm is validated using a numerical phantom and is illustrated using renal (99Tcm DTPA) and cardiac (equilibrium gated angiography) dynamic scintigraphies. The results show that the algorithm is able to recognize the bladder, the renal cavities and the renal parenchyma on the renal series, and the ventricules and the atria on the cardiac series.

Algorithms↗

Structure and function of gustatory neurons in the nucleus of the solitary tract. III. Classification of terminals using cluster analysis.

In sensory systems, insight into synaptic arrangements on cells of known physiological response properties has helped our understanding of the structural basis for these properties. To carry out these types of studies, however, synaptic types in the region of interest must be defined. Unfortunately, defining synaptic types in the brainstem has proved to be a challenging enterprise. Our study was done to classify synapses in the gustatory part of the nucleus solitarius using objective quantitative criteria and a cluster analysis procedure. Cluster analysis allows classification of a population of objects, such as synaptic terminals, into groups that exhibit similar characteristics. Six terminal types were identified using cluster analysis and subsequent analyses of variance and post hoc tests. Unlike classification schemes used for the cerebral cortex, where synaptic apposition density thickness and shape of vesicles is useful (Gray's Type I and II synapses), the concentration of vesicles in a terminal was a more useful measurement with which to classify terminals in the nucleus solitarius. To validate that vesicle density (vesicles/microm2) is a useful defining characteristic to classify terminals in the nucleus solitarius, terminals of a known type were used. GABAergic terminals were identified using postembedding immunohistochemical techniques, and their vesicle density was determined. GABAergic terminals fall into the range of two of the terminal types defined by the cluster analysis and, based on vesicle density, two types of GABAergic terminals were identified. We conclude that vesicle density is a helpful means to identify synapses in this brainstem nucleus.

Analysis of Variance↗

Using cluster analysis to examine dietary patterns: nutrient intakes, gender, and weight status differ across food pattern clusters.

OBJECTIVE: This study explored the usefulness of cluster analysis in identifying food choice patterns of three groups of adults in relation to their energy intake. DESIGN: Food frequency data were converted to percentage of total energy from 38 food groups and entered into a cluster analysis procedure. Subjects in the emerging food group patterns were compared in terms of weight status, demographics, and the nutrition composition of their usual diet. SETTING: Data were collected as part of three studies in two US metropolitan areas using identical protocols. Participants were university employees (103 women and 99 men) who volunteered for a reliability study of health behavior questionnaires and moderately obese volunteers (223 women and 101 men) to two weight-loss studies who were recruited by newspaper advertisements. STATISTICAL ANALYSIS PERFORMED: Subjects were clustered according to food energy sources using the FASTCLUS procedure in the Statistical Analysis System. One-way analysis of variance and chi 2 analysis were then performed to compared the weight status, nutrient intakes, and demographics of the food patterns. RESULTS: Six food pattern clusters were identified. Subjects in the two clusters associated with high consumption of pastry and meat had significantly higher fat intakes (P = .0001). Subjects in two other clusters, those associated with high intake of skim milk and a broad distribution of energy sources had significantly higher micronutrient levels (P = .0001). Body mass index and the distribution of gender were also significantly different across clusters. CONCLUSIONS: The success of cluster analysis in identifying dietary exposure categories with unique demographic and nutritional correlates suggests that the approach may be useful in epidemiologic studies that examine conditions such as obesity, and in the design of nutrition interventions.

Adult↗

Identification of andrologic patient groups by cluster analysis.

To test the validity of the current andrologic classification system, we performed cluster analysis in 317 andrologic patients involuntarily barren for more than 1 year. For cluster analysis, the spermatologic parameters sperm density, motility, and morphologic features were used since only these parameters contributed significantly to group identification, as revealed by stepwise discriminant function analysis. The optimal number clusters determined by calculation of the variance criterion was five. The resulting five groups partly correspond to the prevailing descriptive classification system as far as the extreme groups of "high-grade oligoteratoasthenozoospermia" and "polyzoospermia" are concerned. Surprisingly, cluster analysis distinguished between two groups of normozoospermia that differed in their mean sperm density. Cluster analysis may prove to become a powerful tool for andrologic classification.

Humans↗

Differential gene expression in premalignant human epidermis revealed by cluster analysis of serial analysis of gene expression (SAGE) libraries.

Serial analysis of gene expression (SAGE) has been used for quantitative analysis of gene expression. We applied cluster analysis on multiple SAGE libraries derived from premalignant epidermal tissue (actinic keratosis), normal human epidermis, and cultured keratinocytes. The samples were obtained from skin biopsies without contamination by dermal tissue or blood. A total of 60,000 transcripts (tags) were analyzed. Two-way cluster analysis was applied to both the transcripts and the tissues, resulting in separation of the cultured cells from the epidermal samples, and clustering of many, presumably coregulated, genes. Two clusters of genes, strongly up-regulated in the tumor tissue compared with normal epidermis, were investigated in more detail. The differential expression of genes could be confirmed in actinic keratosis from four patients. Several of these genes have been previously associated with carcinogenesis or are likely to be important on the basis of their presumed function. Automated literature search tools show that a subgroup of these genes is coexpressed in other tissues and is part of an epidermal differentiation gene cluster on chromosome 1q21. We conclude that cluster analysis on large data sets uncovers clear partitions and correlations that could be confirmed by independent methods. We predict that these partitions will lead to biological interpretations that can be relevant for understanding the processes of carcinogenesis and tumor progression.

Blotting, Northern↗

A new method for classifying patterns of prenatal care utilization using cluster analysis.

OBJECTIVES: The objectives of this study were: to 1) define patterns of prenatal care utilization using cluster analysis, 2) describe two alternative cluster solutions and compare these groupings to the Adequacy of Prenatal Care Utilization Index (APNCU), 3) compare the cluster solutions and the APNCU with respect to maternal age and prematurity, and 4) discuss advantages and disadvantages of using cluster analysis to study prenatal care. METHODS: The study sample included 3544 women in the 1988 National Maternal and Infant Health Survey for whom complete prenatal care visit data were available. Clustering was carried out in two stages, first employing nearest centroid sorting (the k means method), a nonhierarchical approach, and then using Ward's Minimum Variance Method, a hierarchical clustering technique. RESULTS: Patterns of prenatal care defined by cluster analysis varied by timing of the first visit, total number of visits, and the rate of accumulation of visits, but this variation was different compared to that seen for the APNCU. While the cluster solutions and the APNCU identified a similar normative pattern of care, other patterns identified were quite different. In particular, the six-cluster solution differentiated among women who entered care at similar times, but accumulated visits at differing rates and experienced differing rates of preterm delivery. CONCLUSION: Cluster analysis is a new tool for studying prenatal care. Further studies are needed to refine the method and test whether the alternative perspective it provides will lead to new findings concerning the relationship of prenatal care and birth outcomes.

Birth Certificates↗

Cluster analysis methods help to clarify the activity-BMI relationship of Chinese youth.

OBJECTIVE: To use cluster analysis to create patterns of overall activity and inactivity in a diverse sample of Chinese youth and to evaluate their use in predicting overweight status. RESEARCH METHODS AND PROCEDURES: The study populations were drawn from the 1997 and 2000 years of the longitudinal China Health and Nutrition Survey, comprised of 2702 and 2641 schoolchildren in the 1997 and 2000 cross-sectional samples, respectively, and 1175 children in the longitudinal cohort. Cluster analysis was used to group children into nonoverlapping activity/inactivity "clusters" that were subsequently used in models of prevalent and incident overweight. Results were compared with traditional models, with activity and inactivity coded separately, to assess whether further insight was gained with the cluster analysis methodology. RESULTS: Moderately and highly active youth were shown to have significantly decreased odds of overweight in both cross-sectional and longitudinal analyses using cluster analysis. In incident longitudinal models, youth in the high activity/high inactivity cluster had the lowest odds of overweight [odds ratio=0.12 (0.03, 0.44)]; in contrast, results from traditional models failed to show any significant relationship between overweight and activity or inactivity. DISCUSSION: Cluster analysis methods allow researchers to simultaneously capture activity and inactivity in new ways. In this comparative study, only with the clustering methodology did we find a significant effect of activity on incident overweight, furthering our ability to examine this complex relationship. Interestingly, no effect of increasing levels of inactivity was observed using either method, indicating that activity seems to be the more important determinant of overweight in this population.

Adolescent↗

Cluster analysis of flow cytometric list mode data on a personal computer.

A cluster analysis algorithm, dedicated to analysis of flow cytometric data is described. The algorithm is written in Pascal and implemented on an MS-DOS personal computer. It uses k-means, initialized with a large number of seed points, followed by a modified nearest neighbor technique to reduce the large number of subclusters. Thus we combine the advantage of the k-means (speed) with that of the nearest neighbor technique (accuracy). In order to achieve a rapid analysis, no complex data transformations such as principal components analysis were used. Results of the cluster analysis on both real and artificial flow cytometric data are presented and discussed. The results show that it is possible to get very good cluster analysis partitions, which compare favorably with manually gated analysis in both time and in reliability, using a personal computer.

Algorithms↗

Differential aging: an exploratory approach using cluster analysis.

This article uses cluster analysis to identify different patterns of personal resources within a random sample of the well, elderly population. Ten such patterns or natural groupings are identified and their implications for coping and successful aging are discussed. It is apparent that there are a number of ways both of aging well and aging badly, and that these patterns cannot be predicted solely on the basis of structural data. The article poses a number of questions on the performance of cluster members over time and draws attention to the importance of longitudinal data.

Adaptation, Psychological↗

[Configuration cluster analysis as an alternative to configuration frequency analysis].

The paper discusses configural cluster analysis (CCA) as an alternative to configural frequency analysis (CFA). CFA defines types as deviations from the assumption of no interactions among the variables. CCA defines clusters as deviations from the stronger assumption of a total lack of effects, in the contingency table. Variations of CCA as aggregating CCA, hierarchical CCA, m-sample CCA, and CCA of profile shifts are discussed. An example analyzes data from research on depression.

Cluster Analysis↗