PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cluster analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Cluster analysis of mass spectrometry data reveals a novel component of SAGA.

The SAGA histone acetyltransferase and TFIID complexes play key roles in eukaryotic transcription. Using hierarchical cluster analysis of mass spectrometry data to identify proteins that copurify with components of the budding yeast TFIID transcription complex, we discovered that an uncharacterized protein corresponding to the YPL047W open reading frame significantly associated with shared components of the TFIID and SAGA complexes. Using mass spectrometry and biochemical assays, we show that YPL047W (SGF11, 11-kDa SAGA-associated factor) is an integral subunit of SAGA. However, SGF11 does not appear to play a role in SAGA-mediated histone acetylation. DNA microarray analysis showed that SGF11 mediates transcription of a subset of SAGA-dependent genes, as well as SAGA-independent genes. SAGA purified from a sgf11 Delta deletion strain has reduced amounts of Ubp8p, and a ubp8 Delta deletion strain shows changes in transcription similar to those seen with the sgf11 Delta deletion strain. Together, these data show that Sgf11p is a novel component of the yeast SAGA complex and that SGF11 regulates transcription of a subset of SAGA-regulated genes. Our data suggest that the role of SGF11 in transcription is independent of SAGA's histone acetyltransferase activity but may involve Ubp8p recruitment to or stabilization in SAGA.

Acetyltransferases↗

An investigation of protistan phylogeny using a numerical taxonomy (cluster) analysis of mitotic systems.

Fourteen mitotic characteristics of 83 species, or groups of species, representing 134 species spread among the "fungi", "algae", "protozoa" and "related" organisms are analyzed by computer-aided cluster analysis techniques. The resultant clusters are represented as phenograms that indicate the degree of similarity among the mitotic systems considered. When compared with other criteria, some clusters correlate well with established taxa but others contain a heterogeneous array of species. Some established taxa are widely dispersed. Some of these "deviant" clusters are probably artefacts of the methodology but others are considered to indicate novel but true common ancestry. It is concluded that this type of analysis and data have phylogenetic value but must be used with considerable care.

Animals↗

Cluster analysis and data visualization of large-scale gene expression data.

The discovery of any new gene requires an analysis of the expression context for that gene. Now that the cDNA and genomic sequencing projects are progressing at such a rapid rate, high throughput gene expression screening approaches are beginning to appear to take advantage of that data. We present a strategy for the analysis for large-scale quantitative gene expression measurement data from time course experiments. Our approach takes advantage of cluster analysis and graphical visualization methods to reveal correlated patterns of gene expression from time series data. The coherence of these patterns suggests an order that conforms to a notion of shared pathways and control processes that can be experimentally verified.

Base Sequence↗

Applying cluster analysis to test a typology of homelessness by pattern of shelter utilization: results from the analysis of administrative data.

This study tests a typology of homelessness using administrative data on public shelter use in New York City (1988-1995) and Philadelphia (1991-1995). Cluster analysis is used to produce three groups (transitionally, episodically, and chronically homeless) by number of shelter days and number of shelter episodes. Results show that the transitionally homeless, who constitute approximately 80% of shelter users in both cities, are younger, less likely to have mental health, substance abuse, or medical problems, and to overrepresent Whites relative to the other clusters. The episodically homeless, who constitute 10% of shelter users, are also comparatively young, but are more likely to be non-White, and to have mental health, substance abuse, and medical problems. The chronically homeless, who account for 10% of shelter users, tend to be older, non-White, and to have higher levels of mental health, substance abuse, and medical problems. Differences in health status between the episodically and chronically homeless are smaller, and in some cases the chronically homeless have lower rates (substance abuse in New York; serious mental illness in Philadelphia). Despite their relatively small number, the chronically homeless consume half of the total shelter days. Results suggest that program planning would benefit from application of this typology, possibly targeting the transitionally homeless with preventive and resettlement assistance, the episodically homeless with transitional housing and residential treatment, and the chronically homeless with supported housing and long-term care programs.

Adult↗

Deciphering protein sequence information through hydrophobic cluster analysis (HCA): current status and perspectives.

Ten years after the idea of hydrophobic cluster analysis (HCA) was conceived and first published, theoretical and practical experience has shown this unconventional method of protein sequence analysis to be particularly efficient and sensitive, especially with families of sequences sharing low levels of sequence identity. This extreme sensitivity has made it possible to predict the functions of genes whose sequence similarities are hardly if at all detectable by current one-dimensional (1D) methods alone, and offers a new way to explore the enormous amount of data generated by genome sequencing. HCA also provides original tools to understand fundamental features of protein stability and folding. Since the last review of HCA published in 1990 [1], significant improvements have been made and several new facets have been addressed. Here we wish to update and summarize this information.

Amino Acid Sequence↗

Cluster analysis of a forensic population with antisocial personality disorder regarding PCL-R scores: differentiation of two patterns of criminal profiles.

Fifty six cases of a forensic population were submitted to a cluster analysis to observe the aglomerative behavior in relation to the total scores of the items comprising the PCL-R Psychopathy Checklist Revised [R.D. Hare, Manual for the Hare Psychopathy Checklist-Revised, Multi-Health System, Toronto, 1991]. The analysis indicated two independent types of antisocial personality disorders, not identified in the PCL-R in its standardized form, one of them being strongly associated with criminal conduct and the other with psychopathic personality. Such clusters were stable when the analysis was replicated with other hierarchical algorithms, and also, they were independently extracted via the k-means method without having previously fixed the value for k. One of the clusters concentrated the PCL-R highest scores, indicating that it is the prototypical psychopathic character determinant.

Adolescent↗

Cluster analysis of human and animal pathogenic Microsporum species and their teleomorphic states, Arthroderma species, based on the DNA sequences of nuclear ribosomal internal transcribed spacer 1.

We performed a cluster analysis of human and animal pathogenic Microsporum species and their teleomorphic states, Arthroderma species, including A. otae-related species (M. canis, M. audouinii, M. distortum, M. equinum, M. langeronii, and M. ferrugineum) and M. gypseum complex (A. fulvum, A. gypseum, and A. incurvatum) using DNA sequences of nuclear ribosomal internal transcribed spacer 1 (ITS1). The dendrogram showed the members of A. otae-related species to be monophyletic and to construct an extremely closely related cluster with a long horizontal branch. This ITS1-homologous group of A. otae was organized in 6 unique genotypes, while sequences of the members of the ITS1-homologous group of M. gypseum complex are more diverse. This ITS1-based database of Microsporum species and their teleomorphic states will provide a useful and reliable species identification system: it is time-saving (takes two to three days), accurate and applicable even to strains with atypical morphological features or in a non-culturable state.

Animals↗

Model-based cluster analysis of microarray gene-expression data.

BACKGROUND: Microarray technologies are emerging as a promising tool for genomic studies. The challenge now is how to analyze the resulting large amounts of data. Clustering techniques have been widely applied in analyzing microarray gene-expression data. However, normal mixture model-based cluster analysis has not been widely used for such data, although it has a solid probabilistic foundation. Here, we introduce and illustrate its use in detecting differentially expressed genes. In particular, we do not cluster gene-expression patterns but a summary statistic, the t-statistic. RESULTS: The method is applied to a data set containing expression levels of 1,176 genes of rats with and without pneumococcal middle-ear infection. Three clusters were found, two of which contain more than 95% genes with almost no altered gene-expression levels, whereas the third one has 30 genes with more or less differential gene-expression levels. CONCLUSIONS: Our results indicate that model-based clustering of t-statistics (and possibly other summary statistics) can be a useful statistical tool to exploit differential gene expression for microarray data.

Animals↗

[The cluster analysis of trace elements in fructus cnidii from different region].

The contents of trace elements in Fructus Cnidii from different region were assayed by atom absorption spectrum and analyzed by cluster analysis methods. The results showed the contents of trace elements from different regions were different, which had some relativity with habitats, but should be further studied.

Cluster Analysis↗

Development of intracranial relations in patients aged 10 to 18 years with clefts of the lip and palate, using cluster analysis.

The investigation is based on a longitudinal cephalometric investigation of lateral teleroentgenographic pictures of male patients with a complete unilateral cleft of the lip and palate. Using cluster analysis the authors investigated the relationship of 75 craniofacial characteristics of size, shape and position during the time interval from 10 to 18 years of age. The main objective of the work was to characterize the development of intracranial relations during the pubertal spurt and compare the final condition in adulthood with a control group. The angle of the cranial base and its effect on the position of the mandibular joint did not change during the investigation period. The relationship between the rotation of the mandible and the inclination of the upper alveolar process with the protrusion of different parts of the skeletal profile also remained constant. Up to adulthood, the rotation of the mandible developed independently of the sagittal intermaxillary relations. The relationship between the sagittal intermaxillary relations and other parts of the face did, however, change. Before the onset of puberty it was influenced most by the reduced length of the maxilla. The inadequacy of maxillary growth was balanced during this period by a change in the shape and position of the mandible. Its adaptative capacities could not compensate later for this uneven development of the jaws potentiated by the pubertal growth spurt. Due to this the intermaxillary relations deteriorated at the end of development in the majority of patients. The association of the restricted vertical maxillary growth with its retroposition was manifested only in adulthood. Intracranial relations of the control group differed from those in the group with clefts. Sagittal intermaxillary and dental relations were not associated in healthy men. As individual probands were not linked by any restriction of growth or development, no close relationship developed between the shape characteristics of the lower jaw, which is the main compensatory adaptative mechanism.

Adolescent↗

Effectiveness of environmental cluster analysis in representing regional species diversity.

A major challenge of regional conservation planning is the identification of sets of sites that together represent the overall biodiversity of the relevant region. Environmental cluster analysis (ECA) has been proposed as a potential tool for efficient selection of conservation sites, but the consequences of methodological decisions involved in its application have not been tested so far. We evaluated the performance of ECA with respect to two such decisions: the choice of the clustering algorithm (single linkage, complete linkage, unweighted arithmetic average, unweighted centroid, Ward's minimum variance, and the ALOC algorithm) and the weight given to different groups of environmental variables (rainfall, temperature, and lithology). Specifically we tested how these decisions affect the spatial configuration of clusters of sites defined by the ECA, whether and how they affect the effectiveness of the ECA (i.e., its ability to represent regional species diversity), and whether the effectiveness of alternative methods of hierarchical clustering can be predicted a priori based on the cophenetic correlation. We used an extensive database of the flora of Israel to test these questions. Differences in both the clustering algorithm and the weighting regime had considerable effects on the spatial configuration of the ECA clusters. The single-linkage algorithm produced mostly single-cell clusters plus a single large-sized cluster and was therefore found inappropriate for environmental regionalization. The effectiveness of the ECA was also sensitive to changes in the clustering algorithm and the weighting regime. Yet, most combinations of clustering algorithms and weighting regimes performed significantly better in capturing regional biodiversity than random null models. The main deviation was classifications based on Ward's minimum variance algorithm, which performed less well relative to all other algorithms. The two algorithms that showed the highest effectiveness (unweighted average and unweighted centroid clustering) also exhibited the highest values of the cophenetic correlation, suggesting that this index may serve as a potential indicator for the effectiveness of alternative ECA algorithms.

Algorithms↗

Searching for potential drug targets in two-component and phosphorelay signal-transduction systems using three-dimensional cluster analysis.

Two-component and phosphorelay signal transduction systems are central components in the virulence and antimicrobial resistance responses of a number of bacterial and fungal pathogens; in some cases, these systems are essential for bacterial growth and viability. Herein, we analyze in detail the conserved surface residue clusters in the phosphotransferase domain of histidine kinases and the regulatory domain of response regulators by using complex structure-based three-dimensional cluster analysis. We also investigate the protein-protein interactions that these residue clusters participate in. The Spo0B-Spo0F complex structure was used as the reference structure, and the multiple aligned sequences of phosphotransferases and response regulators were paired correspondingly. The results show that a contiguous conserved residue cluster is formed around the active site, which crosses the interface of histidine kinases and response regulators. The conserved residue clusters of phosphotransferase and the regulatory domains are directly involved in the functional implementation of two-component signal transduction systems and are good targets for the development of novel antimicrobial agents.

Binding Sites↗

Expression of p34(cdc2) and cyclins A and B compared to other proliferative features of non-Hodgkin's lymphomas: a multivariate cluster analysis.

In view of recent knowledge on proteins regulating the cell cycle, we re-evaluated proliferative features of 98 diffusely growing non-Hodgkin's lymphomas. The combined use of 5 proliferation-associated variables (mitotic indices and percentages of Ki-67(+), p34(cdc2+), cyclin A(+) and cyclin B(+) cells) and their entry into a multivariate cluster analysis separated, without overlaps, the entire cohort into 3 groups (clusters) with (1) low, (2) intermediate and (3) high proliferative activity. Conversely, bivariate plots exposed considerable cluster overlaps. Multivariate stepwise discriminant analysis of all cases revealed a decreasing order of discriminant power for % Ki-67(+) cells > % p34(cdc2+) cells > mitotic index > % cyclin A(+) cells > % cyclin B(+) cells. The combined use of 2 variables only, mitotic index and % p34(cdc2+) cells, allowed a clear-cut separation of clusters 2 and 3. In bivariate plots, correlations were best between % Ki-67(+) cells and % cyclin A(+) cells and between mitotic indices and % cyclin B(+) cells. Except for chronic lymphocytic leukemias, immunocytomas and marginal zone lymphomas (all in cluster 1), individual lymphoma entities were distributed among at least 2 clusters. There was, however, a marked preponderance of mantle cell lymphomas and diffuse follicular center lymphomas in cluster 1 and of diffuse large B-cell lymphomas and peripheral T-cell lymphomas in cluster 2. Anaplastic large-cell lymphomas predominated in cluster 3 and responded best to therapy.

Adolescent↗

Cluster analysis and disease mapping--why, when, and how? A step by step guide.

Growing public awareness of environmental hazards has led to an increased demand for public health authorities to investigate geographical clustering of diseases. Although such cluster analysis is nearly always ineffective in identifying causes of disease, it often has to be used to address public concern about environmental hazards. Interpreting the resulting data is not straightforward, however, and this paper presents a guide for the non-specialist. The pitfalls include the fact that cluster analyses are usually done post hoc, and not as a result of a prior hypothesis. This is particularly true for investigations prompted by reported clusters, which have the inherent danger of overestimating the disease rate through "boundary shrinkage" of the population from which the cases are assumed to have arisen. In disease surveillance the problem of making multiple comparisons can be overcome by testing for clustering and autocorrelation. When rates of disease are illustrated in disease maps undue focus on areas where random fluctuation is greatest can be minimised by smoothing techniques. Despite the fact that cluster analyses rarely prove fruitful in identifying causation, they may-like single case reports-have the potential to generate new knowledge.

Cluster Analysis↗

Social phobia subtypes in the general population revealed by cluster analysis.

BACKGROUND: Epidemiological data on subtypes of social phobia are scarce and their defining features are debated. Hence, the present study explored the prevalence and descriptive characteristics of empirically derived social phobia subgroups in the general population. METHODS: To reveal subtypes, data on social distress, functional impairment, number of social fears and criteria fulfilled for avoidant personality disorder were extracted from a previously published epidemiological study of 188 social phobics and entered into an hierarchical cluster analysis. Criterion validity was evaluated by comparing clusters on the Social Phobia Scale (SPS) and the Social Interaction Anxiety Scale (SIAS). Finally, profile analyses were performed in which clusters were compared on a set of sociodemographic and descriptive characteristics. RESULTS: Three clusters emerged, consisting of phobics scoring either high (generalized subtype), intermediate (non-generalized subtype) or low (discrete subtype) on all variables. Point prevalence rates were 2.0%, 5.9% and 7.7% respectively. All subtypes were distinguished on both SPS and SIAS. Generalized or severe social phobia tended to be over-represented among individuals with low levels of educational attainment and social support. Overall, public-speaking was the most common fear. CONCLUSIONS: Although categorical distinctions may be used, the present data suggest that social phobia subtypes in the general population mainly differ dimensionally along a mild moderate-severe continuum, and that the number of cases declines with increasing severity.

Adult↗

Correlation and cluster analysis of sensory, pain, and reflex thresholds to various stimulus modalities in symptom-free subjects.

OBJECTIVE: In order to evaluate the possible relation between the psychophysical response and a motor reflex, sensory and pain thresholds to various stimuli were analyzed in combination with the occurrence threshold of the late masseteric exteroceptive suppression (ES2) period. METHODS: Twenty men and 20 women participated. The tactile detection threshold and the filament-prick pain detection threshold were measured on the cheek skin overlying the left masseter muscles. The pressure pain threshold and pressure pain tolerance threshold were measured at the left masseter muscle. The surface EMG was recorded from the left masseter muscle, while electrical stimuli with 13 fixed intensities were applied to the skin above the left mental nerve. The stimulation intensity at which the ES2 appeared for the first time and the lowest stimulus intensity at which the subjects reported to be painful were defined as the ES2 and pain threshold, respectively. RESULTS: There were significant positive correlations between the tactile detection threshold and the pain thresholds determined using the different stimulus modalities, and the ES2 threshold was also significantly correlated with the pain thresholds (P<0.05). Cluster analysis could significantly discriminate two distinct groups with high versus low tactile, pain and ES2 thresholds (P<0.05). CONCLUSIONS: The present findings suggested that the ES2 reflex response has a relation with the individual sensory and pain sensitivity in symptom-free subjects. SIGNIFICANCE: Combined examination of sensory, pain, and ES2 thresholds might provide complementary information on the pathophysiology underlying orofacial pain.

Adult↗

Use of cluster analysis for gait pattern classification of patients in the early and late recovery phases following stroke.

The mixture of gait deviations seen in patients following a stroke is remarkably variable. An objective system for classification of gait patterns for this population could be used to guide treatment planning. Quantitated gait analysis was conducted for 47 individuals at admission to in-patient rehabilitation and again at 6 months post-stroke for 42 subjects. Non-hierarchical cluster analysis was used to classify the gait patterns of patients based on the temporal-spatial and kinematic parameters of walking. Four clusters of patients were identified at both assessment intervals. At the admission test walking velocity, peak knee extension in mid stance and peak dorsiflexion in swing were the three factors that best characterized the groups. At 6 months the explanatory variables were velocity, knee extension in terminal stance, and knee flexion in pre swing. Differences in muscle strength and muscle activation patterns during walking were identified between groups.

Adult↗

Energy intake profile in the triathlon competition by means of cluster analysis.

The subjects were 18 male triathletes competing in the 4th Kaike Triathlon held at Tottori in 1984. Answers to a questionnaire on dietary and water intake before and during the competition were analysed. Cluster analysis using the average distance method was applied to 21 variables relating to order, age, stature, body weight, total calorie intake and each nutrient element ingested during three periods: after swimming, during cycling and during the marathon. The arrival order was clustered with stature and weight, and its similarity showed 69.7 when the first cluster formed. Age along with carbohydrate, protein or water intake during the marathon formed secondary clusters with order, stature, and weight (similarity: 60.0). In the upper place group, breakfast calories occupied the largest portion of 26.1% and the smallest portion of 17.2% during the marathon running. Furthermore the group ingested more energy than the other groups during cycling.

Adult↗