PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multiple clustering”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A multiple sclerosis cluster associated with a small, north-central Illinois community.

The authors investigated a reported incidence cluster of multiple sclerosis (MS) cases in a small, north-central Illinois community to determine validity and statistical significance. DePue, Illinois--a small, north-central Illinois community--has previously been the site of significant environmental heavy-metal exposure from a zinc smelter. Significant contamination of soil and water with zinc and other metals has been documented in this community during the time period of interest. In the mid-1990s, several cases of MS were reported to the Illinois Department of Public Health within the geographic limits of this community. Available medical records from purported MS cases reported to the Illinois Department of Public Health were reviewed, and living individuals were seen and examined. Statistical analyses were conducted with clinically definite MS cases; onset dates were determined by first symptom, and expected incidence rates were determined from published epidemiologic studies. Nine new cases of clinically definite MS occurred among residents of DePue, Illinois, during the period between 1971 and 1990. Seven of the 8 living subjects included in the final analyses were examined by one author (RS). The computed incidence rate deriving from these cases within DePue Township, Illinois, represented a statistically significant excess of new MS cases over expected. During the period from 1971 through 1990, a significant excess of MS cases occurred within the population of DePue, Illinois. Significant exposure of this population to mitogenic trace metals, including zinc, was also documented during this time period.

Adolescent↗

Universal scaling functions for bond percolation on planar-random and square lattices with multiple percolating clusters.

Percolation models with multiple percolating clusters have attracted much attention in recent years. Here we use Monte Carlo simulations to study bond percolation on L1xL2 planar random lattices, duals of random lattices, and square lattices with free and periodic boundary conditions, in vertical and horizontal directions, respectively, and with various aspect ratios L(1)/L(2). We calculate the probability for the appearance of n percolating clusters, W(n); the percolating probabilities P; the average fraction of lattice bonds (sites) in the percolating clusters, (n) ( (n)), and the probability distribution function for the fraction c of lattice bonds (sites), in percolating clusters of subgraphs with n percolating clusters, f(n)(c(b)) [f(n)(c(s))]. Using a small number of nonuniversal metric factors, we find that W(n), P, (n) ( (n)), and f(n)(c(b)) [f(n)(c(s))] for random lattices, duals of random lattices, and square lattices have the same universal finite-size scaling functions. We also find that nonuniversal metric factors are independent of boundary conditions and aspect ratios.

Journal Article↗

Multiple myeloma: clusters, clues, and dioxins.

Multiple myeloma (MM) is a B-cell neoplasm of unknown etiology. We searched for etiological clues by examining the literature on geographic clusters of MM. We searched the MEDLINE database from 1966 to 1996 for spatial occurrences of MM that were significantly greater than expected (spatial "clusters"). Eight clusters with verified diagnoses of MM were identified. All of the eight clusters occurred in locations that were proximate to a body of water. Six of these bodies of water are known to have been contaminated with dioxins. We hypothesize that the observed association between MM and proximity to bodies of water is caused by exposure to dioxins in individuals who consume local fish and seafood. This hypothesis is consistent with the significantly elevated risks for MM in groups with high consumption of dioxin-contaminated fish, e.g., Baltic Sea fishermen and Alaskan Indians, and among persons accidentally exposed to dioxins in Seveso, Italy. Dioxins are immunotoxic and inhibit the differentiation of B cells. Thus, dioxins are plausible myelomagens. A dioxin hypothesis could illuminate many epidemiological features of MM and may suggest new avenues for analytic research.

Cluster Analysis↗

Endemic clustering of multiple sclerosis in time and place, 1934-1984. Confirmation of a hypothesis.

The occurrence of multiple sclerosis in clustered groups of cases that are often related to others in time and place has been observed on several occasions in the last 50 years. Selected clusters are here reviewed in relation to suspected sources of heavy metal (mercury, lead) poisoning as background for the analysis of the 1983-1985 "outbreak" of 30-40 cases of multiple sclerosis in Key West, Florida. Evidence is presented that the time-place clustering resulted from environmental pollution stemming from a nearby dump pile of rocky debris. The probable mechanism is discussed.

Adult↗

Clustering of multiple specific genes and gene-rich R-bands around SC-35 domains: evidence for local euchromatic neighborhoods.

Typically, eukaryotic nuclei contain 10-30 prominent domains (referred to here as SC-35 domains) that are concentrated in mRNA metabolic factors. Here, we show that multiple specific genes cluster around a common SC-35 domain, which contains multiple mRNAs. Nonsyntenic genes are capable of associating with a common domain, but domain "choice" appears random, even for two coordinately expressed genes. Active genes widely separated on different chromosome arms associate with the same domain frequently, assorting randomly into the 3-4 subregions of the chromosome periphery that contact a domain. Most importantly, visualization of six individual chromosome bands showed that large genomic segments ( approximately 5 Mb) have striking differences in organization relative to domains. Certain bands showed extensive contact, often aligning with or encircling an SC-35 domain, whereas others did not. All three gene-rich reverse bands showed this more than the gene-poor Giemsa dark bands, and morphometric analyses demonstrated statistically significant differences. Similarly, late-replicating DNA generally avoids SC-35 domains. These findings suggest a functional rationale for gene clustering in chromosomal bands, which relates to nuclear clustering of genes with SC-35 domains. Rather than random reservoirs of splicing factors, or factors accumulated on an individual highly active gene, we propose a model of SC-35 domains as functional centers for a multitude of clustered genes, forming local euchromatic "neighborhoods."

Cell Line↗

SPATCLUS: an R package for arbitrarily shaped multiple spatial cluster detection for case event data.

This paper describes an R package, named SPATCLUS that implements a method recently proposed for spatial cluster detection of case event data. This method is based on a data transformation. This transformation is achieved by the definition of a trajectory, which allows to attribute to each point a selection order and the distance to its nearest neighbour. The nearest point is searched among the points which have not yet been selected in the trajectory. Due to the trajectory effects, the distance is weighted by the expected distance under the uniform distribution hypothesis. Potential clusters are located by using multiple structural change models and a dynamic programming algorithm. The double maximum test allows to select the best model. The significativity of potential clusters is determined by Monte Carlo simulations. This method makes it possible the detection of multiple clusters of any shape.

Algorithms↗

Clustering of multiple transgene integrations in highly-unstable Ascobolus immersus transformants.

A large proportion of Ascobolus immersus transformants are highly unstable in crosses: the phenotype conferred by the transgene is not transmitted to the progeny, irrespective of the endogenous or foreign origin of the transgene. They all have integrated multiple transgene copies, clustered at a single chromosomal site or at tightly-linked sites. Clustered non-homologous integrations are always rearranged. Yet they never escape the "methylation induced premeiotically" (MIP) process. This always results in gene silencing, even when the transgene is partially repeated, accounting for the high instability of these transformants.

Ascomycota↗

A generalized clustering problem, with application to DNA microarrays.

We think of cluster analysis as class discovery. That is, we assume that there is an unknown mapping called clustering structure that assigns a class label to each observation, and the goal of cluster analysis is to estimate this clustering structure, that is, to estimate the number of clusters and cluster assignments. In traditional cluster analysis, it is assumed that such unknown mapping is unique. However, since the observations may cluster in more than one way depending on the variables used, it is natural to permit the existence of more than one clustering structure. This generalized clustering problem of estimating multiple clustering structures is the focus of this paper. We propose an algorithm for finding multiple clustering structures of observations which involves clustering both variables and observations. The number of clustering structures is determined by the number of variable clusters. The dissimilarity measure for clustering variables is based on nearest-neighbor graphs. The observations are clustered using weighted distances with weights determined by the clusters of the variables. The motivating application is to gene expression data.

Algorithms↗

Multiple temporal cluster detection.

This article proposes a simple method to determine single or multiple temporal clustering on a variable size population. By a transformation of the data set, the method based on a regression model allows consideration of a variable population size during the time of study. A model selection procedure and a resampling method are used to select the number of clusters. The results have applications in epidemiological studies of rare diseases.

Cluster Analysis↗

Robust multi-scale clustering of large DNA microarray datasets with the consensus algorithm.

MOTIVATION: Hierarchical and relocation clustering (e.g. K-means and self-organizing maps) have been successful tools in the display and analysis of whole genome DNA microarray expression data. However, the results of hierarchical clustering are sensitive to outliers, and most relocation methods give results which are dependent on the initialization of the algorithm. Therefore, it is difficult to assess the significance of the results. We have developed a consensus clustering algorithm, where the final result is averaged over multiple clustering runs, giving a robust and reproducible clustering, capable of capturing small signal variations. The algorithm preserves valuable properties of hierarchical clustering, which is useful for visualization and interpretation of the results. RESULTS: We show for the first time that one can take advantage of multiple clustering runs in DNA microarray analysis by collecting re-occurring clustering patterns in a co-occurrence matrix. The results show that consensus clustering obtained from clustering multiple times with Variational Bayes Mixtures of Gaussians or K-means significantly reduces the classification error rate for a simulated dataset. The method is flexible and it is possible to find consensus clusters from different clustering algorithms. Thus, the algorithm can be used as a framework to test in a quantitative manner the homogeneity of different clustering algorithms. We compare the method with a number of state-of-the-art clustering methods. It is shown that the method is robust and gives low classification error rates for a realistic, simulated dataset. The algorithm is also demonstrated for real datasets. It is shown that more biological meaningful transcriptional patterns can be found without conservative statistical or fold-change exclusion of data. AVAILABILITY: Matlab source code for the clustering algorithm ClusterLustre, and the simulated dataset for testing are available upon request from T.G. and O.W.

Algorithms↗

Abdominal adiposity and clustering of multiple metabolic syndrome in White, Black and Hispanic americans.

PURPOSE: The aim of this study was to evaluate the association of abdominal adiposity assessed by waist circumference (WC) with clustering of multiple metabolic syndromes (MMS) in White, Black and Hispanic Americans. MMS was defined as the occurrence of two or more of either hypertension, type 2 diabetes mellitus, dyslipidemia, hypertriglyceridemia or hyperinsulinemia. METHODS: The number of MMS and fasting insulin (a surrogate measure of MMS) were each used as dependent variables in gender-specific multiple linear regression models, adjusting for age, smoking and alcohol intake. The contribution of WC to interethnic differences in clustering of MMS and fasting insulin concentration was assessed in gender-specific linear regression models. The risk of MMS due to large waist was estimated by comparing odds ratio for men with WC >/= 102 cm with those with WC < 102, and women with WC >/= 88 cm with women with WC < 88 cm in the logistic regression model adjusting for age, smoking and alcohol intake. RESULTS: WC was positively and independently associated with clustering of MMS and increased fasting insulin concentration adjusting for age, smoking and alcohol intake in the three ethnic groups (p < 0.01). Black ethnicity was associated with clustering of MMS and fasting insulin concentration (p < 0.01). Hispanic ethnicity was also associated with clustering of MMS in men and associated with fasting insulin concentration in both men and women (p < 0.01). In both men and women, the risk of MMS clustering was strongly associated with increased WC in all ethnic groups independent of BMI. CONCLUSION: WC appears to be a marker for multiple metabolic syndromes in these ethnic groups. The results of this investigation lend support to the view that waist measurement should be considered as a clinical variable for assessing the risk of cardiovascular diseases.

Abdomen↗

Extraction of correlated gene clusters by multiple graph comparison.

This paper presents a new method to extract a set of correlated genes with respect to multiple biological features. Relationships among genes on a specific feature are encoded as a graph structure whose nodes correspond to genes. For example, the genome is a graph representing positional correlations of genes on the chromosome, the pathway is a graph representing functional correlations of gene products, and the expression profile is a graph representing gene expression similarities. When a set of genes are localized in a single graph, such as a gene cluster on the chromosome, an enzyme cluster in the metabolic pathway, or a set of coexpressed genes in the microarray gene expression profile, this may suggest a functional link among those genes. The functional link would become stronger when the clusters are correlated; namely, when a set of corresponding genes form clusters in multiple graphs. The newly introduced heuristic algorithm extracts such correlated gene clusters as isomorphic subgraphs in multiple graphs by using inter-graph links that are defined based on biological relevance. Using the method, we found E.coli correlated gene clusters in which genes are related with respect to the positions in the genome and the metabolic pathway, as well as the 3D structural similarity. We also analyzed protein-protein interaction data by two-hybrid experiments and gene coexpression data by microarrays in S.cerevisiae, and estimated the possibility of utilizing our method for screening the datasets that are likely to contain many false positive relations.

Algorithms↗

Persistence of multiple cardiovascular risk clustering related to syndrome X from childhood to young adulthood. The Bogalusa Heart Study.

BACKGROUND: Cardiovascular risk factors are known to persist over time and to cluster both in childhood and adulthood. Less is known about the persistence of clustering of multiple cardiovascular risk factors comprising adverse levels of systolic blood pressure, the ratio of total cholesterol to high-density lipoprotein cholesterol, and plasma insulin from childhood to young adulthood. METHODS: In a community study of cardiovascular risk, 1176 individuals (52% female, 44% black) aged 5 through 17 years at baseline were followed up for 8 years. RESULTS: Calculated as sum of the race-, sex-, and age-specific rankings of systolic blood pressure, insulin level, and total to high-density lipoprotein cholesterol ratio, the multiple risk index was shown to track in all four race-sex groups (year 1 vs year 8 correlations, .54 to .67). The magnitude of the overall multiple risk index tracking correlation (r = .64) was significantly stronger than that noted for individual risk factors (r = .34 to .57). Among subjects who were initially in the highest quintile of the multiple risk index, 61% remained there 8 years later. Tracking of the multiple risk index increased progressively with age and ponderal index (weight/[height3]). In a step-wise regression analysis, baseline multiple risk index score, baseline ponderal index, change in ponderal index, and change in height were predictive of the multiple risk index score on follow-up. These predictors explained 45% to 60% of the variability in multiple risk index scores among the race-sex groups. CONCLUSIONS: The persistence of multiple cardiovascular risk clustering from childhood to adulthood and the impact of obesity in this regard point to the need for preventive measures aimed at developing healthy lifestyles early in life.

Adolescent↗

A single subtype of Epstein-Barr virus in members of multiple sclerosis clusters.

OBJECTIVES: Epidemiological studies strongly indicate an infectious involvement in multiple sclerosis (MS). Epstein-Barr virus (EBV), to which all multiple sclerosis patients are seropositive, is also interesting from an epidemiological point of view. We have reported a cluster of MS patients with 8 members from a small Danish community called Fjelsø. To further evaluate the role of EBV in MS we have investigated the distribution of EBV subtypes in cluster members and in control cohorts. MATERIALS AND METHODS: Blood mononuclear cells were isolated from cluster members, unrelated MS patients, healthy controls, including healthy schoolmates to the Fjelsø cluster patients and finally from persons with autoimmune diseases in order to investigate the number of 39 bp repeats in the EBNA 6-coding region in the EBV seropositive individuals. RESULTS: We observed a preponderance of the subtype with 3 39 bp repeats in the EBNA 6-coding region both in the MS patients and the healthy controls. In the Fjels cluster all 8 cluster members were harbouring this subtype, which is significantly different from the finding in healthy controls (n = 16), which include 8 schoolmates to the cluster members and 8 randomly selected healthy persons (Fischer's exact test P = 0.0047), and also compared to all non-clustered individuals studied (P = 0.017). CONCLUSION: Infection with the same subtype of EBV links together the 8 persons from the Fjelsø cluster who later developed MS. This finding adds to the possibility that development of MS is linked to infection with EBV.

Adult↗

Clusters of multiple different small nucleolar RNA genes in plants are expressed as and processed from polycistronic pre-snoRNAs.

Small nucleolar RNAs (snoRNAs) are involved in many aspects of rRNA processing and maturation. In animals and yeast, a large number of snoRNAs are encoded within introns of protein-coding genes. These introns contain only single snoRNA genes and their processing involves exonucleolytic release of the snoRNA from debranched intron lariats. In contrast, some U14 genes in plants are found in small clusters and are expressed polycistronically. An examination of U14 flanking sequences in maize has identified four additional snoRNA genes which are closely linked to the U14 genes. The presence of seven and five snoRNA genes respectively on 2.05 and 0.97 kb maize genomic fragments further emphasizes the novel organization of plant snoRNA genes as clusters of multiple different genes encoding both box C/D and box H/ACA snoRNAs. The plant snoRNA gene clusters are transcribed as a polycistronic pre-snoRNA transcript from an upstream promoter. The lack of exon sequences between the genes suggests that processing of polycistronic pre-snoRNAs involves endonucleolytic activity. Consistent with this, U14 snoRNAs can be processed from both non-intronic and intronic transcripts in tobacco protoplasts such that processing is splicing independent.

Base Sequence↗

A genetic marker and family history study of the upstate New York multiple sclerosis cluster.

We report nine additional cases of new-onset multiple sclerosis (MS) among employees of an upstate New York manufacturing plant that uses zinc as a primary metal. These cases, identified during the decade 1980 to 1989, had clinical onset of the disease between 1979 and 1987. The new cases confirm the increased incidence of MS previously reported in the plant population for the 1970 to 1979 decade. The MS subjects in this occupationally based cluster do not seem different from other MS patients with regard to rates of familial MS or the frequencies of alleles for human leukocyte (HLA-DR) antigens or transferrin. The frequency distribution of alleles for transferrin (an iron- and zinc-binding protein) may differ in these and other MS subjects compared with controls.

Adult↗

Horizontally oriented clusters of multiple chondrons in the superficial zone of ankle, but not knee articular cartilage.

Osteoarthritis is a progressive disease that is initiated at the surface of articular cartilage and proceeds to destroy the entire depth of the cartilage. The prevalence of osteoarthritis varies in different joints; e.g., the ankle joint has a very low prevalence of the disease compared to the knee joint. To better understand any inherent differences between the articular cartilage of the ankle and that of the knee that would account for the difference in occurrence of osteoarthritis, studies were undertaken to examine differences between the superficial zones in these two joint cartilages obtained from human donors. Chondrocytes in the superficial zones of the normal ankle (talocrural) and the normal knee (tibiofemoral) joints were identified with a monoclonal antibody specific for the superficial zone protein (SZP). When the chondrocytes from both joints were compared in serial horizontal sections, the chondrocytes in the superficial zone of the knee cartilage were seen either as isolated single cells or as doublets. However, the chondrocytes within the superficial zone of normal ankle cartilage were arranged in planar clusters containing multiple chondrons composed of 2-13 cells. There were no detectable differences in the chondrocyte clusters in the superficial zone of the ankle with respect to age, gender, or site on the cartilage surface. Adjacent to a lesion in an ankle joint with degenerative changes, the clusters were larger, containing up to 22 chondrocytes. This is the first report documenting the presence of multiple chondrons in the superficial zone of normal human adult articular cartilage.

Adolescent↗

Automatic detection of conserved gene clusters in multiple genomes by graph comparison and P-quasi grouping.

We previously reported two graph algorithms for analysis of genomic information: a graph comparison algorithm to detect locally similar regions called correlated clusters and an algorithm to find a graph feature called P-quasi complete linkage. Based on these algorithms we have developed an automatic procedure to detect conserved gene clusters and align orthologous gene orders in multiple genomes. In the first step, the graph comparison is applied to pairwise genome comparisons, where the genome is considered as a one-dimensionally connected graph with genes as its nodes, and correlated clusters of genes that share sequence similarities are identified. In the next step, the P-quasi complete linkage analysis is applied to grouping of related clusters and conserved gene clusters in multiple genomes are identified. In the last step, orthologous relations of genes are established among each conserved cluster. We analyzed 17 completely sequenced microbial genomes and obtained 2313 clusters when the completeness parameter P: was 40%. About one quarter contained at least two genes that appeared in the metabolic and regulatory pathways in the KEGG database. This collection of conserved gene clusters is used to refine and augment ortholog group tables in KEGG and also to define ortholog identifiers as an extension of EC numbers.

Algorithms↗