PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Distinguishing key biological pathways between primary breast cancers and their lymph node metastases by gene function-based clustering analysis.

In order to identify key biological pathways that can distinguish between primary breast cancers and their lymph node metastases, we employed gene expression profiling together with gene function-based clustering analysis. We first acquired gene expression profiles of 9 matched primary tumors and the corresponding metastases that contained at least 75% of tumor cells. Then, we applied a clustering algorithm to the preprocessed data. In order to focus on the most informative genes, we ranked all the genes individually based on their abilities to separate the primary breast tumor and metastases samples. Further, we separated these genes into six functional groups according to the Stanford SOURCE database: 'cell cycle,' 'apoptosis,' 'metabolism,' 'cell adhesion and migration,' 'signal transduction,' and 'transcriptional factor and DNA binding molecules.' Unsupervised clustering analysis using all of the 2,303 genes on the microarrays was not able to separate the primary and metastases samples. Clustering analysis using the most informative genes revealed that primary tumors were more tightly clustered, whereas the metastases samples were relatively heterogeneous. The clustering analysis with the genes belonging to different functional groups showed that different functional gene sets varied in their abilities to separate primary tumors and their metastases. Marked separations were found with genes involved in metabolism, signal transduction, cell cycle, and transcriptional factor and DNA binding molecules. In contrast, apoptosis and cell adhesion and migration genes did not provide a clear separation of the two groups of samples. These results suggest that metastatic cells have different metabolism and signal transduction activities, regulated by transcriptional events, from the primary tumor cells. The results also suggest that the altered cell adhesion and migration potentials that are required for tumors to metastasize already exist in the primary tumors as a whole.

Biomarkers, Tumor↗

Standardization and interlaboratory reproducibility assessment of pulsed-field gel electrophoresis-generated fingerprints of Acinetobacter baumannii.

A standard procedure for pulsed-field gel electrophoresis (PFGE) of macrorestriction fragments of Acinetobacter baumannii was set up and validated for its interlaboratory reproducibility and its potential for use in the construction of an Internet-based database for international monitoring of epidemic strains. The PFGE fingerprints of strains were generated at three different laboratories with ApaI as the restriction enzyme and by a rigorously standardized procedure. The results were analyzed at the respective laboratories and also centrally at a national reference institute. In the first phase of the study, 20 A. baumannii strains, including 3 isolates each from three well-characterized hospital outbreaks and 11 sporadic strains, were distributed blindly to the participating laboratories. The local groupings of the isolates in each participating laboratory were identical and allowed the identification of the epidemiologically related isolates as belonging to three clusters and identified all unrelated strains as distinct. Central pattern analysis by using the band-based Dice coefficient and the unweighted pair group method with mathematical averaging as the clustering algorithm showed 95% matching of the outbreak strains processed at each local laboratory and 87% matching of the corresponding strains if they were processed at different laboratories. In the second phase of the study, 30 A. baumannii isolates representing 10 hospital outbreaks from different parts of Europe (3 isolates per outbreak) were blindly distributed to the three laboratories, so that each laboratory investigated 10 epidemiologically independent outbreak isolates. Central computer-assisted cluster analysis correctly identified the isolates according to their corresponding outbreak at an 87% clustering threshold. In conclusion, the standard procedure enabled us to generate PFGE fingerprints of epidemiologically related A. baumannii strains at different locations with sufficient interlaboratory reproducibility to set up an electronic database to monitor the geographic spread of epidemic strains.

Acinetobacter Infections↗

A fully automatic multimodality image registration algorithm.

OBJECTIVE: A fully automatic multimodality image registration algorithm is presented. The method is primarily designed for 3D registration of MR and PET images of the brain. However, it has also been successfully applied to CT-PET, MR-CT, and MR-SPECT registrations. MATERIALS AND METHODS: The head contour is detected on the MR image using a gradient threshold method. The head region in the MR image is then segmented into a set of connected components using the K-means clustering algorithm. When the two image sets are registered, the segmentation of the MR image indirectly generates a segmentation of the PET image. The best registration is taken to be the one that optimizes the segmentation induced on the PET image. In this article, the K-means minimum variance criterion is used as a cost function, and the optimization is performed using the method of coordinate descent. RESULTS: The algorithm was tested on 80 H2 15O PET and MR image pairs from 10 subjects. Qualitatively correct results were obtained in all cases. With use of external markers visible in both image modalities, the average registration error was estimated to be < 3 mm. CONCLUSION: The algorithm presented in this article requires no user interaction and can be applied to a wide range of registration problems. Quantitative and qualitative evaluations of the algorithm indicate a high degree of accuracy.

Algorithms↗

Erythroid-induced commitment of K562 cells results in clusters of differentially expressed genes enriched for specific transcription regulatory elements.

Understanding regulation of fetal and embryonic hemoglobin expression is critical, since their expression decreases clinical severity in sickle cell disease and beta-thalassemia. K562 cells, a human erythroleukemia cell line, can differentiate along erythroid or megakaryocytic lineages and serve as a model for regulation of fetal/embryonic globin expression. We used microarray expression profiling to characterize transcriptomes from K562 cells treated for various times with hemin, an inducer of erythroid commitment. Approximately 5,000 genes were expressed irrespective of treatment. Comparative expression analysis (CEA) identified 899 genes as differentially expressed; analysis by the self-organizing map (SOM) algorithm clustered 425 genes into 8 distinct expression patterns, 322 of which were shared by both analyses. Differential expression of a subset of genes was validated by real-time RT-PCR. Analysis of 5'-flanking regions from differentially expressed genes by PAINT v3.0 software showed enrichment in specific transcription regulatory elements (TREs), some localizing to different expression clusters. This finding suggests coordinate regulation of cluster members by specific TREs. Finally, our findings provide new insights into rate-limiting steps in the appearance of heme-containing hemoglobin tetramers in these cells.

5' Flanking Region↗

Trained artificial neural network for glaucoma diagnosis using visual field data: a comparison with conventional algorithms.

PURPOSE: To evaluate and confirm the performance of an artificial neural network (ANN) trained to recognize glaucomatous visual field defects, and compare its diagnostic accuracy with that of other algorithms proposed for the detection of visual field loss. METHODS: SITA Standard 30-2 visual fields, from 100 glaucoma patients and 116 healthy participants, formed the data set. Our ANN was a previously described fully trained network using scored pattern deviation probability maps as input data. Its diagnostic accuracy was compared to that of the Glaucoma Hemifield Test, the Pattern Standard Deviation index at the P<5% and <1%, and also to a technique based on the recognizing clusters of significantly depressed test points. RESULTS: The included tests had early to moderate visual field loss (median MD=-6.16 dB). ANN achieved a sensitivity of 93% at a specificity level of 94% with an area under the receiver operating characteristic curve of 0.984. Glaucoma Hemifield Test attained a sensitivity of 92% at 91% specificity. Pattern Standard Deviation, with a cut off level at P<5% had a sensitivity of 89% with a specificity of 93%, whereas at P<1% the sensitivity and specificity was 72% and 97%, respectively. The cluster algorithm yielded a sensitivity of 95% and a specificity of 82%. CONCLUSIONS: The high diagnostic performance of our ANN based on refined input visual field data was confirmed in this independent sample. Its diagnostic accuracy was slightly to considerably better than that of the compared algorithms. The results indicate the large potential for ANN as an important clinical glaucoma diagnostic tool.

Adult↗

Advances in the segmentation of multi-component microanalytical images.

Segmenting multi-component microanalytical images consists in trying to find zones of the specimen with approximate homogeneous composition, representing different chemical phases. This can be done through pixel clustering. We first highlight some limitations of classical clustering algorithms (C-means and fuzzy C-means). Then, we describe a new algorithm we have contributed to develop: the Parzen-watersheds algorithm. This algorithm is based on the estimation of the probability density function of the whole data set in the feature space (through the Parzen approach) and its partitioning using a method inherited from mathematical morphology: the watersheds method. Next, we introduce a fuzzy version of this approach, where the pixels are characterized by their grades of membership to the different classes. Finally, we show how the definition of the grades of membership can be used to improve the results of clustering, through probabilistic relaxation in the image space. The different methods presented are illustrated through an example in the field of electron energy loss mapping, where four elemental maps are concentrated in a single chemical phase map.

Journal Article↗

Numerical Taxonomic Analysis of Some Strains of Rhizobium spp. That Uses a Qualitative Coding of Immunodiffusion Reactions.

Antigenic relationships among seven strains of Bradyrhizobium japonicum were examined by immunodiffusion reactions, in which cells of each strain were reacted against each of the seven corresponding antisera. Similar analyses were performed with Rhizobium trifolii (28 strains), Rhizobium meliloti (9 strains), and rhizobia of the cowpea miscellany (13 strains). Antigens and antisera were reacted within each species only; serological interspecies cross-reactions were not performed. The results, scored qualitatively as reactions of identity, cross-reactions, or no reaction, were formed into datum matrices and used to analyze the relationships between strains by applying the association measure of Bray and Curtis (J. R. Bray and J. T. Curtis, Ecol. Monogr. 27:325-349, 1957) and the UPGMA clustering algorithm (P. H. A. Sneath and R. R. Sokal, Numerical Taxonomy, 1973). No two strains were regarded as being serologically identical unless each gave the same results as the other in each immunodiffusion reaction against every antiserum. Despite the high level of cross-reactions and reactions of identity (totalling 93% of all cell-antiserum combinations) among strains of R. trifolii and R. meliloti, no strains were identical by the criterion described above; however, the strains of these species clustered rapidly and fused at the 70% similarity level. The B. japonicum strains and the rhizobia of the cowpea miscellany were much less cross-reactive (67 and 86% of all combinations were negative, respectively), and they clustered more slowly. The strains of B. japonicum fused completely only at the 4% similarity level, whereas of the 13 cowpea-nodulating strains, 4 reacted as two pairs of identical strains and 6 remained unfused.

Journal Article↗

A nonparametric quantification of neural response field structures.

The response fields of higher cortical neurons are usually approximated with smooth mathematical functions for the purpose of population parameterization or theoretical modeling. We used instead two nonparametric methods (principal component analysis and independent component analysis), which provided a basis for the response field clustering. Although both methods performed satisfactorily, the principal component analysis space is more straightforward to calculate. It also gave a clear preference toward the smallest number of functional response field classes. Clustering was performed with both K-means and superparamagnetic clustering algorithms with similar results. We also show that the shapes of the eigenvectors remain consistent regardless of the response field data sets size. This finding reflects the fact that the response fields were generated by the same neural network and encode the same underlying process.

Brain Mapping↗

Unique gene expression profiles of human macrophages and dendritic cells to phylogenetically distinct parasites.

Monocyte-derived dendritic cells (DCs) and macrophages (Ms) generated in vitro from the same individual blood donors were exposed to 5 different pathogens, and gene expression profiles were assessed by microarray analysis. Responses to Mycobacterium tuberculosis and to phylogenetically distinct protozoan (Leishmania major, Leishmania donovani, Toxoplasma gondii) and helminth (Brugia malayi) parasites were examined, each of which produces chronic infections in humans yet vary considerably in the nature of the immune responses they trigger. In the absence of microbial stimulation, DCs and Ms constitutively expressed approximately 4000 genes, 96% of which were shared between the 2 cell types. In contrast, the genes altered transcriptionally in DCs and Ms following pathogen exposure were largely cell specific. Profiling of the gene expression data led to the identification of sets of tightly coregulated genes across all experimental conditions tested. A newly devised literature-based clustering algorithm enabled the identification of functionally and transcriptionally homogenous groups of genes. A comparison of the responses induced by the individual pathogens by means of this strategy revealed major differences in the functionally related gene profiles associated with each infectious agent. Although the intracellular pathogens induced responses clearly distinct from the extracellular B malayi, they each displayed a unique pattern of gene expression that would not necessarily be predicted on the basis of their phylogenetic relationship. The association of characteristic functional clusters with each infectious agent is consistent with the concept that antigen-presenting cells have prewired signaling patterns for use in the response to different pathogens.

Animals↗

Sliding window discretization: a new method for multiple band matching of bacterial genotyping fingerprints.

Microbiologists have traditionally applied hierarchical clustering algorithms as their mathematical tool of choice to unravel the taxonomic relationships between micro-organisms. However, the interpretation of such hierarchical classifications suffers from being subjective, in that a variety of ad hoc choices must be made during their construction. On the other hand, the application of more profound and objective mathematical methods--such as the minimization of stochastic complexity--for the classification of bacterial genotyping fingerprints data is hampered by the prerequisite that such methods only act upon vectorized data. In this paper we introduce a new method, coined sliding window discretization, for the transformation of genotypic fingerprint patterns into binary vector format. In the context of an extensive amplified fragment length polymorphism (AFLP) data set of 507 strains from the Vibrionaceae family that has previously been analysed, we demonstrate by comparison with a number of other discretization methods that this new discretization method results in minimal loss of the original information content captured in the banding patterns. Finally, we investigate the implications of the different discretization methods on the classification of bacterial genotyping fingerprints by minimization of stochastic complexity, as it is implemented in the BinClass software package for probabilistic clustering of binary vectors. The new taxonomic insights learned from the resulting classification of the AFLP patterns will prove the value of combining sliding window discretization with minimization of stochastic complexity, as an alternative classification algorithm for bacterial genotyping fingerprints.

Bacteria↗

Selection of surrogate marker genes in primary central nervous system lymphomas for radio-chemotherapy by DNA array analysis of gene expression profiles.

Primary central nervous system lymphomas (PCNSLs) are extra nodal B-cell non-Hodgkin's lymphomas with primary manifestation in the brain, and their incidence has been increasing among both immunocompetent and immunocompromised populations. Samples of oligodendroglioma (n=5), glioblastoma (n=7), PCNSL (n=6), and normal brain (n=3) were studied (total of 21 samples) using cDNA array technology. The hierarchical clustering algorithm was used to obtain a phylogenetic tree, and it revealed a striking feature: PCNSL was clearly separated. The genes encoding laminin receptor 2, thioredoxin peroxidase, and elongation factor-1 were selected as specific genes in PCNSL by principal component analysis (PCA). When Mann-Whitney tests were performed to identify genes responsible for the differences between responders and non-responders to the treatment schedule for PCNSL, 76 known genes were found to show significantly different expression patterns between the two groups at the P<0.01 level. The two groups were clearly separated by the re-clustering method using the selected genes related to response to chemo-radiotherapy. This is the first report describing the gene expression profiles of PCNSL. In conclusion, accumulation of data with respect to the expression profiles of PCNSL specimens, clinicopathological data, susceptibility to treatment, and outcome will provide information for identifying optimal therapeutic modalities for individual patients and novel therapeutic targets.

Adult↗

Studies of potential cerebrospinal fluid molecular markers for Alzheimer's disease.

There is a need for a reliable, molecular-based ante mortem diagnostic test for Alzheimer's disease (AD). In this study, we examined the use of two-dimensional protein electrophoresis for generating molecular barcodes which may be useful for the clinical differentiation of AD patients from normals. We compared cerebrospinal fluid samples taken from AD patients with confirmed post mortem pathology to comparable specimens from normal volunteers. Using canonical correlation analysis, a panel of nine molecular markers were identified which segregated diseased cases from normal controls. Using the scaled volume image analysis variable, a principal factor analysis was also used to distinguish normal from AD spinal fluid, based on molecular markers identified using a heuristic clustering algorithm. The use of panels of molecular markers derived from proteomic analysis may offer the best prospect for developing molecular diagnostic tests for complex neurodegenerative disorders such as AD.

Adult↗

Gene expression profiles in esophageal adenocarcinoma.

BACKGROUND: The incidence of esophageal adenocarcinoma (EAC) has risen dramatically in the last two decades. As with other malignancies, changes in gene expression play a key role in the development and progression of these tumors. METHODS: Microarray analysis was used to study gene expression of 12,000 genes in EAC specimens. Adenocarcinoma tissue samples (n = 10) and controls of normal stomach (n = 6) and esophageal (n = 7) mucosa were collected fresh, then rapidly frozen in liquid nitrogen. The messenger ribonucleic acid (mRNA) from the samples was isolated, reverse transcribed, and used to generate biotin-labeled mRNA fragments, which were hybridized to Affymetrix U95 gene chips (AME Bioscience, Norway) for analysis. Additional samples analyzed included tissue containing dysplastic Barrett's epithelium from three patients, metastatic lymph nodes from two patients with EAC, one squamous carcinoma, and two esophageal cancer cell lines. Samples were segregated into groups with similar patterns of gene expression using clustering algorithms and gene sets that differentiated tumors from normal tissue were generated. RESULTS: There were 150 genes that were fourfold up regulated and 183 genes that were fourfold down regulated in the esophageal adenocarcinoma specimens, as compared to normal esophageal mucosa tissue controls. Using paired specimens (n = 5) and the paired t-test (p Value of 0.05) as a filter, only 64 genes were fourfold up regulated and 110 were fourfold down regulated. These groups included cytoskeletal, cell adhesion, tumor suppressor, and signal transduction genes. Hierarchical clustering segregated the samples into the expected divisions. The esophageal cancer cell lines, OE19 and OE33, clustered separately from the EAC specimens. Extremely high gene expression levels of the ERBB2 gene, seen in the microarray analysis of the 2 cell lines, correlated with amplification of the gene determined by Southern blotting. CONCLUSIONS: Gene expression patterns from a small subset of genes distinguish EAC specimens from normal controls. This technique can rapidly identify genes for targeted chemotherapeutic approaches to cancer treatment.

Adenocarcinoma↗

Properties of learning of a Fuzzy ART Variant.

This paper discusses a variation of the Fuzzy ART algorithm referred to as the Fuzzy ART Variant. The Fuzzy ART Variant is a Fuzzy ART algorithm that uses a very large choice parameter value. Based on the geometrical interpretation of the weights in Fuzzy ART, useful properties of learning associated with the Fuzzy ART Variant are presented and proven. One of these properties establishes an upper bound on the number of list presentations required by the Fuzzy ART Variant to learn an arbitrary list of input patterns. This bound is small and demonstrates the short-training time property of the Fuzzy ART Variant. Through simulation, it is shown that the Fuzzy ART Variant is as good a clustering algorithm as a Fuzzy ART algorithm that uses typical (i.e. small) values for the choice parameter.

Journal Article↗

Profiling of genes associated with transcriptional responses in mouse hippocampus after transient forebrain ischemia using high-density oligonucleotide DNA array.

Several cascades of changes in gene expression have been shown to be involved in the neuronal injury after transient cerebral ischemia; however, little is known about the profile of genes showing alteration of expression in a mouse model of transient forebrain ischemia. We analyzed the gene expression profile in the mouse hippocampus during 24 h of reperfusion, after 20 min of transient forebrain ischemia, using a high-density oligonucleotide DNA array. Using statistical filtration (Welch's ANOVA and Welch's t-test), we identified 25 genes with a more than 3.0-fold higher or lower level of expression on average, with statistical significance set at p<0.05, in at least one ischemia-reperfusion group than in the sham group. Using unsupervised clustering methods (hierarchical clustering and k-means clustering algorithms), we identified four types of gene expression pattern that may be associated with the response of cell populations in the hippocampus to an ischemic insult in this mouse model. Functional classification of the 25 genes demonstrated alterations of expression of several kinds of biological pathways, regulating transcription (Bhlhb2, Jun, c-fos, Egr1, Egr2, Fosb, Junb, Ifrd1, Neurod6), the cell cycle (c-fos, Fosb, Jun, Junb, Dusp1), stress response (Dusp1, Dnajb1, Dnaja4), chaperone activity (Dnajb1, Dnaja4) and cell death (Ptgs2, Gadd45g, Tdag51), in the mouse hippocampus by 24 h of reperfusion. Using hierarchical clustering analysis, we also found that the same 25 genes clearly discriminated between the sham group and the ischemia-reperfusion groups. The alteration of expression of 25 genes identified in this study suggests the involvement of these genes in the transcriptional response of cell populations in the mouse hippocampus after transient forebrain ischemia.

Analysis of Variance↗

Periodograms and pulse detection methods for pulsatile hormone data.

Pulse detection algorithms and spectral analysis are the two most common methods for analysing pulsatile hormone data. We compared a popular high quality pulse detection algorithm (CLUSTER) to spectral analysis on a data set comparing luteinizing hormone data in depressed and control women. For these data, periodogram analysis methods, in particular Fisher's periodicity test, were superior in distinguishing the groups. Extending the pulse detection method to include measures of intra-individual variability improved its discriminatory performance. The two methods complement each other.

Adult↗

Identity structure, narrative accounts, and commitment to a volunteer role.

Degree of commitment was explored in relation to core self and role-identity. Thirty-one American emergency medical technicians (EMTs) described themselves in the EMT role (EMT now) and the way they anticipated they would be in the future (EMT future) by selecting items from an adjective checklist. Participants also described "real me," "ideal me," and "ought me." Ratings of commitment and extranormative activity were also obtained. Finally, participants described a positive and a negative episode they had experienced as an EMT in an open-ended question that was coded for task and relational content. Each participant's checklist data set was individually analyzed using HICLAS, a clustering algorithm for binary data (P. DeBoeck, S. Rosenberg, & I. Van Mechelen, 1993). Results indicate that the similarity between EMT now and real me best predicted activity and the similarity between EMT future and real me best predicted commitment (positive correlations in both cases). Older, more experienced EMTs tended to describe positive episodes in relational terms, whereas younger, less experienced EMTs described positive experiences in task-oriented terms.

Adult↗

Prediction of functional modules based on gene distributions in microbial genomes.

We present a computational method for prediction of functional modules that can be directly applied to the newly sequenced microbial genomes for predicting gene functions and the component genes of biological pathways. We first quantify the functional relatedness among genes based on their distribution (i.e., their existences and orders) across multiple microbial genomes, and obtain a gene network in which every pair of genes is associated with a score representing their functional relatedness. We then apply a threshold-based clustering algorithm to this gene network, and obtain modules for each of which the number of genes is bounded from above by a pre-specified value and the component genes are more strongly functionally related to each other than genes across the predicted modules. Particularly, when the module size is bounded by 130, we obtain 167 functional modules covering 813 genes for Escherichia coli K12, and 138 functional modules covering 731 genes for Bacillus subtilis subsp. subtilis str. 168. We have used the gene ontology (GO) information to assess the prediction results. The GO similarities among the genes of the same functional module are compared with the GO similarities among the genes that are randomly clustered together. This comparison reveals that our predicted functional modules are statistically and biologically significant, and the genes of the same functional module share more commonality in terms of biological process than in terms of molecular function or cellular component. We have also examined the predicted functional modules that are common to both Escherichia coli K12 and Bacillus subtilis subsp. subtilis str. 168, and provide explanations for some functional modules.

Cluster Analysis↗