PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Transcriptional profiling of wheat caryopsis development using cDNA microarrays.

The expression of 7,835 genes in developing wheat caryopses was analyzed using cDNA arrays. Using a mixed model analysis of variance (ANOVA) method, 29% (2,237) of the genes on the array were identified to be differentially expressed at the 6 different time-points examined, which covers the developmental stages from coenocytic endosperm to physiological maturity. Comparison of genes differentially expressed between two time-points revealed a dynamic transcript accumulation profile with major re-programming events that occur at 3-7, 7-14 and 21-28 DPA. A k-means clustering algorithm grouped the differentially expressed genes into 10 clusters, revealing co-expression of genes involved in the same pathway such as carbohydrate and protein synthesis or preparation for desiccation. Functional annotation of genes that show peak expression at specific time-points correlated with the developmental events associated with the respective stages. Results provide information on the temporal expression during caryopsis development for a significant number of differentially expressed genes with unknown function.

DNA, Complementary↗

Measuring spike pattern reliability with the Lempel-Ziv-distance.

Spike train distance measures serve two purposes: to measure neuronal firing reliability, and to provide a metric with which spike trains can be classified. We introduce a novel spike train distance based on the Lempel-Ziv complexity that does not require the choice of arbitrary analysis parameters, is easy to implement, and computationally cheap. We determine firing reliability in vivo by calculating the deviation of the mean distance of spike trains obtained from multiple presentations of an identical stimulus from a Poisson reference. Using both the Lempel-Ziv-distance (LZ-distance) and a distance focussing on coincident firing, the pattern and timing reliability of neuronal firing is determined for spike data obtained along the visual information processing pathway of macaque monkey (LGN, simple and complex cells of V1, and area MT). In combination with the sequential superparamagnetic clustering algorithm, we show that the LZ-distance groups together spike trains with similar but not necessarily synchronized firing patterns. For both applications, we show how the LZ-distance gives additional insights, as it adds a new perspective on the problem of firing reliability determination and allows neuron classifications in cases, where other distance measures fail.

Algorithms↗

Small molecule affinity fingerprinting. A tool for enzyme family subclassification, target identification, and inhibitor design.

Classifying proteins into functionally distinct families based only on primary sequence information remains a difficult task. We describe here a method to generate a large data set of small molecule affinity fingerprints for a group of closely related enzymes, the papain family of cysteine proteases. Binding data was generated for a library of inhibitors based on the ability of each compound to block active-site labeling of the target proteases by a covalent activity based probe (ABP). Clustering algorithms were used to automatically classify a reference group of proteases into subfamilies based on their small molecule affinity fingerprints. This approach was also used to identify cysteine protease targets modified by the ABP in complex proteomes by direct comparison of target affinity fingerprints with those of the reference library of proteases. Finally, experimental data were used to guide the development of a computational method that predicts small molecule inhibitors based on reported crystal structures. This method could ultimately be used with large enzyme families to aid in the design of selective inhibitors of targets based on limited structural/function information.

Algorithms↗

Comparison of gene expression profiling between malignant and normal plasma cells with oligonucleotide arrays.

The DNA microarray technology enables the identification of the large number of genes involved in the complex deregulation of cell homeostasis taking place in cancer. Using Affymetrix microarrays, we have compared the gene expression profiles of highly purified malignant plasma cells from nine patients with multiple myeloma (MM) and eight myeloma cell lines to those of highly purified nonmalignant plasma cells (eight samples) obtained by in vitro differentiation of peripheral blood B cells. Two unsupervised clustering algorithms classified these 25 samples into two distinct clusters: a malignant plasma cell cluster and a normal plasma cell cluster. Two hundred and fifty genes were significantly up-regulated and 159 down-regulated in malignant plasma samples compared to normal plasma samples. For some of these genes, an overexpression or downregulation of the encoded protein was confirmed (cyclin D1, c-myc, BMI-1, cystatin c, SPARC, RB). Two genes overexpressed in myeloma cells (ABL and cystathionine beta synthase) code for enzymes that could be a therapeutic target with specific drugs. These data provide a new insight into the understanding of myeloma disease and prefigure that the development of DNA microarray could help to develop an 'à la carte' treatment in cancer disease.

Adult↗

A classification-based machine learning approach for the analysis of genome-wide expression data.

Three important areas of data analysis for global gene expression analysis are class discovery, class prediction, and finding dysregulated genes (biomarkers). The clinical application of microarray data will require marker genes whose expression patterns are sufficiently well understood to allow accurate predictions on disease subclass membership. Commonly used methods of analysis include hierarchical clustering algorithms, t-, F-, and Z-tests, and machine learning approaches. We describe an approach called the maximum difference subset (MDSS) algorithm that combines classification algorithms, classical statistics, and elements of machine learning and provides a coherent framework. By integrating prediction accuracy, the MDSS algorithm learns the critical threshold of statistical significance (the alpha or P-value), eliminating the arbitrariness of setting a threshold of statistical significance and minimizing the effect of the normality assumptions. To reduce the false positive rate and to increase external validity of the predictive gene set, a jackknife step is used. This step identifies and removes genes in the initial MDSS with low combined predictive utility. The overall MDSS provides a prediction that is less dependent on an arbitrary study design (sample inclusion or exclusion) and should thus have high external validity. We demonstrate that this approach, unlike other published methods, identifies biomarkers capable of predicting the outcome of anthracycline-cytarabine chemotherapy in cases of acute myeloid leukemia. By incorporating two criteria-statistical significance and predictive utility-the approach learns the significance level relevant for a given data set. The MDSS approach can be used with any test and classifier operator pair.

Acute Disease↗

Analysis of the weighting exponent in the FCM.

The fuzzy c-means (FCM) algorithm is one of the most frequently used clustering algorithms. The weighting exponent m is a parameter that greatly influences the performance of the FCM. But there has been no theoretical basis for selecting the proper weighting exponent in the literature. In this paper, we develop a new theoretical approach to selecting the weighting exponent in the FCM. Based on this approach, we reveal the relation between the stability of the fixed points of the FCM and the data set itself. This relation provides the theoretical basis for selecting the weighting exponent in the FCM. The numerical experiments verify the effectiveness of our theoretical conclusion.

Journal Article↗

Metal artifact reduction in CT using tissue-class modeling and adaptive prefiltering.

High-density objects such as metal prostheses, surgical clips, or dental fillings generate streak-like artifacts in computed tomography images. We present a novel method for metal artifact reduction by in-painting missing information into the corrupted sinogram. The information is provided by a tissue-class model extracted from the distorted image. To this end the image is first adaptively filtered to reduce the noise content and to smooth out streak artifacts. Consecutively, the image is segmented into different material classes using a clustering algorithm. The corrupted and missing information in the original sinogram is completed using the forward projected information from the tissue-class model. The performance of the correction method is assessed on phantom images. Clinical images featuring a broad spectrum of metal artifacts are studied. Phantom and clinical studies show that metal artifacts, such as streaks, are significantly reduced and shadows in the image are eliminated. Furthermore, the novel approach improves detectability of organ contours. This can be of great relevance, for instance, in radiation therapy planning, where images affected by metal artifacts may lead to suboptimal treatment plans.

Algorithms↗

Automated diagnosis of brain tumours astrocytomas using probabilistic neural network clustering and support vector machines.

A computer-aided diagnosis system was developed for assisting brain astrocytomas malignancy grading. Microscopy images from 140 astrocytic biopsies were digitized and cell nuclei were automatically segmented using a Probabilistic Neural Network pixel-based clustering algorithm. A decision tree classification scheme was constructed to discriminate low, intermediate and high-grade tumours by analyzing nuclear features extracted from segmented nuclei with a Support Vector Machine classifier. Nuclei were segmented with an average accuracy of 86.5%. Low, intermediate, and high-grade tumours were identified with 95%, 88.3%, and 91% accuracies respectively. The proposed algorithm could be used as a second opinion tool for the histopathologists.

Algorithms↗

cDNA microarray analysis of bovine embryo gene expression profiles during the pre-implantation period.

BACKGROUND: After fertilization, embryo development involves differentiation, as well as development of the fetal body and extra-embryonic tissues until the moment of implantation. During this period various cellular and molecular changes take place with a genetic origin, e.g. the elongation of embryonic tissues, cell-cell contact between the mother and the embryo and placentation. To identify genetic profiles and search for new candidate molecules involved during this period, embryonic gene expression was analyzed with a custom designed utero-placental complementary DNA (cDNA) microarray. METHODS: Bovine embryos on days 7, 14 and 21, extra-embryonic membranes on day 28 and fetuses on days 28 were collected to represent early embryo, elongating embryo, pre-implantation embryo, post-implantation extra-embryonic membrane and fetus, respectively. Gene expression at these different time points was analyzed using our cDNA microarray. Two clustering algorithms such as k-means and hierarchical clustering methods identified the expression patterns of differentially expressed genes across pre-implantation period. Novel candidate genes were confirmed by real-time RT-PCR. RESULTS: In total, 1,773 individual genes were analyzed by complete k-means clustering. Comparison of day 7 and day 14 revealed most genes increased during this period, and a small number of genes exhibiting altered expression decreased as gestation progressed. Clustering analysis demonstrated that trophoblast-cell-specific molecules such as placental lactogens (PLs), prolactin-related proteins (PRPs), interferon-tau, and adhesion molecules apparently all play pivotal roles in the preparation needed for implantation, since their expression was remarkably enhanced during the pre-implantation period. The hierarchical clustering analysis and RT-PCR data revealed new functional roles for certain known genes (dickkopf-1, NPM, etc) as well as novel candidate genes (AW464053, AW465434, AW462349, AW485575) related to already established trophoblast-specific genes such as PLs and PRPs. CONCLUSIONS: A large number of genes in extra-embryonic membrane increased up to implantation and these profiles provide information fundamental to an understanding of extra-embryonic membrane differentiation and development. Genes in significant expression suggest novel molecules in trophoblast differentiation.

Animals↗

NIPALSTREE: a new hierarchical clustering approach for large compound libraries and its application to virtual screening.

A hierarchical clustering algorithm--NIPALSTREE--was developed that is able to analyze large data sets in high-dimensional space. The result can be displayed as a dendrogram. At each tree level the algorithm projects a data set via principle component analysis onto one dimension. The data set is sorted according to this one dimension and split at the median position. To avoid distortion of clusters at the median position, the algorithm identifies a potentially more suited split point left or right of the median. The procedure is recursively applied on the resulting subsets until the maximal distance between cluster members exceeds a user-defined threshold. The approach was validated in a retrospective screening study for angiotensin converting enzyme (ACE) inhibitors. The resulting clusters were assessed for their purity and enrichment in actives belonging to this ligand class. Enrichment was observed in individual branches of the dendrogram. In further retrospective virtual screening studies employing the MDL Drug Data Report (MDDR), COBRA, and the SPECS catalog, NIPALSTREE was compared with the hierarchical k-means clustering approach. Results show that both algorithms can be used in the context of virtual screening. Intersecting the result lists obtained with both algorithms improved enrichment factors while losing only few chemotypes.

Algorithms↗

Optimization of tissue segmentation of brain MR images based on multispectral 3D feature maps.

The purpose of this work was to optimize and increase the accuracy of tissue segmentation of the brain magnetic resonance (MR) images based on multispectral 3D feature maps. We used three sets of MR images as input to the in-house developed semi-automated 3D tissue segmentation algorithm: proton density (PD) and T2-weighted fast spin echo and, T1-weighted spin echo. First, to eliminate the random noise, non-linear anisotropic diffusion type filtering was applied to all the images. Second, to reduce the nonuniformity of the images, we devised and applied a correction algorithm based on uniform phantoms. Following these steps, the qualified observer "seeded" (identified training points) the tissue of interest. To reduce the operator dependent errors, cluster optimization was also used; this clustering algorithm identifies the densest clusters pertaining to the tissues. Finally, the images were segmented using k-NN (k-Nearest Neighborhood) algorithm and a stack of color-coded segmented images were created along with the connectivity algorithm to generate the entire surface of the brain. The application of pre-processing optimization steps substantially improved the 3D tissue segmentation methodology.

Anisotropy↗

A cluster-based strategy for assessing the overlap between large chemical libraries and its application to a recent acquisition.

We report on the structural comparison of the corporate collections of Johnson & Johnson Pharmaceutical Research & Development (JNJPRD) and 3-Dimensional Pharmaceuticals (3DP), performed in the context of the recent acquisition of 3DP by JNJPRD. The main objective of the study was to assess the druglikeness of the 3DP library and the extent to which it enriched the chemical diversity of the JNJPRD corporate collection. The two databases, at the time of acquisition, collectively contained more than 1.1 million compounds with a clearly defined structural description. The analysis was based on a clustering approach and aimed at providing an intuitive quantitative estimate and visual representation of this enrichment. A novel hierarchical clustering algorithm called divisive k-means was employed in combination with Kelley's cluster-level selection method to partition the combined data set into clusters, and the diversity contribution of each library was evaluated as a function of the relative occupancy of these clusters. Typical 3DP chemotypes enriching the diversity of the JNJPRD collection were catalogued and visualized using a modified maximum common substructure algorithm. The joint collection of JNJPRD and 3DP compounds was also compared to other databases of known medicinally active or druglike compounds. The potential of the methodology for the analysis of very large chemical databases is discussed.

Algorithms↗

Statistical mechanics of histories: a cluster Monte Carlo algorithm.

We present an efficient computational approach to sample the histories of nonlinear stochastic processes. This framework builds upon recent work on casting a d-dimensional stochastic dynamical system into a (d+1)-dimensional equilibrium system using the path-integral approach. We introduce a cluster algorithm that efficiently samples histories and discuss how to include measurements that are available into the estimate of the histories. This allows our approach to be applicable to the simulation of rare events and to optimal state and parameter estimation. We demonstrate the utility of this approach for Phi4 Langevin dynamics in two spatial dimensions where our algorithm improves sampling efficiency up to an order of magnitude.

Journal Article↗

Phenotypic study by numerical taxonomy of strains belonging to the genus Aeromonas.

AIMS: This study was undertaken to cluster and identify a large collection of Aeromonas strains. METHODS AND RESULTS: Numerical taxonomy was used to analyse phenotypic data obtained on 54 new isolates taken from water, fish, snails, sputum and 99 type and reference strains. Each strain was tested for 121 characters but only the data for 71 were analysed using the 'SSM' and 'SJ' coefficients, and the UPGMA clustering algorithm. At SJ values of > or = 81.6% the strains clustered into 22 phenons which were identified as Aer. jandaei, Aer. hydrophila, Aer. encheleia, Aer. veronii biogroup veronii, Aer. trota, Aer. caviae, Aer. eucrenophila, Aer. ichthiosmia, Aer. sobria, Aer. allosaccharophila, Aer. media, Aer. schubertii and Aer. salmonicida. The species Aer. veronii biogroup sobria was represented by several clusters which formed two phenotypic cores, the first related to reference strain CECT 4246 and the second related to CECT 4835. A good correlation was generally observed among this phenotypic clustering and previous genomic and phylogenetic data. In addition, three new phenotypic groups were found, which may represent new Aeromonas species. CONCLUSIONS: The phenetic approach was found to be a necessary tool to delimitate and identify the Aeromonas species. SIGNIFICANCE AND IMPACT OF THE STUDY: Valuable traits for identifying Aeromonas as well as the possible existence of new Aeromonas species or biotypes are indicated.

Aeromonas↗

The comparison of clinical imaging devices with respect to parallel readings in both devices.

OBJECTIVE: Many proposals for the comparison of diagnostic devices refer to the computation of ROC curves or sensitivity / specificity-based parameters, thereby strictly assuming the presence of a reliably parameterized clinical reference method. When none of the devices under consideration can be regarded as a reference, Cohen's kappa coefficient for assessing the methods' relative agreement becomes increasingly popular. If, however, not only the agreement between two diagnostic devices, but also the devices' reliability must be taken into account (for example, if multiple parallel readings are obtained from one or both of the devices), no corresponding coefficients can be obtained from standard software. Bearing the recent modifications in the German Medicinal Devices Law (Medizinproduktegesetz) in mind, such methods will soon become necessary and strongly demanded for the sake of immediate re-evaluation of previously certified medicinal devices. METHODS: Generalizations of Cohen's kappa (kappa) for complex multi reader designs can be found by estimating weighted averages of the observed and expected agreement among subsets of parallel readings. A flexible, although instructive, strategy for designing kappa coefficients in the context of method comparison trials is proposed, which measures the two methods' overall agreement while correcting for each method's underlying inter / intra observer reliability. Cluster algorithms will be outlined, which allow to identify (in)compatible clusters of readings. Their application will be illustrated by means of the intraindividual comparison of two different strategies in radiographical imaging, where none of the underlying imaging methods can be regarded as a reference. RESULTS: The algorithms are illustrated by the comparison of two radiological imaging devices R and F, where none of these imaging methods could be considered as a valid reference, i.e. replicate readings by three independent radiologists were taken from each device, respectively. The setting allowed for intraindividual comparison of the imaging methods, since each of the three involved radiologists took one reading from both devices on each of 120 individuals. The algorithm identifies a subset of compatible reading patterns with an overall agreement of kappa = 0.83 (95% confidence interval 0.78 - 0.88) despite the fact, that the underlying readings arose from two different imaging devices. An obvious interpretation suggests, that the gradient in experience between the readers was more relevant to their reading patterns' outcome than any difference between the imaging devices. CONCLUSIONS: The generalized kappa coefficients can be modified according to the study design at hand to instructively identify (in)compatible clusters of multiple parallel reading patterns; the relative agreement of imaging methods can be estimated as well as each imaging method's internal reliability as assessed by parallel readings from the respective methods.

Algorithms↗

Biosphere: the interoperation of web services in microarray cluster analysis.

UNLABELLED: The growing use of DNA microarrays in biomedical research has led to the proliferation of analysis tools. These software programs address different aspects of analysis (e.g. normalisation and clustering within and across individual arrays) as well as extended analysis methods (e.g. clustering, annotation and mining of multiple datasets). Therefore, microarray data analysis typically requires the interoperability of multiple software programs involving different analysis types and methods. Such interoperation is often hampered by the heterogeneity inherent in the software tools (which may function by implementing different interfaces and using different programming languages). To address this problem, we employed the simple object access protocol (SOAP)-based web service approach that provides a uniform programmatic interface to these heterogeneous software components. To demonstrate this approach in the microarray context, we created a web server application, Biosphere, which interoperates a number of web services that are geographically widely distributed. These web services include a clustering web service, which is a suite of different clustering algorithms for analysing microarray data; XEMBL, developed at the European Bioinformatics Institute (EBI) for retrieving EMBL Nucleotide Sequence Database sequence data; and three gene annotation web services: GetGO, GetHAPI and GetUMLS. GetGO allows retrieval of Gene Ontology (GO) annotation, and the other two web services retrieve annotation from the biomedical literature that is indexed based on the Medical Subject Headings (MeSH) terms. With these web services, Biosphere allows the users to do the following: (i) cluster gene expression data using seven different algorithms; (ii) visualise the clustering results that are grouped statistically in colour; and (iii) retrieve sequence, annotation and citation data for the genes of interest. AVAILABILITY: Biosphere and its web services described in Web Service Description Language (WSDL) can be accessed at http://rook.cecid.hku.hk:8280/BiosphereServer.

Cluster Analysis↗

Multivariate analysis of antibiograms for typing Pseudomonas aeruginosa.

A method for typing Pseudomonas aeruginosa using antibiotic susceptibility patterns is presented, which allows recognition of clusters of the same strain among clinical isolates from different patients, thus indicating whether cross infection has occurred. An index of similarity (the euclidean or the oblique distance), which includes all the differences of disk zone sizes among isolates, is computed and then elaborated by a clustering algorithm that successively groups all the isolates in larger clusters. The results of clustering are presented as dendrograms, whose terminal branches are pruned down to a level below which differences are casual; isolates that still appear on a common branch are considered identical. The reliability of this technique for detecting nosocomial cross infections was assessed by comparing its results with that of serotyping and pyocin typing. Only 2 of 31 (6.4%) clusters detected by multivariate analysis were not confirmed, while 4 of 33 (12.1%) clusters were recognized by serotyping and pyocin typing, but not by multivariate analysis. In at least two instances the differences in susceptibility patterns were due to cytoplasmic R factors. The routine use of antibiogram data for typing purposes should be considered an essential part of nosocomial infection control.

Anti-Bacterial Agents↗

Gene-Ontology-based clustering of gene expression data.

UNLABELLED: The expected correlation between genetic co-regulation and affiliation to a common biological process is not necessarily the case when numerical cluster algorithms are applied to gene expression data. GO-Cluster uses the tree structure of the Gene Ontology database as a framework for numerical clustering, and thus allowing a simple visualization of gene expression data at various levels of the ontology tree. AVAILABILITY: The 32-bit Windows application is freely available at http://www.mpibpc.mpg.de/go-cluster/

Cluster Analysis↗