PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Adaptable fuzzy C-Means for improved classification as a preprocessing procedure of brain parcellation.

Parcellation, one of several brain analysis methods, is a procedure popular for subdividing the regions identified by segmentation into smaller topographically defined units. The fuzzy clustering algorithm is mainly used to preprocess parcellation into several segmentation methods, because it is very appropriate for the characteristics of magnetic resonance imaging (MRI), such as partial volume effect and intensity inhomogeneity. However, some gray matter, such as basal ganglia and thalamus, may be misclassified into the white matter class using the conventional fuzzy C-Means (FCM) algorithm. Parcellation has been nearly achieved through manual drawing, but it is a tedious and time-consuming process. We propose improved classification using successive fuzzy clustering and implementing the parcellation module with the modified graphic user interface (GUI) for the convenience of users.

Algorithms↗

Particle track structure and its correlation with radiobiological endpoint.

One of the possible ways to classify track structures is application of the conventional partition techniques of analysis of multidimensional data to the track structure. Using these cluster algorithms this paper attempts to find characteristics of radiation reflecting the spatial distribution of ionizations in the primary particle track. Absolute frequency distributions of clusters giving the mean number of clusters produced by radiation per unit of deposited energy have been computed for radiation of different qualities. The results were compared with the published experimental data of cell inactivation. For particular biological objects the critical properties of radiation correlating with the cell inactivation can be found and it seems that the occurrence of a cluster of at least four ionizations formed in a domain of approximately 2-3 nm correlates with the induction of double strand break.

Ions↗

Adaptive neuro-fuzzy inference system: an instant and architecture-free predictor for improved QSAR studies.

The application of an adaptive neuro-fuzzy inference system (ANFIS) has been developed for obtaining sufficient quantitative structure-activity relationships (QSAR) with high accuracy. To this end, a data set of 68 pyrimidines derivatives as DHFR inhibitors, described first in the excellent independent studies of Hansch et al. (J. Med. Chem. 1982, 25, 777-784 and J. Med. Chem. 1991, 34, 46-54) and later by So and Richards (J. Med. Chem. 1992, 35, 3201-3207), was examined. The ANFIS system, first time applied in the literature to QSAR studies, was trained using a hybrid algorithm consisting of back-propagation and least-squares estimation while the optimum number and shape of membership functions were obtained through the subtractive clustering algorithm. Prior to the development and evaluation of the ANFIS system, geometry optimization of the examined compounds was performed, deriving a series of diverse descriptors from which the best subset was selected by using a hybrid genetic algorithm system. The predictive abilities of the resulting models compared to those produced from classical multivariate regression such as linear and nonlinear (quadratic) partial least squares regression (PLS and QPLS, respectively). The ANFIS method outperformed both the PLS models as well as the published results, leading to substantial gain in both the prediction ability and the computation speed (almost instant training).

Algorithms↗

Cluster analysis applied to symptom ratings of psychiatric patients: an evaluation of its predictive ability.

Rating on 39 symptoms were examined for patients admitted to the Neuropsychiatric Institute of the University of Michigan Medical Center. A detailed evaluation was made of the clusters derived by a hierarchical clustering algorithm, using complete linkage and a simple matching coefficient on the binary variables of presence or absence of symptoms. The four groups of patients suggested by the cluster analysis can be characterized as follows: (1) generalized multiplicity of symptoms; (2) capacity to cope except for orientation apart from generally held norms; (3) activity level and thought processes speeded up, intensified, and unselected; (4) inwardly punitive, slowed down and distressed. It is shown that these groups received significantly different treatment and that the effect of treatment was significantly different, while no such differences were noted for groups defined in terms of diagnoses. By means of linear discriminant functions, rules are suggested for assigning other psychiatric patients to one of these four groups.

Antipsychotic Agents↗

Critical behavior of the long-range Ising chain from the largest-cluster probability distribution.

Monte Carlo simulations of the one-dimensional Ising model with ferromagnetic interactions decaying with distance r as 1/r(1+sigma) are performed by applying the Swendsen-Wang cluster algorithm with cumulative probabilities. The critical behavior in the nonclassical critical regime corresponding to 0.5<sigma<1 is derived from finite-size scaling analysis of the largest cluster.

Journal Article↗

Significance and statistical errors in the analysis of DNA microarray data.

DNA microarrays are important devices for high throughput measurements of gene expression, but no rational foundation has been established for understanding the sources of within-chip statistical error. We designed a specialized chip and protocol to investigate the distribution and magnitude of within-chip errors and discovered that, as expected from theoretical expectations, measurement errors follow a Lorentzian-like distribution, which explains the widely observed but unexplained ill-reproducibility in microarray data. Using this specially designed chip, we examined a data set of repeated measurements to extract estimates of the distribution and magnitude of statistical errors in DNA microarray measurements. Using the common "ratio of medians" method, we find that the measurements follow a Lorentzian-like distribution, which is problematic for subsequent analysis. We show that a method of analysis dubbed "median of ratios" yields a more Gaussian-like distribution of errors. Finally, we show that the bootstrap algorithm can be used to extract the best estimates of the error in the measurement. Quantifying the statistical error in such measurements has important applications for estimating significance levels, clustering algorithms, and process optimization.

Algorithms↗

An analysis of auditory alphabet confusions.

The present study, using the nonhierarchical overlapping clustering algorithm MAPCLUS to fit the Shepard-Arabie (1979) ADCLUS model, attempted to derive a set of features that would accurately describe the auditory alphabet confusions present in the data matrices of Conrad (1964) and Hull (1973). Separate nine-cluster solutions accounted for 80% and 89% of the variance in the matrices, respectively. The clusters revealed that the most frequently confused letter names contained common vowels and phonetically similar consonants. Further analyses using INDCLUS, an individual differences extension of the MAPCLUS algorithm and ADCLUS model, indicated that while the patterns of errors in the two matrices were remarkably similar, some differences were also apparent. These differences reflected the differing amounts of background noise present in the two studies.

Adult↗

Determination of protein tertiary structure class from circular dichroism spectra.

Fifty-three circular dichroism (CD) spectra consisting of the spectra of 46 native proteins, 3 denatured proteins, and one oligopeptide (the spectra of two denatured proteins and oligopeptide were taken at two different temperatures) were investigated in order to examine the correlation between the shape of the CD spectrum and the tertiary structure class of the protein. Five classes were considered--all -alpha, all -beta, alpha+beta, alpha/beta, and denatured proteins. Spectra from 190 to 236 nm with 2 nm interval were described as points in 24-dimensional hyperspace, where coordinates were values of ellipticities at fixed wavelengths. This allows the spectra to be treated as patterns and subsequently analyzed using pattern recognition algorithms. Cluster analysis, which does not need predefined information about protein structure, divides spectra into several compact groups or clusters with good correlation with tertiary structure class. To visualize these results, orthogonalization procedures were imposed on the original data set in 24-dimensional space. The new 3-dimensional coordinate system demonstrated well-separated all-beta class and denatured proteins. Regions corresponding to all -alpha and especially alpha+beta and alpha/beta proteins were not as well resolved. The following approach was then applied to the original data set to obtain an objective mathematical algorithm for the determination of a protein's tertiary structure class from its CD spectrum. Regions in 24-dimensional hyperspace corresponding to all of the tertiary structure classes were found by calculating the decision functions, or equations of hyperplanes, which separate groups of spectral patterns of different classes.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

An artificial intelligent algorithm for tumor detection in screening mammogram.

Cancerous tumor mass is one of the major types of breast cancer. When cancerous masses are embedded in and camouflaged by varying densities of parenchymal tissue structures, they are very difficult to be visually detected on mammograms. This paper presents an algorithm that combines several artificial intelligent techniques with the discrete wavelet transform (DWT) for detection of masses in mammograms. The AI techniques include fractal dimension analysis, multiresolution markov random field, dogs-and-rabbits algorithm, and others. The fractal dimension analysis serves as a preprocessor to determine the approximate locations of the regions suspicious for cancer in the mammogram. The dogs-and-rabbits clustering algorithm is used to initiate the segmentation at the LL subband of a three-level DWT decomposition of the mammogram. A tree-type classification strategy is applied at the end to determine whether a given region is suspicious for cancer. We have verified the algorithm with 322 mammograms in the Mammographic Image Analysis Society Database. The verification results show that the proposed algorithm has a sensitivity of 97.3% and the number of false positive per image is 3.92.

Algorithms↗

Analysis of methotrexate treatment effect in a longitudinal observational study: utility of cluster analysis.

We studied 235 patients with rheumatoid arthritis (RA) beginning therapy with methotrexate utilizing a k-means clustering algorithm. Four groups were identified: mild RA (Group 3), very severe RA (Group 4), and 2 groups intermediate in severity (Groups 1 and 2). Group 2, the largest of the clusters (n = 89), appeared to have greater tolerability of RA as measured by severity and psychological variables, and took the drug almost twice as long as other groups, although improvement was not greater nor side effects fewer. All groups improved over a mean of 1.9 years, and the degree of improvement was not related to the initial severity classification. Improvement occurred almost equally in all clusters, and the relative ranking of the groups was maintained at study closure.

Arthritis, Rheumatoid↗

Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays.

Oligonucleotide arrays can provide a broad picture of the state of the cell, by monitoring the expression level of thousands of genes at the same time. It is of interest to develop techniques for extracting useful information from the resulting data sets. Here we report the application of a two-way clustering method for analyzing a data set consisting of the expression patterns of different cell types. Gene expression in 40 tumor and 22 normal colon tissue samples was analyzed with an Affymetrix oligonucleotide array complementary to more than 6,500 human genes. An efficient two-way clustering algorithm was applied to both the genes and the tissues, revealing broad coherent patterns that suggest a high degree of organization underlying gene expression in these tissues. Coregulated families of genes clustered together, as demonstrated for the ribosomal proteins. Clustering also separated cancerous from noncancerous tissue and cell lines from in vivo tissues on the basis of subtle distributed patterns of genes even when expression of individual genes varied only slightly between the tissues. Two-way clustering thus may be of use both in classifying genes into functional groups and in classifying tissues based on gene expression.

Adenocarcinoma↗

The 5S ribosomal RNA sequences of a red algal rhodoplast and a gymnosperm chloroplast. Implications for the evolution of plastids and cyanobacteria.

The 5S ribosomal RNA sequences have been determined for the rhodoplast of the red alga Porphyra umbilicalis and the chloroplast of the conifer Juniperus media. The 5S RNA sequence of the Vicia faba chloroplast is corrected with respect to a previous report. A survey of the known sequences and secondary structures of 5S RNAs from plastids and cyanobacteria shows a close structural similarity between all 5S RNAs from land plant chloroplasts. The algal plastid 5S RNAs on the other hand show much more structural diversity and have certain structural features in common with bacterial 5S RNAs. A dendrogram constructed from the aligned sequences by a clustering algorithm points to a common ancestor for the present-living cyanobacteria and the land plant plastids. However, the algal plastids branch off at an early stage within the plastid-cyanobacteria cluster, before the divergence between cyanobacteria and land plant chloroplasts. This evolutionary picture points to the occurrence of multiple endosymbiotic events, with the ancestors of the present algal plastids already established as photosynthetic endosymbionts at a time when the ancestors of the present land plant chloroplasts were still free-living cells.

Base Sequence↗

Gene expression profiles of cutaneous B cell lymphoma.

We studied gene expression profiles of 17 cutaneous B cell lymphomas that were collected with 4-6 mm skin punch biopsies. We also included tissue from two cases of mycosis fungoides, three normal skin biopsies, and three tonsils to create a framework for further interpretation. A hierarchical cluster algorithm was applied for data analysis. Our results indicate that small amounts of skin tissue can be used successfully to perform microarray analysis and result in distinct gene expression patterns. Duplicate specimens clustered together demonstrating a reproducible technique. Within the cutaneous B cell lymphoma specimens two specific B cell differentiation stage signatures of germinal center B cells and plasma cells could be identified. Primary cutaneous follicular and primary cutaneous diffuse large B cell lymphomas had a germinal center B cell signature, whereas a subset of marginal zone lymphomas demonstrated a plasma cell signature. Primary and secondary follicular B cell lymphoma of the skin were closely related, despite previously reported genetic and phenotypic differences. In contrast primary and secondary cutaneous diffuse large B cell lymphoma were less related to each other. This pilot study allows a first glance into the complex and unique microenvironment of B cell lymphomas of the skin and provides a basis for future studies, which may lead to the identification of potential histologic and prognostic markers as well as therapeutic targets.

Adult↗

Joint entropy maximization in kernel-based topographic maps.

A new learning algorithm for kernel-based topographic map formation is introduced. The kernel parameters are adjusted individually so as to maximize the joint entropy of the kernel outputs. This is done by maximizing the differential entropies of the individual kernel outputs, given that the map's output redundancy, due to the kernel overlap, needs to be minimized. The latter is achieved by minimizing the mutual information between the kernel outputs. As a kernel, the (radial) incomplete gamma distribution is taken since, for a gaussian input density, the differential entropy of the kernel output will be maximal. Since the theoretically optimal joint entropy performance can be derived for the case of nonoverlapping gaussian mixture densities, a new clustering algorithm is suggested that uses this optimum as its "null" distribution. Finally, it is shown that the learning algorithm is similar to one that performs stochastic gradient descent on the Kullback-Leibler divergence for a heteroskedastic gaussian mixture density model.

Algorithms↗

The phylogenetic diversity of eukaryotic transcription.

Eukaryotic transcription is a highly regulated process involving interactions between large numbers of proteins. To analyse the phylogenetic distribution of the components of this process, six crown eukaryote group genomes were queried with a reference set of transcription-associated (TA) proteins. On average, one in 10 proteins encoded by these genomes were found to be homologous to sequences in the reference set. Analysis of families identified using an accurate sequence clustering algorithm and containing both TA proteins and eukaryotic sequences showed that in two-thirds of the families the homologues originate from a single kingdom. Furthermore, in only 15% of the fungal-specific clusters are the homologues present in both budding and fission yeast, as compared with the metazoan-specific clusters where 53% of the homologues originate from two or more species. Families whose members comprise general transcription factor or RNA polymerase subunits exhibit a low degree of taxon specificity, suggesting that the transcription initiation complex is highly conserved. This contrasts with transcriptional regulator families, that are primarily taxon-specific, indicating proteins controlling gene activation exhibit considerable sequence diversity across the eukaryotic domain.

Animals↗

5S ribosomal ribonucleic acid sequences in Bacteroides and Fusobacterium: evolutionary relationships within these genera and among eubacteria in general.

The 5S ribosomal ribonucleic acid (rRNA) sequences were determined for Bacteroides fragilis, Bacteroides thetaiotaomicron, Bacteroides capillosus, Bacteroides veroralis, Porphyromonas gingivalis, Anaerorhabdus furcosus, Fusobacterium nucleatum, Fusobacterium mortiferum, and Fusobacterium varium. A dendrogram constructed by a clustering algorithm from these sequences, which were aligned with all other hitherto known eubacterial 5S rRNA sequences, showed differences as well as similarities with respect to results derived from 16S rRNA analyses. In the 5S rRNA dendrogram, Bacteroides clustered together with Cytophaga and Fusobacterium, as in 16S rRNA analyses. Intraphylum relationships deduced from 5S rRNAs suggested that Bacteroides is specifically related to Cytophaga rather than to Fusobacterium, as was suggested by 16S rRNA analyses. Previous taxonomic considerations concerning the genus Bacteroides, based on biochemical and physiological data, were confirmed by the 5S rRNA sequence analysis.

Algorithms↗

Metric and multidimensional scaling: efficient tools for clustering molecular conformations.

The application of metric and multidimensional scaling to conformer ensembles was demonstrated in this work. An automated process was devised to cluster and assign group memberships and cluster representatives. The method allows rapid clustering, leading to intuitive results that can be visually inspected. Multidimensional scaling was found to be superior to metric scaling for clustering conformers. The performance of different hierarchical clustering algorithms was compared using multidimensional plots, and the group average method was found to perform best.

Amino Acid Sequence↗

Docking of flexible molecules using multiscale ligand representations.

Structural genomics will yield an immense number of protein three-dimensional structures in the near future. Automated theoretical methodologies are needed to exploit this information and are likely to play a pivotal role in drug discovery. Here, we present a fully automated, efficient docking methodology that does not require any a priori knowledge about the location of the binding site or function of the protein. The method relies on a multiscale concept where we deal with a hierarchy of models generated for the potential ligand. The models are created using the k-means clustering algorithm. The method was tested on seven protein-ligand complexes. In the largest complex, human immunodeficiency virus reverse transcriptase/nevirapin, the root mean square deviation value when comparing our results to the crystal structure was 0.29 A. We demonstrate on an additional 25 protein-ligand complexes that the methodology may be applicable to high throughput docking. This work reveals three striking results. First, a ligand can be docked using a very small number of feature points. Second, when using a multiscale concept, the number of conformers that require to be generated can be significantly reduced. Third, fully flexible ligands can be treated as a small set of rigid k-means clusters.

Algorithms↗