PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Genetic diversity analysis in valencia peanut (Arachis hypogaea L.) using microsatellite markers.

Cultivated peanut or groundnut (Arachis hypogaea L) is an important source of oil and protein. Considerable variation has been recorded for morphological, physiological and agronomic traits, whereas few molecular variations have been recorded for this crop. The identification and understanding of molecular genetic diversity in cultivated peanut types will help in effective genetic conservation along with efficient breeding programs in this crop. The New Mexico breeding program has embarked upon a program of improvement of Valencia peanut (belonging to the sub species fastigiata), because efforts to improve the yield potential are lacking due to lack of identified divergent exotic types. For the first time, this study has shown molecular diversity using microsatellite markers in the cultivated Valencia peanut (sub spp. fastigiata) from around the globe. In this investigation, 48 cultivated Valencia peanut genotypes have been selected and analyzed using 18 fluorescently labeled SSR (f-SSR) primer pairs. These primer pairs amplified 120 polymorphic loci among the genotypes screened and amplified from 3 to 19 alleles with an average of 6.9 allele per primer pair. The f-SSR marker data was further analyzed using cluster algorithms and principal component analysis. The results indicated that (1) considerable genetic variations were discovered among the analyzed genotypes; (2) The f-SSR based clustering could identify the putative pedigree types of the present Valencia types of diverse origins, and (3) The f-SSR in general is sufficient to obtain estimates of genetic divergence for the material in study. The results are being utilized in our breeding program for parental selection and linkage map construction.

Arachis↗

Statistical methods for characterizing diversity of microbial communities by analysis of terminal restriction fragment length polymorphisms of 16S rRNA genes.

The analysis of terminal restriction fragment length polymorphisms (T-RFLP) of 16S rRNA genes has proven to be a facile means to compare microbial communities and presumptively identify abundant members. The method provides data that can be used to compare different communities based on similarity or distance measures. Once communities have been clustered into groups, clone libraries can be prepared from sample(s) that are representative of each group in order to determine the phylogeny of the numerically abundant populations in a community. In this paper methods are introduced for the statistical analysis of T-RFLP data that include objective methods for (i) determining a baseline so that 'true' peaks in electropherograms can be identified; (ii) a means to compare electropherograms and bin fragments of similar size; (iii) clustering algorithms that can be used to identify communities that are similar to one another; and (iv) a means to select samples that are representative of a cluster that can be used to construct 16S rRNA gene clone libraries. The methods for data analysis were tested using simulated data with assumptions and parameters that corresponded to actual data. The simulation results demonstrated the usefulness of these methods in their ability to recover the true microbial community structure generated under the assumptions made. Software for implementing these methods is available at http://www.ibest.uidaho.edu/tools/trflp_stats/index.php.

Bacteria↗

Cluster analysis applied to symptom ratings of psychiatric patients: an evaluation of its predictive ability.

Rating on 39 symptoms were examined for patients admitted to the Neuropsychiatric Institute of the University of Michigan Medical Center. A detailed evaluation was made of the clusters derived by a hierarchical clustering algorithm, using complete linkage and a simple matching coefficient on the binary variables of presence or absence of symptoms. The four groups of patients suggested by the cluster analysis can be characterized as follows: (1) generalized multiplicity of symptoms; (2) capacity to cope except for orientation apart from generally held norms; (3) activity level and thought processes speeded up, intensified, and unselected; (4) inwardly punitive, slowed down and distressed. It is shown that these groups received significantly different treatment and that the effect of treatment was significantly different, while no such differences were noted for groups defined in terms of diagnoses. By means of linear discriminant functions, rules are suggested for assigning other psychiatric patients to one of these four groups.

Antipsychotic Agents↗

Infrared target detection with probability density functions of wavelet transform subbands.

We report the development of a wavelet multiresolution texture-based algorithm that uses the probability density functions (PDFs) of the subband of the wavelet decomposition of an image. The moments of these pdfs are used in a clustering algorithm to segment the targets from their background clutter. Using the tools of experimental methodology, we evaluate the performance of this algorithm on real infrared imagery under varying algorithm parameter sets as well as scene, image, and false-alarm conditions. We estimate a set of multidimensional predictive analytic performance models that relate the detection probabilities as functions of false alarm, algorithm internal parameter, target pixel number, target-to-background interference ratio, target-interference ratio, and Fechner-Weber and local entropy metrics in the scene. These models can be used to predict performance in regions were no data are available and to optimize performance by selection of the optimum parameter and constant false-alarm values in regions with known scene and metric conditions.

Journal Article↗

Metagenes and molecular pattern discovery using matrix factorization.

We describe here the use of nonnegative matrix factorization (NMF), an algorithm based on decomposition by parts that can reduce the dimension of expression data from thousands of genes to a handful of metagenes. Coupled with a model selection mechanism, adapted to work for any stochastic clustering algorithm, NMF is an efficient method for identification of distinct molecular patterns and provides a powerful method for class discovery. We demonstrate the ability of NMF to recover meaningful biological information from cancer-related microarray data. NMF appears to have advantages over other methods such as hierarchical clustering or self-organizing maps. We found it less sensitive to a priori selection of genes or initial conditions and able to detect alternative or context-dependent patterns of gene expression in complex biological systems. This ability, similar to semantic polysemy in text, provides a general method for robust molecular pattern discovery.

Algorithms↗

Critical behavior of the long-range Ising chain from the largest-cluster probability distribution.

Monte Carlo simulations of the one-dimensional Ising model with ferromagnetic interactions decaying with distance r as 1/r(1+sigma) are performed by applying the Swendsen-Wang cluster algorithm with cumulative probabilities. The critical behavior in the nonclassical critical regime corresponding to 0.5<sigma<1 is derived from finite-size scaling analysis of the largest cluster.

Journal Article↗

Significance and statistical errors in the analysis of DNA microarray data.

DNA microarrays are important devices for high throughput measurements of gene expression, but no rational foundation has been established for understanding the sources of within-chip statistical error. We designed a specialized chip and protocol to investigate the distribution and magnitude of within-chip errors and discovered that, as expected from theoretical expectations, measurement errors follow a Lorentzian-like distribution, which explains the widely observed but unexplained ill-reproducibility in microarray data. Using this specially designed chip, we examined a data set of repeated measurements to extract estimates of the distribution and magnitude of statistical errors in DNA microarray measurements. Using the common "ratio of medians" method, we find that the measurements follow a Lorentzian-like distribution, which is problematic for subsequent analysis. We show that a method of analysis dubbed "median of ratios" yields a more Gaussian-like distribution of errors. Finally, we show that the bootstrap algorithm can be used to extract the best estimates of the error in the measurement. Quantifying the statistical error in such measurements has important applications for estimating significance levels, clustering algorithms, and process optimization.

Algorithms↗

An analysis of auditory alphabet confusions.

The present study, using the nonhierarchical overlapping clustering algorithm MAPCLUS to fit the Shepard-Arabie (1979) ADCLUS model, attempted to derive a set of features that would accurately describe the auditory alphabet confusions present in the data matrices of Conrad (1964) and Hull (1973). Separate nine-cluster solutions accounted for 80% and 89% of the variance in the matrices, respectively. The clusters revealed that the most frequently confused letter names contained common vowels and phonetically similar consonants. Further analyses using INDCLUS, an individual differences extension of the MAPCLUS algorithm and ADCLUS model, indicated that while the patterns of errors in the two matrices were remarkably similar, some differences were also apparent. These differences reflected the differing amounts of background noise present in the two studies.

Adult↗

Determination of protein tertiary structure class from circular dichroism spectra.

Fifty-three circular dichroism (CD) spectra consisting of the spectra of 46 native proteins, 3 denatured proteins, and one oligopeptide (the spectra of two denatured proteins and oligopeptide were taken at two different temperatures) were investigated in order to examine the correlation between the shape of the CD spectrum and the tertiary structure class of the protein. Five classes were considered--all -alpha, all -beta, alpha+beta, alpha/beta, and denatured proteins. Spectra from 190 to 236 nm with 2 nm interval were described as points in 24-dimensional hyperspace, where coordinates were values of ellipticities at fixed wavelengths. This allows the spectra to be treated as patterns and subsequently analyzed using pattern recognition algorithms. Cluster analysis, which does not need predefined information about protein structure, divides spectra into several compact groups or clusters with good correlation with tertiary structure class. To visualize these results, orthogonalization procedures were imposed on the original data set in 24-dimensional space. The new 3-dimensional coordinate system demonstrated well-separated all-beta class and denatured proteins. Regions corresponding to all -alpha and especially alpha+beta and alpha/beta proteins were not as well resolved. The following approach was then applied to the original data set to obtain an objective mathematical algorithm for the determination of a protein's tertiary structure class from its CD spectrum. Regions in 24-dimensional hyperspace corresponding to all of the tertiary structure classes were found by calculating the decision functions, or equations of hyperplanes, which separate groups of spectral patterns of different classes.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

A test for the consecutive ones property on noisy data--application to physical mapping and sequence assembly.

A (0,1)-matrix satisfies the consecutive ones property (COP) for the rows if there exists a column permutation such that the ones in each row of the resultant matrix are consecutive. The consecutive ones test is useful for physical mapping and DNA sequence assembly, for example, in the STS content mapping of YAC library, and in the Bactig assembly based on STS as well as EST markers. The linear time algorithm by Booth and Lueker (1976) for this problem has a serious drawback: the data must be error free. However, laboratory work is never flawless. We devised a new iterative clustering algorithm for this problem, which has the following advantages: 1. If the original matrix satisfies the COP, then the algorithm will produce a column ordering realizing it without any fill-in. 2. Under moderate assumptions, the algorithm can accommodate the following four types of errors: false negatives, false positives, nonunique probes, and chimeric clones. Note that in some cases (low quality EST marker identification), NPs occur because of repeat sequences. 3. In case some local data is too noisy, our algorithm could likely discover that and suggest additional lab work to reduce the degree of ambiguity in that part. 4. A unique feature of our algorithm is that, rather than forcing all probes to be included and ordered in the final arrangement, our algorithm would delete some noisy probes. Thus, it could produce more than one contig. The gaps are created mostly by noisy probes.

Algorithms↗

An artificial intelligent algorithm for tumor detection in screening mammogram.

Cancerous tumor mass is one of the major types of breast cancer. When cancerous masses are embedded in and camouflaged by varying densities of parenchymal tissue structures, they are very difficult to be visually detected on mammograms. This paper presents an algorithm that combines several artificial intelligent techniques with the discrete wavelet transform (DWT) for detection of masses in mammograms. The AI techniques include fractal dimension analysis, multiresolution markov random field, dogs-and-rabbits algorithm, and others. The fractal dimension analysis serves as a preprocessor to determine the approximate locations of the regions suspicious for cancer in the mammogram. The dogs-and-rabbits clustering algorithm is used to initiate the segmentation at the LL subband of a three-level DWT decomposition of the mammogram. A tree-type classification strategy is applied at the end to determine whether a given region is suspicious for cancer. We have verified the algorithm with 322 mammograms in the Mammographic Image Analysis Society Database. The verification results show that the proposed algorithm has a sensitivity of 97.3% and the number of false positive per image is 3.92.

Algorithms↗

Screening anti-inflammatory compounds in injured spinal cord with microarrays: a comparison of bioinformatics analysis approaches.

Inflammatory responses contribute to secondary tissue damage following spinal cord injury (SCI). A potent anti-inflammatory glucocorticoid, methylprednisolone (MP), is the only currently accepted therapy for acute SCI but its efficacy has been questioned. To search for additional anti-inflammatory compounds, we combined microarray analysis with an explanted spinal cord slice culture injury model. We compared gene expression profiles after treatment with MP, acetaminophen, indomethacin, NS398, and combined cytokine inhibitors (IL-1ra and soluble TNFR). Multiple gene filtering methods and statistical clustering analyses were applied to the multi-dimensional data set and results were compared. Our analysis showed a consistent and unique gene expression profile associated with NS398, the selective cyclooxygenase-2 (COX-2) inhibitor, in which the overall effect of these upregulated genes could be interpreted as neuroprotective. In vivo testing demonstrated that NS398 reduced lesion volumes, unlike MP or acetaminophen, consistent with a predicted physiological effect in spinal cord. Combining explanted spinal cultures, microarrays, and flexible clustering algorithms allows us to accelerate selection of compounds for in vivo testing.

Algorithms↗

Analysis of methotrexate treatment effect in a longitudinal observational study: utility of cluster analysis.

We studied 235 patients with rheumatoid arthritis (RA) beginning therapy with methotrexate utilizing a k-means clustering algorithm. Four groups were identified: mild RA (Group 3), very severe RA (Group 4), and 2 groups intermediate in severity (Groups 1 and 2). Group 2, the largest of the clusters (n = 89), appeared to have greater tolerability of RA as measured by severity and psychological variables, and took the drug almost twice as long as other groups, although improvement was not greater nor side effects fewer. All groups improved over a mean of 1.9 years, and the degree of improvement was not related to the initial severity classification. Improvement occurred almost equally in all clusters, and the relative ranking of the groups was maintained at study closure.

Arthritis, Rheumatoid↗

Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays.

Oligonucleotide arrays can provide a broad picture of the state of the cell, by monitoring the expression level of thousands of genes at the same time. It is of interest to develop techniques for extracting useful information from the resulting data sets. Here we report the application of a two-way clustering method for analyzing a data set consisting of the expression patterns of different cell types. Gene expression in 40 tumor and 22 normal colon tissue samples was analyzed with an Affymetrix oligonucleotide array complementary to more than 6,500 human genes. An efficient two-way clustering algorithm was applied to both the genes and the tissues, revealing broad coherent patterns that suggest a high degree of organization underlying gene expression in these tissues. Coregulated families of genes clustered together, as demonstrated for the ribosomal proteins. Clustering also separated cancerous from noncancerous tissue and cell lines from in vivo tissues on the basis of subtle distributed patterns of genes even when expression of individual genes varied only slightly between the tissues. Two-way clustering thus may be of use both in classifying genes into functional groups and in classifying tissues based on gene expression.

Adenocarcinoma↗

Identification of hair cycle-associated genes from time-course gene expression profile data by using replicate variance.

The hair-growth cycle is an example of a cyclic process that is well characterized morphologically but understood incompletely at the molecular level. As an initial step in discovering regulators in hair-follicle morphogenesis and cycling, we used DNA microarrays to profile mRNA expression in mouse back skin from eight representative time points. We developed a statistical algorithm to identify the set of genes expressed within skin that are associated specifically with the hair-growth cycle. The methodology takes advantage of higher replicate variance during asynchronous hair cycles in comparison with synchronous cycles. More than one-third of genes with detectable skin expression showed hair-cycle-related changes in expression, suggesting that many more genes may be associated with the hair-growth cycle than have been identified in the literature. By using a probabilistic clustering algorithm for replicated measurements, these genes were grouped into 30 time-course profile clusters, which fall into four major classes. Distinct genetic pathways were characteristic for the different time-course profile clusters, providing insights into the regulation of hair-follicle cycling and suggesting that this approach is useful for identifying hair follicle regulators. In addition to revealing known hair-related genes, we identified genes that were not previously known to be hair cycle-associated and confirmed their temporal and spatial expression patterns during the hair-growth cycle by quantitative real-time PCR and in situ hybridization. The same computational approach should be generally useful for identifying genes associated with cyclic processes from complex tissues.

Algorithms↗

The 5S ribosomal RNA sequences of a red algal rhodoplast and a gymnosperm chloroplast. Implications for the evolution of plastids and cyanobacteria.

The 5S ribosomal RNA sequences have been determined for the rhodoplast of the red alga Porphyra umbilicalis and the chloroplast of the conifer Juniperus media. The 5S RNA sequence of the Vicia faba chloroplast is corrected with respect to a previous report. A survey of the known sequences and secondary structures of 5S RNAs from plastids and cyanobacteria shows a close structural similarity between all 5S RNAs from land plant chloroplasts. The algal plastid 5S RNAs on the other hand show much more structural diversity and have certain structural features in common with bacterial 5S RNAs. A dendrogram constructed from the aligned sequences by a clustering algorithm points to a common ancestor for the present-living cyanobacteria and the land plant plastids. However, the algal plastids branch off at an early stage within the plastid-cyanobacteria cluster, before the divergence between cyanobacteria and land plant chloroplasts. This evolutionary picture points to the occurrence of multiple endosymbiotic events, with the ancestors of the present algal plastids already established as photosynthetic endosymbionts at a time when the ancestors of the present land plant chloroplasts were still free-living cells.

Base Sequence↗

Gene expression profiles of cutaneous B cell lymphoma.

We studied gene expression profiles of 17 cutaneous B cell lymphomas that were collected with 4-6 mm skin punch biopsies. We also included tissue from two cases of mycosis fungoides, three normal skin biopsies, and three tonsils to create a framework for further interpretation. A hierarchical cluster algorithm was applied for data analysis. Our results indicate that small amounts of skin tissue can be used successfully to perform microarray analysis and result in distinct gene expression patterns. Duplicate specimens clustered together demonstrating a reproducible technique. Within the cutaneous B cell lymphoma specimens two specific B cell differentiation stage signatures of germinal center B cells and plasma cells could be identified. Primary cutaneous follicular and primary cutaneous diffuse large B cell lymphomas had a germinal center B cell signature, whereas a subset of marginal zone lymphomas demonstrated a plasma cell signature. Primary and secondary follicular B cell lymphoma of the skin were closely related, despite previously reported genetic and phenotypic differences. In contrast primary and secondary cutaneous diffuse large B cell lymphoma were less related to each other. This pilot study allows a first glance into the complex and unique microenvironment of B cell lymphomas of the skin and provides a basis for future studies, which may lead to the identification of potential histologic and prognostic markers as well as therapeutic targets.

Adult↗

Application of SNPs for assessing biodiversity and phylogeny among yeast strains.

We examined the efficacy of single-nucleotide polymorphism (SNP) markers for the assessment of the phylogeny and biodiversity of Saccharomyces strains. Each of 32 Saccharomyces cerevisiae strains was genotyped at 30 SNP loci discovered by sequence alignment of the S. cerevisiae laboratory strain SK1 to the database sequence of strain S288c. In total, 10 SNPs were selected from each of the following three categories: promoter regions, nonsynonymous and synonymous sites (in open reading frames). The strains in this study included 11 haploid laboratory strains used for genetic studies and 21 diploids. Three non-cerevisiae species of Saccharomyces (sensu stricto) were used as an out-group. A Bayesian clustering-algorithm, Structure, effectively identified four different strain groups: laboratory, wine, other diploids and the non-cerevisiae species. Analysing haploid and diploid strains together caused problems for phylogeny reconstruction, but not for the clustering produced by Structure. The ascertainment bias introduced by the SNP discovery method caused difficulty in the phylogenetic analysis; alternative options are proposed. A smaller data set, comprising only the nine most polymorphic loci, was sufficient to obtain most features of the results.

Biodiversity↗