PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Sectorization of the central 30 degrees visual field in glaucoma.

PURPOSE: To determine an optimal sector pattern of the central 30 degrees visual field in glaucoma by mathematically analyzing the visual field data of primary open-angle glaucoma (POAG) without any assumption such as the retinal nerve fiber layer anatomy. METHODS: One hundred three visual fields of the 30-2 program of the Humphrey Field Analyzer obtained from 103 POAG patients of early to moderately advanced stage were included. Based on the interpoint correlation of deviation of the measured threshold value from the age-corrected normal reference value (total deviation, STATPAC), test points of the 30-2 program were mathematically clustered using the VARCLUS procedure, a new clustering algorithm developed by the SAS Institute. The sector value, which summarizes the visual field performance of the clustered test points (sector), also was calculated. RESULTS: The 30 degrees central visual field was divided into 15 sectors consisting of at least 3 points. The distribution of sectors was compatible with the projection of nerve fiber layers. There was no sector extending over the horizontal meridian, but the sector pattern was not completely symmetrical around it. Linear regression analysis of the sector values against the mean deviation (STATPAC) suggested that the index is useful in following visual field performance of each sector. CONCLUSION: The sector pattern and sector values obtained were considered useful in studying the visual field data of glaucoma.

Algorithms↗

Data mining the NCI cancer cell line compound GI(50) values: identifying quinone subtypes effective against melanoma and leukemia cell classes.

Using data mining techniques, we have studied a subset (1400) of compounds from the large public National Cancer Institute (NCI) compounds data repository. We first carried out a functional class identity assignment for the 60 NCI cancer testing cell lines via hierarchical clustering of gene expression data. Comprised of nine clinical tissue types, the 60 cell lines were placed into six classes-melanoma, leukemia, renal, lung, and colorectal, and the sixth class was comprised of mixed tissue cell lines not found in any of the other five classes. We then carried out supervised machine learning, using the GI(50) values tested on a panel of 60 NCI cancer cell lines. For separate 3-class and 2-class problem clustering, we successfully carried out clear cell line class separation at high stringency, p < 0.01 (Bonferroni corrected t-statistic), using feature reduction clustering algorithms embedded in RadViz, an integrated high dimensional analytic and visualization tool. We started with the 1400 compound GI(50) values as input and selected only those compounds, or features, significant in carrying out the classification. With this approach, we identified two small sets of compounds that were most effective in carrying out complete class separation of the melanoma, non-melanoma classes and leukemia, non-leukemia classes. To validate these results, we showed that these two compound sets' GI(50) values were highly accurate classifiers using five standard analytical algorithms. One compound set was most effective against the melanoma class cell lines (14 compounds), and the other set was most effective against the leukemia class cell lines (30 compounds). The two compound classes were both significantly enriched in two different types of substituted p-quinones. The melanoma cell line class of 14 compounds was comprised of 11 compounds that were internal substituted p-quinones, and the leukemia cell line class of 30 compounds was comprised of 6 compounds that were external substituted p-quinones. Attempts to subclassify melanoma or leukemia cell lines based upon their clinical cancer subtype met with limited success. For example, using GI(50) values for the 30 compounds we identified as effective against all leukemia cell lines, we could subclassify acute lymphoblastic leukemia (ALL) origin cell lines from non-ALL leukemia origin cell lines without significant overlap from non-leukemia cell lines. Based upon clustering using GI(50) values for the 60 cancer cell lines laid out by the RadViz algorithm, these two compound subsets did not overlap with clusters containing any of the NCI's 92 compounds of known mechanism of action, a few of which are quinones. Given their structural patterns, the two p-quinone subtypes we identified would clearly be expected to possess different redox potentials/substrate specificities for enzymatic reduction in vivo. These two p-quinone subtypes represent valuable information that may be used in the elucidation of pharmacophores for the design of compounds to treat these two cancer tissue types in the clinic.

Algorithms↗

Computer image analysis of two-dimensional crystals of beef heart NADH: ubiquinone oxidoreductase fragments. I. Comparison of crystal structures in various negative stains.

We investigated the structure of two-dimensional crystals from bovine heart mitochondrial NADH: ubiquinone oxidoreductase. A detailed description of uranyl acetate-stained crystals demonstrated that they are composed of fragments in a spatial arrangement according to space group P4212 [J. Brink, S. Hovmöller, C.I. Ragan, M.W.J. Cleeter, E.J. Boekema and E.F.J. van Bruggen, European J. Biochem. 166 (1987) 287]. To gain more structural information on the crystal structure and to assess the effects of various negative stains on the structure preservation and appearance, we examined stained crystals by means of electron microscopy and image analysis. The space group P4212 appeared to be present for several stains tested, i.e. ammonium molybdate, uranyl acetate, uranyl nitrate and uranyl sulphate. Use of phosphotungstic acid and silicotungstate resulted in a reduction of symmetry to pseudo-P4212 or p4. Use of sodium tungstate led to a considerable loss of resolution to 3.8 nm at best, whereas otherwise 1.5 to 1.9 nm could be demonstrated. The lattice vectors were not affected by the stains; they were determined as a = b = 14.9 +/- 0.25 nm with gamma = 89.8 degrees +/- 0.6 degrees. Image analysis showed the presence of similar structures with the molybdate and uranyl compounds. Differences were observed in the case of the tungstate type of stains. Furthermore, the analysis revealed the complete absence of the four small pores of 2.0 nm diameter in the unit cell. This effect was observed irrespective of the type of stain and supporting film, and could be ascribed only to the glow-discharge treatment of the supporting film. The observed difference must be caused by changed interactions between the protein, stain and supporting film. Application of correspondence analysis and clustering algorithms to the various reconstructed images of the crystals showed that they could be separated into several clusters. Each of these clusters corresponded on the average to only one type of stain, whereas a further division according to the specific uranyl compounds was observed. This study therefore shows that under identical preparation conditions subtle differences between individual stains can be detected.

Animals↗

Continuous representations of time-series gene expression data.

We present algorithms for time-series gene expression analysis that permit the principled estimation of unobserved time points, clustering, and dataset alignment. Each expression profile is modeled as a cubic spline (piecewise polynomial) that is estimated from the observed data and every time point influences the overall smooth expression curve. We constrain the spline coefficients of genes in the same class to have similar expression patterns, while also allowing for gene specific parameters. We show that unobserved time points can be reconstructed using our method with 10-15% less error when compared to previous best methods. Our clustering algorithm operates directly on the continuous representations of gene expression profiles, and we demonstrate that this is particularly effective when applied to nonuniformly sampled data. Our continuous alignment algorithm also avoids difficulties encountered by discrete approaches. In particular, our method allows for control of the number of degrees of freedom of the warp through the specification of parameterized functions, which helps to avoid overfitting. We demonstrate that our algorithm produces stable low-error alignments on real expression data and further show a specific application to yeast knock-out data that produces biologically meaningful results.

Algorithms↗

A funny thing happened to us on the way to the latent entities.

Inferred latent entities, whether those of psychoanalysis, factor analysis, or cluster analysis, have declined in value for many clinical psychologists, both as tools of practice and as objects of theoretical interest. Behavior modification, rational-emotive therapy, crisis intervention, psycho-pharmacology, and actuarial prediction all tend to minimize reliance on latent entities in favor of purely dispositional concepts. Behavior genetics is, however, a powerful movement to the contrary. As regards categorical entities (types, taxa, syndromes, diseases), history reveals no impressive examples of their discovery by cluster algorithms; whereas organic medicine and psychopathology have both discovered many taxonic entities without reliance on formal (statistical) cluster methods. I offer eight reasons for this strange condition, with associated suggestions for ameliorating it. Adopting a realist instead of a fictionist approach to taxonomy, I give high priority to theory-based mathematical derivation of quantitative consistency tests for all taxometric results. I urge a large scale cooperative survey of taxometric methods based on Monte Carlo runs, biological pseudoproblems where the true axon is independently known, and live problem in genetics, organic medicine, and psychopathology. An empirical example of taxometric bootstrapping and consistency testing was presented from my own current research on schizotypy.

Humans↗

Molecular profiles of allograft rejection following inhibition of CD40 ligand costimulation differentiated by cluster analysis.

Recent technological advances in biomedical research, such as genome sequences and DNA microarrays, have dramatically increased the size of relevant databases. A major challenge is the extraction of a limited number of parameters from these databases that can differentiate and diagnose complex biological states. In a model of cardiac transplantation investigating immunosuppression by inhibition of CD40 ligand costimulation, we have applied a combination of cluster algorithms and self-organizing maps to analyze a panel of 60 candidate genes. Dendrograms generated by cluster analysis distinguished different molecular bases of rejection. Using self-organizing maps, we identified nine genes (CD4, CCR3, CCR5, LT beta, MIP-1 alpha, MIP-2, CD8 alpha, IP-10, and RANTES), each with a unique profile of transcriptional expression, that reproduce the differentiation of states of rejection in dendrograms. Using histology and immunohistochemistry, we correlated differential regulation of CD4 and CD8 at the levels of mRNA and protein. Our strategy of data reduction successfully decreased the number of genes to nine, which are sufficient to differentiate distinct states of rejection in our experimental protocol.

Animals↗

Cardiac MR image segmentation and left ventricle surface reconstruction based on level set method.

A two-stage segmentation algorithm is presented to solve the problems of inhomogeneity, weak edges and artifacts exhibited in the magnetic resonance imaging (MRI) images. First, the K-mean clustering algorithm is applied to classify the objects. Then, a speed function based on the clustering results is defined in order to search the rough boundary. Secondly, a speed function of the gradient intensity is constructed to locate the boundary accurately. Due to the lack of deformation information of the boundaries between MR slices, a deformable model is used to reconstruct the shape of the LV: a dynamic equation governing the surface deformation is given; from the slice data, external forces are constructed and elastic forces are provided with mean curvatures of the deformation surface. The level set method is applied to solve the dynamic equation for the LV shape. Experimental results demonstrate the effectiveness of the algorithm listed in the paper.

Algorithms↗

Prediction and statistics of pseudoknots in RNA structures using exactly clustered stochastic simulations.

Ab initio RNA secondary structure predictions have long dismissed helices interior to loops, so-called pseudoknots, despite their structural importance. Here we report that many pseudoknots can be predicted through long-time-scale RNA-folding simulations, which follow the stochastic closing and opening of individual RNA helices. The numerical efficacy of these stochastic simulations relies on an O(n2) clustering algorithm that computes time averages over a continuously updated set of n reference structures. Applying this exact stochastic clustering approach, we typically obtain a 5- to 100-fold simulation speed-up for RNA sequences up to 400 bases, while the effective acceleration can be as high as 105-fold for short, multistable molecules (<or=150 bases). We performed extensive folding statistics on random and natural RNA sequences and found that pseudoknots are distributed unevenly among RNA structures and account for up to 30% of base pairs in G+C-rich RNA sequences (online RNA-folding kinetics server including pseudoknots: http://kinefold.u-strasbg.fr).

Algorithms↗

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing↗

Integrated biclustering of heterogeneous genome-wide datasets for the inference of global regulatory networks.

BACKGROUND: The learning of global genetic regulatory networks from expression data is a severely under-constrained problem that is aided by reducing the dimensionality of the search space by means of clustering genes into putatively co-regulated groups, as opposed to those that are simply co-expressed. Be cause genes may be co-regulated only across a subset of all observed experimental conditions, biclustering (clustering of genes and conditions) is more appropriate than standard clustering. Co-regulated genes are also often functionally (physically, spatially, genetically, and/or evolutionarily) associated, and such a priori known or pre-computed associations can provide support for appropriately grouping genes. One important association is the presence of one or more common cis-regulatory motifs. In organisms where these motifs are not known, their de novo detection, integrated into the clustering algorithm, can help to guide the process towards more biologically parsimonious solutions. RESULTS: We have developed an algorithm, cMonkey, that detects putative co-regulated gene groupings by integrating the biclustering of gene expression data and various functional associations with the de novo detection of sequence motifs. CONCLUSION: We have applied this procedure to the archaeon Halobacterium NRC-1, as part of our efforts to decipher its regulatory network. In addition, we used cMonkey on public data for three organisms in the other two domains of life: Helicobacter pylori, Saccharomyces cerevisiae, and Escherichia coli. The biclusters detected by cMonkey both recapitulated known biology and enabled novel predictions (some for Halobacterium were subsequently confirmed in the laboratory). For example, it identified the bacteriorhodopsin regulon, assigned additional genes to this regulon with apparently unrelated function, and detected its known promoter motif. We have performed a thorough comparison of cMonkey results against other clustering methods, and find that cMonkey biclusters are more parsimonious with all available evidence for co-regulation.

Algorithms↗

GRAM and genfragII: solving and testing the single-digest, partially ordered restriction map problem.

GRAM (Genomic Restriction map AsseMbly) takes as input single-digest restriction fragments for a set of overlapping clones and outputs one or more plausible partially ordered restriction maps. For each restriction map, GRAM shows the corresponding alignment of the input clone fragments. Due to the error and uncertainty in experimental data, this problem is computationally difficult to solve; therefore, the principle objective in the design of GRAM is to facilitate man-machine collaborative problem solving. GRAM quickly approximates a solution, as follows. (i) A clustering algorithm determines a probable set of restriction fragments. (ii) An assembly algorithm permutes the set of restriction fragments such that the maximal number of clone fragments are contiguous. The output of the GRAM algorithm is displayed for the user to query and edit. This paper describes the stochastic assembly algorithm and shows how it works with the interactive graphics to support man-machine problem solving. In order to test and verify the performance of GRAM, we have developed a program called genfragII to simulate the digestion of clones and fragments; this program is described and results are presented. GRAM is also being used for a number of genome mapping projects.

Algorithms↗

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics↗

Identification of gene expression signatures in autoimmune disease without the influence of familial resemblance.

Even though autoimmune diseases are heterogeneous, believed to result from the interaction between genetic and environmental components, patients with these disorders exhibit reproducible patterns of gene expression in their peripheral blood mononuclear cells. A portion of this gene expression profile is a property of familial resemblance rather than autoimmune disease. Here, we wanted to identify the portion of this gene expression profile that is independent of familial resemblance and determine whether it is a product of disease duration, disease onset or other factors. By employing supervised clustering algorithms, we identified 100 genes whose expression profiles are shared in individuals with various autoimmune diseases but are not shared by unaffected family members of individuals with autoimmune disease or by controls. Individuals with early disease (1 year after onset) and established disease (10 years after onset) exhibit a near-identical expression pattern, suggesting that this unique profile is a product of disease onset rather than disease duration.

Algorithms↗

Consensus clustering and functional interpretation of gene-expression data.

Microarray analysis using clustering algorithms can suffer from lack of inter-method consistency in assigning related gene-expression profiles to clusters. Obtaining a consensus set of clusters from a number of clustering methods should improve confidence in gene-expression analysis. Here we introduce consensus clustering, which provides such an advantage. When coupled with a statistically based gene functional analysis, our method allowed the identification of novel genes regulated by NFkappaB and the unfolded protein response in certain B-cell lymphomas.

Cluster Analysis↗

Clinical assessment of hand-arm vibration syndrome.

The clinical assessment of patients thought to be suffering from hand-arm vibration syndrome (HAVS) requires the use of multiple vascular and sensory tests. In a family physician's office, Adson's, Allen's and cold water immersion of the hands are the only feasible vascular tests, while the sensory tests have to be limited to assessing impairment of skin sensitivity and manipulative dexterity. This paper reviews the laboratory tests deemed to be useful in a hospital or clinic facility, and reports on the investigation of 364 patients exposed to hand-arm vibration who were examined in Toronto, Canada during the period 1989-92. A statistical clustering algorithm was used to categorise 138 male subjects according to the results of their diagnostic tests. From the cluster analysis, four vascular and four sensorineural categories of impairment were recognised in patients suffering from HAVS. The Stockholm vascular classification stages and the four vascular clusters were found to correspond. The Stockholm sensorineural classification (Stages 1, 2, and 3) correlated with clusters formed from the sensory tests evaluating the sensitivity of the nerve endings and the distal digital branches of the median and ulnar nerves. When the myelinated nerve fibres were affected, as detected by abnormal Tinel's, Phalen's, and nerve conduction tests, an additional cluster group emerged. The subjects with abnormal nerve conduction test results constituted a distinct group with increased impairment, so there is a need for them to be categorised separately i.e. as a Stage 4. It is suggested that a Stage 4 be included in the Stockholm sensorineural classification.

Arm↗

Data-driven analysis approach for biomarker discovery using molecular-profiling technologies.

High-throughput molecular-profiling technologies provide rapid, efficient and systematic approaches to search for biomarkers. Supervised learning algorithms are naturally suited to analyse a large amount of data generated using these technologies in biomarker discovery efforts. The study demonstrates with two examples a data-driven analysis approach to analysis of large complicated datasets collected in high-throughput technologies in the context of biomarker discovery. The approach consists of two analytic steps: an initial unsupervised analysis to obtain accurate knowledge about sample clustering, followed by a second supervised analysis to identify a small set of putative biomarkers for further experimental characterization. By comparing the most widely applied clustering algorithms using a leukaemia DNA microarray dataset, it was established that principal component analysis-assisted projections of samples from a high-dimensional molecular feature space into a few low dimensional subspaces provides a more effective and accurate way to explore visually and identify data structures that confirm intended experimental effects based on expected group membership. A supervised analysis method, shrunken centroid algorithm, was chosen to take knowledge of sample clustering gained or confirmed by the first step of the analysis to identify a small set of molecules as candidate biomarkers for further experimentation. The approach was applied to two molecular-profiling studies. In the first study, PCA-assisted analysis of DNA microarray data revealed that discrete data structures exist in rat liver gene expression and correlated with blood clinical chemistry and liver pathological damage in response to a chemical toxicant diethylhexylphthalate, a peroxisome-proliferator-activator receptor agonist. Sixteen genes were then identified by shrunken centroid algorithm as the best candidate biomarkers for liver damage. Functional annotations of these genes revealed roles in acute phase response, lipid and fatty acid metabolism and they are functionally relevant to the observed toxicities. In the second study, 26 urine ions identified from a GC/MS spectrum, two of which were glucose fragment ions included as positive controls, showed robust changes with the development of diabetes in Zucker diabetic fatty rats. Further experiments are needed to define their chemical identities and establish functional relevancy to disease development.

Algorithms↗

Structural prediction of peptides bound to MHC class I.

An ab initio structure prediction approach adapted to the peptide-major histocompatibility complex (MHC) class I system is presented. Based on structure comparisons of a large set of peptide-MHC class I complexes, a molecular dynamics protocol is proposed using simulated annealing (SA) cycles to sample the conformational space of the peptide in its fixed MHC environment. A set of 14 peptide-human leukocyte antigen (HLA) A0201 and 27 peptide-non-HLA A0201 complexes for which X-ray structures are available is used to test the accuracy of the prediction method. For each complex, 1000 peptide conformers are obtained from the SA sampling. A graph theory clustering algorithm based on heavy atom root-mean-square deviation (RMSD) values is applied to the sampled conformers. The clusters are ranked using cluster size, mean effective or conformational free energies, with solvation free energies computed using Generalized Born MV 2 (GB-MV2) and Poisson-Boltzmann (PB) continuum models. The final conformation is chosen as the center of the best-ranked cluster. With conformational free energies, the overall prediction success is 83% using a 1.00 Angstroms crystal RMSD criterion for main-chain atoms, and 76% using a 1.50 Angstroms RMSD criterion for heavy atoms. The prediction success is even higher for the set of 14 peptide-HLA A0201 complexes: 100% of the peptides have main-chain RMSD values < or =1.00 Angstroms and 93% of the peptides have heavy atom RMSD values < or =1.50 Angstroms. This structure prediction method can be applied to complexes of natural or modified antigenic peptides in their MHC environment with the aim to perform rational structure-based optimizations of tumor vaccines.

Algorithms↗

Distance-based clustering of CGH data.

MOTIVATION: We consider the problem of clustering a population of Comparative Genomic Hybridization (CGH) data samples. The goal is to develop a systematic way of placing patients with similar CGH imbalance profiles into the same cluster. Our expectation is that patients with the same cancer types will generally belong to the same cluster as their underlying CGH profiles will be similar. RESULTS: We focus on distance-based clustering strategies. We do this in two steps. (1) Distances of all pairs of CGH samples are computed. (2) CGH samples are clustered based on this distance. We develop three pairwise distance/similarity measures, namely raw, cosine and sim. Raw measure disregards correlation between contiguous genomic intervals. It compares the aberrations in each genomic interval separately. The remaining measures assume that consecutive genomic intervals may be correlated. Cosine maps pairs of CGH samples into vectors in a high-dimensional space and measures the angle between them. Sim measures the number of independent common aberrations. We test our distance/similarity measures on three well known clustering algorithms, bottom-up, top-down and k-means with and without centroid shrinking. Our results show that sim consistently performs better than the remaining measures. This indicates that the correlation of neighboring genomic intervals should be considered in the structural analysis of CGH datasets. The combination of sim with top-down clustering emerged as the best approach. AVAILABILITY: All software developed in this article and all the datasets are available from the authors upon request. CONTACT: juliu@cise.ufl.edu.

Algorithms↗