PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Sectorization of the central 30 degrees visual field in glaucoma.

PURPOSE: To determine an optimal sector pattern of the central 30 degrees visual field in glaucoma by mathematically analyzing the visual field data of primary open-angle glaucoma (POAG) without any assumption such as the retinal nerve fiber layer anatomy. METHODS: One hundred three visual fields of the 30-2 program of the Humphrey Field Analyzer obtained from 103 POAG patients of early to moderately advanced stage were included. Based on the interpoint correlation of deviation of the measured threshold value from the age-corrected normal reference value (total deviation, STATPAC), test points of the 30-2 program were mathematically clustered using the VARCLUS procedure, a new clustering algorithm developed by the SAS Institute. The sector value, which summarizes the visual field performance of the clustered test points (sector), also was calculated. RESULTS: The 30 degrees central visual field was divided into 15 sectors consisting of at least 3 points. The distribution of sectors was compatible with the projection of nerve fiber layers. There was no sector extending over the horizontal meridian, but the sector pattern was not completely symmetrical around it. Linear regression analysis of the sector values against the mean deviation (STATPAC) suggested that the index is useful in following visual field performance of each sector. CONCLUSION: The sector pattern and sector values obtained were considered useful in studying the visual field data of glaucoma.

Algorithms↗

Computer image analysis of two-dimensional crystals of beef heart NADH: ubiquinone oxidoreductase fragments. I. Comparison of crystal structures in various negative stains.

We investigated the structure of two-dimensional crystals from bovine heart mitochondrial NADH: ubiquinone oxidoreductase. A detailed description of uranyl acetate-stained crystals demonstrated that they are composed of fragments in a spatial arrangement according to space group P4212 [J. Brink, S. Hovmöller, C.I. Ragan, M.W.J. Cleeter, E.J. Boekema and E.F.J. van Bruggen, European J. Biochem. 166 (1987) 287]. To gain more structural information on the crystal structure and to assess the effects of various negative stains on the structure preservation and appearance, we examined stained crystals by means of electron microscopy and image analysis. The space group P4212 appeared to be present for several stains tested, i.e. ammonium molybdate, uranyl acetate, uranyl nitrate and uranyl sulphate. Use of phosphotungstic acid and silicotungstate resulted in a reduction of symmetry to pseudo-P4212 or p4. Use of sodium tungstate led to a considerable loss of resolution to 3.8 nm at best, whereas otherwise 1.5 to 1.9 nm could be demonstrated. The lattice vectors were not affected by the stains; they were determined as a = b = 14.9 +/- 0.25 nm with gamma = 89.8 degrees +/- 0.6 degrees. Image analysis showed the presence of similar structures with the molybdate and uranyl compounds. Differences were observed in the case of the tungstate type of stains. Furthermore, the analysis revealed the complete absence of the four small pores of 2.0 nm diameter in the unit cell. This effect was observed irrespective of the type of stain and supporting film, and could be ascribed only to the glow-discharge treatment of the supporting film. The observed difference must be caused by changed interactions between the protein, stain and supporting film. Application of correspondence analysis and clustering algorithms to the various reconstructed images of the crystals showed that they could be separated into several clusters. Each of these clusters corresponded on the average to only one type of stain, whereas a further division according to the specific uranyl compounds was observed. This study therefore shows that under identical preparation conditions subtle differences between individual stains can be detected.

Animals↗

Continuous representations of time-series gene expression data.

We present algorithms for time-series gene expression analysis that permit the principled estimation of unobserved time points, clustering, and dataset alignment. Each expression profile is modeled as a cubic spline (piecewise polynomial) that is estimated from the observed data and every time point influences the overall smooth expression curve. We constrain the spline coefficients of genes in the same class to have similar expression patterns, while also allowing for gene specific parameters. We show that unobserved time points can be reconstructed using our method with 10-15% less error when compared to previous best methods. Our clustering algorithm operates directly on the continuous representations of gene expression profiles, and we demonstrate that this is particularly effective when applied to nonuniformly sampled data. Our continuous alignment algorithm also avoids difficulties encountered by discrete approaches. In particular, our method allows for control of the number of degrees of freedom of the warp through the specification of parameterized functions, which helps to avoid overfitting. We demonstrate that our algorithm produces stable low-error alignments on real expression data and further show a specific application to yeast knock-out data that produces biologically meaningful results.

Algorithms↗

A funny thing happened to us on the way to the latent entities.

Inferred latent entities, whether those of psychoanalysis, factor analysis, or cluster analysis, have declined in value for many clinical psychologists, both as tools of practice and as objects of theoretical interest. Behavior modification, rational-emotive therapy, crisis intervention, psycho-pharmacology, and actuarial prediction all tend to minimize reliance on latent entities in favor of purely dispositional concepts. Behavior genetics is, however, a powerful movement to the contrary. As regards categorical entities (types, taxa, syndromes, diseases), history reveals no impressive examples of their discovery by cluster algorithms; whereas organic medicine and psychopathology have both discovered many taxonic entities without reliance on formal (statistical) cluster methods. I offer eight reasons for this strange condition, with associated suggestions for ameliorating it. Adopting a realist instead of a fictionist approach to taxonomy, I give high priority to theory-based mathematical derivation of quantitative consistency tests for all taxometric results. I urge a large scale cooperative survey of taxometric methods based on Monte Carlo runs, biological pseudoproblems where the true axon is independently known, and live problem in genetics, organic medicine, and psychopathology. An empirical example of taxometric bootstrapping and consistency testing was presented from my own current research on schizotypy.

Humans↗

Molecular profiles of allograft rejection following inhibition of CD40 ligand costimulation differentiated by cluster analysis.

Recent technological advances in biomedical research, such as genome sequences and DNA microarrays, have dramatically increased the size of relevant databases. A major challenge is the extraction of a limited number of parameters from these databases that can differentiate and diagnose complex biological states. In a model of cardiac transplantation investigating immunosuppression by inhibition of CD40 ligand costimulation, we have applied a combination of cluster algorithms and self-organizing maps to analyze a panel of 60 candidate genes. Dendrograms generated by cluster analysis distinguished different molecular bases of rejection. Using self-organizing maps, we identified nine genes (CD4, CCR3, CCR5, LT beta, MIP-1 alpha, MIP-2, CD8 alpha, IP-10, and RANTES), each with a unique profile of transcriptional expression, that reproduce the differentiation of states of rejection in dendrograms. Using histology and immunohistochemistry, we correlated differential regulation of CD4 and CD8 at the levels of mRNA and protein. Our strategy of data reduction successfully decreased the number of genes to nine, which are sufficient to differentiate distinct states of rejection in our experimental protocol.

Animals↗

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing↗

GRAM and genfragII: solving and testing the single-digest, partially ordered restriction map problem.

GRAM (Genomic Restriction map AsseMbly) takes as input single-digest restriction fragments for a set of overlapping clones and outputs one or more plausible partially ordered restriction maps. For each restriction map, GRAM shows the corresponding alignment of the input clone fragments. Due to the error and uncertainty in experimental data, this problem is computationally difficult to solve; therefore, the principle objective in the design of GRAM is to facilitate man-machine collaborative problem solving. GRAM quickly approximates a solution, as follows. (i) A clustering algorithm determines a probable set of restriction fragments. (ii) An assembly algorithm permutes the set of restriction fragments such that the maximal number of clone fragments are contiguous. The output of the GRAM algorithm is displayed for the user to query and edit. This paper describes the stochastic assembly algorithm and shows how it works with the interactive graphics to support man-machine problem solving. In order to test and verify the performance of GRAM, we have developed a program called genfragII to simulate the digestion of clones and fragments; this program is described and results are presented. GRAM is also being used for a number of genome mapping projects.

Algorithms↗

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics↗

Clinical assessment of hand-arm vibration syndrome.

The clinical assessment of patients thought to be suffering from hand-arm vibration syndrome (HAVS) requires the use of multiple vascular and sensory tests. In a family physician's office, Adson's, Allen's and cold water immersion of the hands are the only feasible vascular tests, while the sensory tests have to be limited to assessing impairment of skin sensitivity and manipulative dexterity. This paper reviews the laboratory tests deemed to be useful in a hospital or clinic facility, and reports on the investigation of 364 patients exposed to hand-arm vibration who were examined in Toronto, Canada during the period 1989-92. A statistical clustering algorithm was used to categorise 138 male subjects according to the results of their diagnostic tests. From the cluster analysis, four vascular and four sensorineural categories of impairment were recognised in patients suffering from HAVS. The Stockholm vascular classification stages and the four vascular clusters were found to correspond. The Stockholm sensorineural classification (Stages 1, 2, and 3) correlated with clusters formed from the sensory tests evaluating the sensitivity of the nerve endings and the distal digital branches of the median and ulnar nerves. When the myelinated nerve fibres were affected, as detected by abnormal Tinel's, Phalen's, and nerve conduction tests, an additional cluster group emerged. The subjects with abnormal nerve conduction test results constituted a distinct group with increased impairment, so there is a need for them to be categorised separately i.e. as a Stage 4. It is suggested that a Stage 4 be included in the Stockholm sensorineural classification.

Arm↗

Inference from clustering with application to gene-expression microarrays.

There are many algorithms to cluster sample data points based on nearness or a similarity measure. Often the implication is that points in different clusters come from different underlying classes, whereas those in the same cluster come from the same class. Stochastically, the underlying classes represent different random processes. The inference is that clusters represent a partition of the sample points according to which process they belong. This paper discusses a model-based clustering toolbox that evaluates cluster accuracy. Each random process is modeled as its mean plus independent noise, sample points are generated, the points are clustered, and the clustering error is the number of points clustered incorrectly according to the generating random processes. Various clustering algorithms are evaluated based on process variance and the key issue of the rate at which algorithmic performance improves with increasing numbers of experimental replications. The model means can be selected by hand to test the separability of expected types of biological expression patterns. Alternatively, the model can be seeded by real data to test the expected precision of that output or the extent of improvement in precision that replication could provide. In the latter case, a clustering algorithm is used to form clusters, and the model is seeded with the means and variances of these clusters. Other algorithms are then tested relative to the seeding algorithm. Results are averaged over various seeds. Output includes error tables and graphs, confusion matrices, principal-component plots, and validation measures. Five algorithms are studied in detail: K-means, fuzzy C-means, self-organizing maps, hierarchical Euclidean-distance-based and correlation-based clustering. The toolbox is applied to gene-expression clustering based on cDNA microarrays using real data. Expression profile graphics are generated and error analysis is displayed within the context of these profile graphics. A large amount of generated output is available over the web.

Computational Biology↗

Multivariate analysis of the ecoregion delineation for aquatic systems.

The ecoregion concept is a popular method of understanding the spatial distribution of the environment', however, it has yet to be adequately demonstrated that the environment is distributed in accordance with these bounded units. In this paper, we generated a testable hypothesis based on the current usage of ecoregions: the ecoregion classification will allow for discrimination between lakes of different water quality. The ecoregion classification should also be more effective better than a comparably scaled classification based on political boundaries, land-use class, or random grouping. To test this hypothesis we used the Environmental Monitoring and Assessment Program (EMAP) lake water chemistry data from the northeast United States. The water chemistry data were reduced to four components using principal component analysis. For comparison to an optimal grouping of these data we used K-means cluster analysis to define the extent at which these lakes could be segregated into distinct classes. Jackknifed discriminant analysis was used to determine the classification rate of ecoregions, the three alternative spatial classification methods, and the clustering algorithm. The classification based on ecoregions was successful for 35% of the lakes included in this study, in comparison to the clustered groups accuracy of 98%. These results suggest that the large scale spatial distribution of ecosystem types is more complicated than that suggested by the present ecoregion boundaries. Further tests of ecoregion delineations are needed and alternative large-scale management strategies should be investigated.

Algorithms↗

Computer-aided detection of clustered microcalcifications on digital mammograms.

A computer-aided diagnosis scheme to assist radiologists in detecting clustered microcalcifications from mammograms is being developed. Starting with a digital mammogram, the scheme consists of three steps. First, the image is filtered so that the signal-to-noise ratio of microcalcifications is increased by suppression of the normal background structure of the breast. Secondly, potential microcalcifications are extracted from the filtered image with a series of three different techniques: a global thresholding based on the grey-level histogram of the full filtered image, an erosion operator for eliminating very small signals, and a local adaptive grey-level thresholding. Thirdly, some false-positive signals are eliminated by means of a texture analysis technique, and a non-linear clustering algorithm is then used for grouping the remaining signals. With this method, the scheme can detect approximately 85% of true clusters, with an average of two false clusters detected per image.

Breast Diseases↗

Computerised intrapartum diagnosis of fetal hypoxia based on fetal heart rate monitoring and fetal pulse oximetry recordings utilising wavelet analysis and neural networks.

OBJECTIVE: To develop a computerised system that will assist the early diagnosis of fetal hypoxia and to investigate the relationship between the fetal heart rate variability and the fetal pulse oximetry recordings. DESIGN: Retrospective off-line analysis of cardiotocogram and FSpO2 recordings. SETTING: The Maternity Unit of the 2nd Department of Obstetrics and Gynaecology, Aretaieion Hospital, University of Athens. POPULATION: Sixty-one women of more than 37 weeks of gestation were monitored throughout labour. METHODS: Multiresolution wavelet analysis was applied in each 10-minute period of second stage of labour focussing on long term variability changes in different frequency ranges and statistical analysis was performed in the associated 10-minute FSpO2 recordings. Self-organising map neural network was used to categorise the different 10-minute fetal heart rate patterns and the associated 10-minute FSpO2 recordings. MAIN OUTCOME MEASURES: Umbilical artery pH of < or = 7.20 and Apgar score at 5 minutes of < or = 7 formed the inclusion criteria of the risk group. RESULTS: After using k-means clustering algorithm, the two-dimensional output layer of the self-organising map neural network was divided into three distinct clusters. All the cases that mapped in cluster 3 belonged in the risk group except one. The sensitivity of the system was 83.3% and the specificity 97.9% for the detection of risk group cases. CONCLUSIONS: A relationship between the fetal heart rate variability in different frequency ranges and the time in which FSpO2 is less than 30% was noticed. Fetal pulse oximetry seems to be an important additional source of information. Computerised analysis of the fetal heart rate monitoring and pulse oximetry recordings is a promising technique in objective intrapartum diagnosis of fetal hypoxia. Further evaluation of this technique is mandatory to evaluate its efficacy and reliability in interpreting fetal heart rate recordings.

Adult↗

ProtoMap: automatic classification of protein sequences, a hierarchy of protein families, and local maps of the protein space.

We investigate the space of all protein sequences in search of clusters of related proteins. Our aim is to automatically detect these sets, and thus obtain a classification of all protein sequences. Our analysis, which uses standard measures of sequence similarity as applied to an all-vs.-all comparison of SWISSPROT, gives a very conservative initial classification based on the highest scoring pairs. The many classes in this classification correspond to protein subfamilies. Subsequently we merge the subclasses using the weaker pairs in a two-phase clustering algorithm. The algorithm makes use of transitivity to identify homologous proteins; however, transitivity is applied restrictively in an attempt to prevent unrelated proteins from clustering together. This process is repeated at varying levels of statistical significance. Consequently, a hierarchical organization of all proteins is obtained. The resulting classification splits the protein space into well-defined groups of proteins, which are closely correlated with natural biological families and superfamilies. Different indices of validity were applied to assess the quality of our classification and compare it with the protein families in the PROSITE and Pfam databases. Our classification agrees with these domain-based classifications for between 64.8% and 88.5% of the proteins. It also finds many new clusters of protein sequences which were not classified by these databases. The hierarchical organization suggested by our analysis reveals finer subfamilies in families of known proteins as well as many novel relations between protein families.

Algorithms↗

Post-acquisition correction of MR inhomogeneities.

Signal inhomogeneities in volumetric head MR scans are a major obstacle to segmentation and neuromorphometry. The fuzzy c-means (FCM) statistical clustering algorithm was extended to estimate and retrospectively correct a multiplicative inhomogeneity field in T1-weighted head MR scans. The method was tested on a mathematically simulated object and on seven whole head 3D MR scans. Once initial parameters governing operation of the algorithm were chosen for this class of images, results were obtained without intervention for individual MR studies. Post-acquisition inhomogeneity correction by extended FCM clustering improved overall image uniformity and separability of gray and white matter intensities.

Algorithms↗

Probabilistic clustering of sequences: inferring new bacterial regulons by comparative genomics.

Genome-wide comparisons between enteric bacteria yield large sets of conserved putative regulatory sites on a gene-by-gene basis that need to be clustered into regulons. Using the assumption that regulatory sites can be represented as samples from weight matrices (WMs), we derive a unique probability distribution for assignments of sites into clusters. Our algorithm, "PROCSE" (probabilistic clustering of sequences), uses Monte Carlo sampling of this distribution to partition and align thousands of short DNA sequences into clusters. The algorithm internally determines the number of clusters from the data and assigns significance to the resulting clusters. We place theoretical limits on the ability of any algorithm to correctly cluster sequences drawn from WMs when these WMs are unknown. Our analysis suggests that the set of all putative sites for a single genome (e.g., Escherichia coli) is largely inadequate for clustering. When sites from different genomes are combined and all the homologous sites from the various species are used as a block, clustering becomes feasible. We predict 50-100 new regulons as well as many new members of existing regulons, potentially doubling the number of known regulatory sites in E. coli.

Bacteria↗

Generalizing the plurality method for forming hospital service areas.

The upcoming Health Care Financing Administration's Fourth Scope of Work for peer review organizations (PROs) envisions much use of geographic analysis of utilization rates and quality of care. Proper analysis of utilization rates requires each PRO to form multiple sets of hospital service areas. The method used most often in the literature is the plurality method. Because this method can create fractured service areas and can leave hospitals without a service area, the service areas and their associated hospitals often are reworked by hand. This last step drastically raises the effort required to form service areas and makes the method nonreproducible. This report defines the generalized plurality method for forming the hospital service areas that are central to the study of use patterns via small area analysis. This new method is a true generalization of the plurality method. Like the plurality method, it forms service areas by allowing geographic areas to "vote" for their preferred hospital. The generalization is achieved by allowing near-ties in the voting to cause clustering of hospitals. Hence it is a clustering algorithm that operates on both the geographic areas (sources) and the hospitals (destinations) at the same time. It is a nonhierarchical, nonagglomerative cluster method. There are several free parameters that may be chosen to adjust the effect of the clustering by adjusting the definition of a near-tie, as well as the sensitivity of the clustering to near-ties from sources with a small number of votes. This automated method enjoys many major advantages over the methods commonly appearing in the literature: it is completely reproducible, it is quick, and it does not require the a priori convening of a panel of experts. It thus can be applied easily to a wide variety of types of care that would not necessarily have the same service areas.

Algorithms↗

Distinctive gene expression patterns in human mammary epithelial cells and breast cancers.

cDNA microarrays and a clustering algorithm were used to identify patterns of gene expression in human mammary epithelial cells growing in culture and in primary human breast tumors. Clusters of coexpressed genes identified through manipulations of mammary epithelial cells in vitro also showed consistent patterns of variation in expression among breast tumor samples. By using immunohistochemistry with antibodies against proteins encoded by a particular gene in a cluster, the identity of the cell type within the tumor specimen that contributed the observed gene expression pattern could be determined. Clusters of genes with coherent expression patterns in cultured cells and in the breast tumors samples could be related to specific features of biological variation among the samples. Two such clusters were found to have patterns that correlated with variation in cell proliferation rates and with activation of the IFN-regulated signal transduction pathway, respectively. Clusters of genes expressed by stromal cells and lymphocytes in the breast tumors also were identified in this analysis. These results support the feasibility and usefulness of this systematic approach to studying variation in gene expression patterns in human cancers as a means to dissect and classify solid tumors.

Algorithms↗