PubMed HealthSearch

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Tumor-immune partitioning and clustering algorithm for identifying tumor-immune cell spatial interaction signatures within the tumor microenvironment.

BACKGROUND: Growing evidence supports the importance of characterizing the organizational patterns of various cellular constituents in the tumor microenvironment in precision oncology. Most existing data on immune cell infiltrates in tumors, which are based on immune cell counts or nearest neighbor-type analyses, have failed to fully capture the cellular organization and heterogeneity. METHODS: We introduce a computational algorithm, termed Tumor-Immune Partitioning and Clustering (TIPC), that jointly measures immune cell partitioning between tumor epithelial and stromal areas and immune cell clustering versus dispersion. As proof-of-principle, we applied TIPC to a prospective cohort incident tumor biobank containing 931 colorectal carcinoma cases. TIPC identified tumor subtypes with unique spatial patterns between tumor cells and T lymphocytes linked to certain molecular pathologic and prognostic features. T lymphocyte identification and phenotyping were achieved using multiplexed (multispectral) immunofluorescence. In a separate hepatocellular carcinoma cohort, we replaced the stromal component with specific immune cell types-CXCR3+CD68+ or CD8+-to profile their spatial relationships with CXCL9+CD68+ cells. RESULTS: Six unsupervised TIPC subtypes based on T lymphocyte distribution patterns were identified, comprising two cold and four hot subtypes. Three of the four hot subtypes were associated with significantly longer colorectal cancer (CRC)-specific survival compared to a reference cold subtype. Our analysis showed that variations in T-cell densities among the TIPC subtypes did not strictly correlate with prognostic benefits, underscoring the prognostic significance of immune cell spatial patterns. Additionally, TIPC revealed two spatially distinct and cell density-specific subtypes among microsatellite instability-high colorectal cancers, indicating its potential to upgrade tumor subtyping. TIPC was also applied to additional immune cell types, eosinophils and neutrophils, identified using morphology and supervised machine learning; here two tumor subtypes with similarly low densities, namely 'cold, tumor-rich' and 'cold, stroma-rich', exhibited differential prognostic associations. Lastly, we validated our methods and results using The Cancer Genome Atlas colon and rectal adenocarcinoma data (n = 570). Moreover, applying TIPC to hepatocellular carcinoma cases (n = 27) highlighted critical cell interactions like CXCL9-CXCR3 and CXCL9-CD8. CONCLUSIONS: Unsupervised discoveries of microgeometric tissue organizational patterns and novel tumor subtypes using the TIPC algorithm can deepen our understanding of the tumor immune microenvironment and likely inform precision cancer immunotherapy.

Humans

Cluster analyses of cardiovascular responsivity to three laboratory stressors.

Seventy-three young normotensive male subjects were tested with an experimental protocol that included a reaction time, a mental arithmetic, and a cold pressor task. Physiological variables that were recorded included heart rate, stroke volume, pre-ejection period, blood pressure, total peripheral resistance, and respiratory sinus arrhythmia. In order to identify subgroups of subjects who differed in their pattern of autonomic responses to the tasks, the physiological change scores from baseline to the tasks for each subject were entered into a cluster analysis for each task. Ward's method was used as the clustering algorithm. The cluster analyses identified four clusters for the reaction time and mental arithmetic tasks, and five clusters for the cold pressor task. Although there was a wide range of patterns exhibited by cluster subgroups, most subjects who were reactive to the tasks showed response patterns that were qualitatively similar to the pattern of overall mean response by all subjects, albeit varying considerably in terms of quantitative response. Little evidence was generated for the consistency of extreme beta-adrenergic response from one task to another, although significant consistency was noted when milder beta-responders were included in the comparisons. Some consistency of alpha-adrenergic response noted across tasks, as well as significant consistency of being relatively nonreactive to the tasks.

Adult

Numerical taxonomy of Actinomadura and related actinomycetes.

One hundred and fifty-six Actinomadura strains, marker strains of related taxa, and related isolates from bagasse and fodder were the subject of numerical phenetic analyses using 90 unit characters. The data were examined using the simple matching (SSM), Jaccard (SJ) and pattern (DP) coefficients and clustering was achieved using both single and average linkage algorithms. Cluster composition was not markedly affected either by the coefficient or clustering algorithms used or by test error, estimated at 4.5%. Actinomadura dassonvillei, Actinomadura madurae and Streptomyces somaliensis formed good taxospecies, but the separation of Actinomadura pelletieri strains into two clusters by SJ and SSM analysis requires further study. The single representatives of Actinomadura helvata, Actinomadura pusilla, Actinomadura roseoviolacea, Actinomadura spadix and Actinomadura verrucosopora seemed to form new centres of variation while Actinomadura citrea and Actinomadura malachitica showed much similarity with Actinomadura madurae. Most of the isolates form bagasse and fodder were recovered in two well-defined phena, provisionally labelled clusters 'A' and 'B' which showed little similarity to either Actinomadura or Nocardia strains. The effect of the different coefficients on the aggregation of clusters is discussed.

Nocardiaceae

Application of stepwise cluster analysis in medical research.

A stepwise clustering algorithm, a method of multivariate statistical analysis, is suggested in this paper. The algorithm is designed for solving problems connected with stepwise regression. It is efficient not only in handling both continuous and discrete variables, but also in the nonlinear relationships between the variables. The above procedure was used in an attempt to find out the causal association of esophageal cancer with its precursors, i.e. nitrates and nitrites of nitrosamines, some of which are known to be carcinogenic. An analysis has been made of the correlation between esophageal cancer as well as severe epithelial hyperplasia of the esophagus and the concentrations of NO3- and NO2- in the drinking water. The samples used were collected from 495 wells in 49 production brigades of the Yaocun Commune in Linxian County, Honan Province. The result indicates that esophageal cancer is definitely connected with the levels of NO3- (summer) and NO2- (spring) in the drinking water. Severe epithelial hyperplasia is defintely connected with the contents of NO2- and NO3- in the drinking water collected in spring, autumn and winter. Our preliminary analysis shows that the stepwise clustering algorithm is a useful statistical method to be used for medical research.

Humans

Comparing multiple RNA secondary structures using tree comparisons.

In a previous paper, an algorithm was presented for analyzing multiple RNA secondary structures utilizing a multiple string alignment algorithm. In this paper we present another approach to the problem of comparing many secondary structures by utilizing a very efficient tree-matching algorithm that will compare two trees in O([T1] X [T2] X L1 X L2) in the worst case and very close to O([T1] X [T2]) for average trees representing secondary structures. The result of the pairwise comparison algorithm is then used with a cluster algorithm to produce a multiple structure clustering which can be displayed in a taxonomy tree to show related structures.

Algorithms

Rare genetic variant risks in patients with sepsis-associated acute respiratory distress syndrome.

BACKGROUND: Acute respiratory distress syndrome (ARDS) is a complex, heterogeneous, and deadly condition often resulting from pulmonary lesions due to sepsis, among other causes. There is a lack of targeted therapies to specifically treat the patients. Common genetic factors in the population (frequency&#x2009;>&#x2009;1%) have been associated with ARDS susceptibility, but systematic genetic screens of the role of rare genetic variants are lacking. We used the network of known molecular interactions to identify ARDS risks from clusters of biologically related genes containing qualifying variants (QVs) with frequency&#x2009;<&#x2009;1% likely affecting function. METHODS: We conducted whole-exome sequencing in sepsis patients from the GEN-SEP cohort (n&#x2009;=&#x2009;822, of which 272 developed ARDS). A network-based heterogeneity clustering algorithm was used to discover significant gene clusters (p&#x2009;<&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5). Gene-set enrichment analysis and logistic regression models aggregating QVs were used for cross-verification to confirm consistency and deepen understanding of the effect sizes of gene clusters. RESULTS: We identified 19 significant clusters (plowest&#x2009;=&#x2009;3.29&#x2009;&#xd7;&#x2009;10&#x2013;10), each containing an average of 102 genes (11.6% mean similarity). QVs in nine gene clusters were associated with sepsis-associated ARDS (plowest&#x2009;=&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5) but were not associated with 28-day survival. Clusters were enriched in several biological pathways, notably the Toll-like receptor cascades. CONCLUSIONS: These results support a marked genetic heterogeneity underlying ARDS susceptibility and the presence of rare risk variants involving multiple biological processes that are associated with sepsis outcomes. Particularly, they underscore the importance of rare variants in genes of the Toll-like receptor cascades in the risk for sepsis-associated ARDS.

Humans

Personality characteristics of substance abusers: an MCMI cluster typology of recreational drug users treated in a therapeutic community and its relationship to length of stay and outcome.

The purpose of this investigation was to define homogeneous personality subtypes among substance abusers treated in a long-term, inpatient, drug-free therapeutic community and to determine how the resulting typology was related to length of stay and treatment outcome. A hierarchical agglomerative cluster analysis was performed on the Millon Clinical Multiaxial Inventory (MCMI) scale scores of 235 admissions to a therapeutic community. Five cluster types emerged, which were similar to typologies found in studies with alcoholic inpatients. A concordant solution evolved when a different clustering algorithm was used with the same sample and when clustering was done with a different group of substance abusers. As hypothesized, clusters of patients with average MCMI elevations that indicated avoidant, schizoid, and antisocial qualities tended to stay in treatment fewer days and relapsed earlier during the 1-year follow-up. The implications for substance abuse treatment are discussed.

Adult

Selection of a representative set of structures from Brookhaven Protein Data Bank.

Reliable structural and statistical analyses of three dimensional protein structures should be based on unbiased data. The Protein Data Bank is highly redundant, containing several entries for identical or very similar sequences. A technique was developed for clustering the known structures based on their sequences and contents of alpha- and beta-structures. First, sequences were aligned pairwise. A representative sample of sequences was then obtained by grouping similar sequences together, and selecting a typical representative from each group. The similarity significance threshold needed in the clustering method was found by analyzing similarities of random sequences. Because three dimensional structures for proteins of same structural class are generally more conserved than their sequences, the proteins were clustered also according to their contents of secondary structural elements. The results of these clusterings indicate conservation of alpha- and beta-structures even when sequence similarity is relatively low. An unbiased sample of 103 high resolution structures, representing a wide variety of proteins, was chosen based on the suggestions made by the clustering algorithm. The proteins were divided into structural classes according to their contents and ratios of secondary structural elements. Previous classifications have suffered from subjective view of secondary structures, whereas here the classification was based on backbone geometry. The concise view lead to reclassification of some structures. The representative set of structures facilitates unbiased analyses of relationships between protein sequence, function, and structure as well as of structural characteristics.

Algorithms

Image analysis and quantification of atherosclerosis using MRI.

This paper describes an image processing, pattern recognition, and computer graphics system for the noninvasive identification and evaluation of atherosclerosis using multidimensional Magnetic Resonance Imaging (MRI). Particular emphasis has been placed on the problem of developing a pattern recognition system for noninvasively identifying the different plaque classes involved in atherosclerosis using minimal a priori information. This pattern recognition technique involves an extension of the ISODATA clustering algorithm to include an information theoretic criterion (Consistent Akaike Information Criterion) to provide a measure of the fit of the cluster composition at a particular iteration to the actual data. A rapid 3-D display system is also described for the simultaneous display of multiple data classes resulting from the tissue identification process. This work demonstrates the feasibility of developing a "high information content" display which will aid in the diagnosis and analysis of the atherosclerotic disease process. Such capability will permit detailed and quantitative studies to assess the effectiveness of therapies, such as drug, exercise, and dietary regimens.

Algorithms

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics

A funny thing happened to us on the way to the latent entities.

Inferred latent entities, whether those of psychoanalysis, factor analysis, or cluster analysis, have declined in value for many clinical psychologists, both as tools of practice and as objects of theoretical interest. Behavior modification, rational-emotive therapy, crisis intervention, psycho-pharmacology, and actuarial prediction all tend to minimize reliance on latent entities in favor of purely dispositional concepts. Behavior genetics is, however, a powerful movement to the contrary. As regards categorical entities (types, taxa, syndromes, diseases), history reveals no impressive examples of their discovery by cluster algorithms; whereas organic medicine and psychopathology have both discovered many taxonic entities without reliance on formal (statistical) cluster methods. I offer eight reasons for this strange condition, with associated suggestions for ameliorating it. Adopting a realist instead of a fictionist approach to taxonomy, I give high priority to theory-based mathematical derivation of quantitative consistency tests for all taxometric results. I urge a large scale cooperative survey of taxometric methods based on Monte Carlo runs, biological pseudoproblems where the true axon is independently known, and live problem in genetics, organic medicine, and psychopathology. An empirical example of taxometric bootstrapping and consistency testing was presented from my own current research on schizotypy.

Humans

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Particle track structure and its correlation with radiobiological endpoint.

One of the possible ways to classify track structures is application of the conventional partition techniques of analysis of multidimensional data to the track structure. Using these cluster algorithms this paper attempts to find characteristics of radiation reflecting the spatial distribution of ionizations in the primary particle track. Absolute frequency distributions of clusters giving the mean number of clusters produced by radiation per unit of deposited energy have been computed for radiation of different qualities. The results were compared with the published experimental data of cell inactivation. For particular biological objects the critical properties of radiation correlating with the cell inactivation can be found and it seems that the occurrence of a cluster of at least four ionizations formed in a domain of approximately 2-3 nm correlates with the induction of double strand break.

Ions

Cluster analysis applied to symptom ratings of psychiatric patients: an evaluation of its predictive ability.

Rating on 39 symptoms were examined for patients admitted to the Neuropsychiatric Institute of the University of Michigan Medical Center. A detailed evaluation was made of the clusters derived by a hierarchical clustering algorithm, using complete linkage and a simple matching coefficient on the binary variables of presence or absence of symptoms. The four groups of patients suggested by the cluster analysis can be characterized as follows: (1) generalized multiplicity of symptoms; (2) capacity to cope except for orientation apart from generally held norms; (3) activity level and thought processes speeded up, intensified, and unselected; (4) inwardly punitive, slowed down and distressed. It is shown that these groups received significantly different treatment and that the effect of treatment was significantly different, while no such differences were noted for groups defined in terms of diagnoses. By means of linear discriminant functions, rules are suggested for assigning other psychiatric patients to one of these four groups.

Antipsychotic Agents

Analysis of methotrexate treatment effect in a longitudinal observational study: utility of cluster analysis.

We studied 235 patients with rheumatoid arthritis (RA) beginning therapy with methotrexate utilizing a k-means clustering algorithm. Four groups were identified: mild RA (Group 3), very severe RA (Group 4), and 2 groups intermediate in severity (Groups 1 and 2). Group 2, the largest of the clusters (n = 89), appeared to have greater tolerability of RA as measured by severity and psychological variables, and took the drug almost twice as long as other groups, although improvement was not greater nor side effects fewer. All groups improved over a mean of 1.9 years, and the degree of improvement was not related to the initial severity classification. Improvement occurred almost equally in all clusters, and the relative ranking of the groups was maintained at study closure.

Arthritis, Rheumatoid

5S rRNA sequences of representatives of the genera Chlorobium, Prosthecochloris, Thermomicrobium, Cytophaga, Flavobacterium, Flexibacter and Saprospira and a discussion of the evolution of eubacteria in general.

5S rRNA sequences were determined for the green sulphur bacteria Chlorobium limicola, Chlorobium phaeobacteroides and Prosthecochloris aestuarii, for Thermomicrobium roseum, which is a relative of the green non-sulphur bacteria, and for Cytophaga aquatilis, Cytophaga heparina, Cytophaga johnsonae, Flavobacterium breve, Flexibacter sp. and Saprospira grandis, organisms allotted to the phylum 'Bacteroides-Cytophaga-Flavobacterium' and relatives as determined by 16S rRNA analyses. By using a clustering algorithm a dendrogram was constructed from these sequences and from all other known eubacterial 5S RNA sequences. The dendrogram showed differences, as well as similarities, with respect to results obtained by 16S RNA analyses. The 5S RNA sequences of green sulphur bacteria were closely related to one another, and to a cluster containing 5S RNA sequences from Bacteroides and its relatives, including Cytophaga aquatilis. 5S RNA sequences of all other representatives of the 'Bacteroides-Cytophaga-Flavobacterium' phylum as distinguished by 16S RNA analysis failed to group with Bacteroides and related clusters. On the basis of 5S RNA sequences, Thermomicrobium roseum clustered with Chloroflexus aurantiacus, as was expected from 16S RNA analysis.

Base Sequence