PubMed HealthSearch

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Multivariate procedures to describe clinical staging of melanoma.

Analyzing multivariate clinical data to identify subclasses of patients being treated for a specific disease may improve patient management and increase understanding of the behavior of disease under clinical conditions. In some cases, patients have been classified on prognostic characteristics using standard risk assessment procedures (e.g., Cox' regression). This requires long term follow-up, differentiates patients only on attributes relevant to survival, and assumes that patients are sampled from a common population. Other approaches involve the use of clustering algorithms to classify patients into categories based on multiple clinical attributes. We illustrate the use of a multivariate statistical procedure to directly characterize patients on multiple clinical characteristics. The procedure is designed to analyze discrete response data with parameters representing individual differences within groups. Its use is illustrated for patients with Stage I melanoma in determining how age is related to treatment response in different patient groups.

Adult

From biopsy to automatic diagnosis.

High resolution two-dimensional gel electrophoresis is a very powerful biochemical tool for analysis of complex protein mixtures. In well defined situations, protein maps, obtained from tissue biopsies or biological fluids by this technique, can be automatically analyzed by computer. Some polypeptide patterns are the fingerprints of diseases. Applying clustering algorithm and learning techniques, the prototype expert system MELANIE recognized patterns and associated the correct diagnosis to the specific pattern.

Diagnosis, Computer-Assisted

[Phylogenetic analysis of partial nucleotide sequences of 18S rRNA for 14 plant species].

The variable 260 base long region from the interior of 18S rRNA of 14 plant species was determined by chain termination method with the use of reverse transcriptase. The hairpin revealed in this region appeared to be conservative in all species compared. Thermodynamic stability of such hairpin is lower than of an alternative structure with different base pairing mode. From sequence data dendrograms were produced by clustering algorithms and by the compatibility method. In addition to the plant sequences these dendrograms included also the homologous regions from yeast and Xenopus 18S rRNAs. The compatibility method seems to be more reliable. Inferences were drawn on relations between gymnosperms and angiosperms, monocots and dicots on the bases of the analysis of this tree.

Base Sequence

Quantification of progressive diabetic macular nonperfusion.

We used the IS-2000 Image Analyzer to estimate the extent of progressive diabetic macular nonperfusion in a patient by means of an automatic clustering algorithm applied to digitized fluorescein angiograms of the patient's macula taken over time. This method may provide an objective and reproducible quantification of progressive macular nonperfusion.

Adult

A detection algorithm for multiform premature ventricular contractions.

This paper reports an algorithm developed to identify and quantify multiform PVCs. The algorithm clusters PVCs of similar morphology using a combination of time-domain and frequency-domain analysis. Initially, PVCs are grouped together on the basis of four time-domain-based morphological feature measurements. However, these time-domain-based clusters many times are nonunique because commonly encountered signal changes can cause substantial variations in the feature measurements of clinically similar beats. These redundant clusters are consolidated using two frequency-domain parameters: The First Spectral Moment (FSM) (center of gravity) of the amplitude spectrum, and the 5-Hz phase angle.

Cardiac Complexes, Premature

Multiple sequence alignment with hierarchical clustering.

An algorithm is presented for the multiple alignment of sequences, either proteins or nucleic acids, that is both accurate and easy to use on microcomputers. The approach is based on the conventional dynamic-programming method of pairwise alignment. Initially, a hierarchical clustering of the sequences is performed using the matrix of the pairwise alignment scores. The closest sequences are aligned creating groups of aligned sequences. Then close groups are aligned until all sequences are aligned in one group. The pairwise alignments included in the multiple alignment form a new matrix that is used to produce a hierarchical clustering. If it is different from the first one, iteration of the process can be performed. The method is illustrated by an example: a global alignment of 39 sequences of cytochrome c.

Algorithms

Clustering cDNA sequences.

A set of programs has been written to quantify the similarities between large numbers of cDNA sequences. This information is used to cluster similar sequences together. The main program can cluster thousands of cDNA sequences per day using a novel, computationally inexpensive algorithm. The clustering information is kept in a small index file so that disk storage requirements are negligible. Using this index file, subsidiary programs create various views and statistical summaries of the entire cDNA sequence collection.

Algorithms

Algorithm for the detection of fine clustered calcifications on film mammograms.

An algorithmic process for the detection and marking of clustered calcifications in digitized film-screen mammograms has been applied to mammograms from 50 clinical cases sampled at two digitization levels, in both the craniocaudal and mediolateral views. In all but one case the detector accurately located suggestive clusters found by radiologists in normal screening. In five cases additional clusters were also found by the detector. The detector has a negligible false-positive rate for the detection of clustered calcifications, although it is sensitive to clusters of emulsion defects displayed as artifactual calcification densities in the original film. The detector is flexible in structure and is easily adapted to various calcification/cluster criteria. The detector shows considerable promise when applied to clinical examples but will require refinement before formal testing.

Algorithms

Atlas-level single-cell integration and clustering-free differential expression analysis with GEDI 2.0.

MOTIVATION: GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS: We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98× mean single-threaded speedups, and up to 11.5× acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.

Single-Cell Analysis

Automatic identification of significant graphoelements in multichannel EEG recordings by adaptive segmentation and fuzzy clustering.

A new approach to visual evaluation of long-term EEG recordings is proposed. The method is based on multichannel adaptive segmentation, subsequent feature extraction, automatic classification of the acquired segments by fuzzy cluster analysis (fuzzy c-means algorithm), and on the distinguishing of thus identified EEG segments by colour directly in the EEG record. The black and white variant of the described automatic system is presented. The method was evaluated by applying it to simulated artificial data and to real EEG recordings; some of the illustrative results are shown. In addition, the performance of this system is evaluated and the first experience with its application to routine EEG recordings is discussed.

Algorithms

Cluster analysis and related techniques in medical research.

In this paper we review methods of cluster analysis in the context of classifying patients on the basis of clinical and/or laboratory type observations. Both hierarchical and non-hierarchical methods of clustering are considered, although the emphasis is on the latter type, with particular attention devoted to the mixture likelihood-based approach. For the purposes of dividing a given data set into g clusters, this approach fits a mixture model of g components, using the method of maximum likelihood. It thus provides a sound statistical basis for clustering. The important but difficult question of how many clusters are there in the data can be addressed within the framework of standard statistical theory, although theoretical and computational difficulties still remain. Two case studies, involving the cluster analysis of some haemophilia and diabetes data respectively, are reported to demonstrate the mixture likelihood-based approach to clustering.

Algorithms

Exploring cross-category relationships between symptoms in people with hypermobile EDS (hEDS) to identify disability patterns.

BACKGROUND: Hypermobile Ehlers-Danlos Syndrome (hEDS) is a connective tissue disorder with variable symptom presentation across multiple organ systems and significant morbidity. Little is known about hEDS etiology and identifying patterns of symptom co-occurrence can reveal previously unidentified relationships between phenotypes and inform studies of underlying disease pathophysiology for symptoms that may share functional biological pathways. In this exploratory analysis, we specifically assessed the distribution of symptoms in case and controls to identify clusters of co-occurring symptoms. METHODS: We have interrogated clinically relevant symptom areas in 47 females with hEDS, 36 age-matched female controls and 8 hypermobile patients without chronic pain. Studied symptoms include general health, mental health, body pain, vitality and energy, autonomic symptoms, bleeding, and gastrointestinal symptoms. We conducted hierarchal clustering on principle components (HCPC) to identify groups and compared the groups for the previously described symptoms. Radial plots were used to identify relationships between severe symptom categories. RESULTS: Our analysis reveals statistically significantly more severe symptoms in all categories in people with hEDS compared with age- and sex-matched controls and asymptomatic hypermobile patients. HCPC identified clearly separated Low, Moderate, and High symptom groups within participants. The Low dysfunction groups include nearly all controls and hypermobile patients without chronic pain. The High dysfunction group includes ~60% of people with hEDS, while around 40% are in the Moderate dysfunction cluster. Cluster solutions for all participants were stable with moderate fit (silhouette 0.64; Jaccard boot mean 0.91). Group level radial plots showed high bleeding severity across all symptom clusters, while disproportional severity of general health, physical function, limitation of role due to physical symptoms, pain, and social functioning deficits differentiates the High from Moderate and Low Dysfunction clusters. CONCLUSION: Using this analysis at the group level has revealed patterns suggesting a progression of disease symptoms. People with hypermobility do not uniformly have severe symptoms but instead have some symptoms that differentiate from non-hypermobile individuals. While exploratory, using a radar multi-symptom analysis may be used to evaluate disproportionately severe symptoms contributing to the patterns of global symptom severity. These include pain but also ability to perform roles, suggesting strong utility of physical and occupational therapies to emphasize coping. This may also allow better targeting of etiological studies and may have additional utility at an individual level to develop symptom management strategies.

Humans

Clustering patterns of behavioral and metabolic risk factors for noncommunicable diseases in Iran: findings from a national STEPS survey.

BACKGROUND: Noncommunicable diseases (NCDs) are the leading cause of mortality in Iran, driven by behavioral and metabolic risk factors that frequently co-occur. OBJECTIVE: To identify patterns of co-occurring behavioral and metabolic NCD risk factors among Iranian adults and characterize their demographic and socioeconomic correlates. METHODS: This cross-sectional study analyzed data from 16,618 adults aged ≥25 years who participated in Iran's 2021 nationally representative STEPS survey. Thirteen behavioral and metabolic variables, including physical activity, nutrition score, smoking frequency, alcohol intake, salt intake, body mass index, blood pressure, fasting plasma glucose, and lipid markers, were entered into a K-means clustering analysis. Clusters were characterized by their risk profiles and demographic/socioeconomic attributes. Multinomial logistic regression examined associations between cluster membership and sociodemographic factors. RESULTS: Five distinct behavioral-metabolic clusters emerged. The smokers-drinkers (SD) cluster (3.1%) comprised mostly older, less-educated men with high smoking and alcohol use. The healthy-low-risk (HLR) cluster (40.3%) showed favorable profiles and included younger, more educated individuals. The physically active (PA) cluster (6.6%) was characterized mainly by younger men with markedly high physical activity levels. The dyslipidemic (DLP) cluster (26.0%) exhibited high dyslipidemia and overweight prevalence, while the hypertensive-diabetic (HTD) cluster (24.0%) had the highest obesity, hypertension, and diabetes rates, common among older urban adults. CONCLUSION: Behavioral and metabolic NCD risk factors in Iran formed five distinct co-occurrence patterns. Nearly half of adults belonged to metabolically high-risk clusters, highlighting the need for targeted prevention strategies that combine lifestyle interventions with screening and management of obesity, hypertension, diabetes, and dyslipidemia.

Humans

Identification of homogeneous geographical areas of mortality for tumours from cluster analysis.

This paper attempts to demonstrate the utility of cluster analysis as a descriptive method of studying mortality in epidemiology. In order to verify which algorithms of clustering best fit the data structure, the method of cophenetic correlation was implemented. Furthermore the probabilistic algorithm proposed by Beale was used to assess the partition. The results show the presence of some striking clusters between Local Sanitary Units of the Emilia Romagna Region for four types of tumour in men.

Algorithms

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n = 38, 74%). Hierarchical clustering (n = 20) and K-means clustering (n = 14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.

Spatial omics technologies have revolutionized the study of tissue architecture and cellular heterogeneity by integrating molecular profiles with spatial localization. In spatially resolved transcriptomics, delineating higher-order anatomical structures is critical for understanding how cellular organization affects function. However, the reliability of current benchmarks of spatially aware clustering (SAC) methods is undermined by their narrow focus on Visium and brain tissue datasets and the incorrect interpretation of manual annotation as ground truth. Here we present SACCELERATOR, a community-driven, extensible framework that standardizes data formatting, method integration and metric evaluation, enabling rapid inclusion of new methods and datasets. Our analysis revealed substantial limitations in the generalizability and reproducibility of SAC methods and shows that anatomical labels commonly used as ground truths are often biased, error prone and unsuitable for benchmarking. Rather than ranking methods, we propose a consensus-guided workflow where descriptive spatial metrics highlight high-entropy regions of method disagreement, enabling targeted feedback for tissue experts. Applied to brain and cancer datasets, this approach uncovered biologically meaningful patterns overlooked by individual SAC methods and manual annotations, highlighting the need for iterative, expert-in-the-loop evaluation.

Benchmarking

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms