PubMed HealthSearch

SEARCH · PubMed Health

Results for “clustering”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family

Dependence of the rates of dissolution of the Fe4S4 clusters of Chromatium vinosum high-potential iron protein and ferredoxin on cluster oxidation state.

The influence of oxidation state on the pH dependence of the dissolution of the Fe(4)S(4) clusters of Chromatium vinosum ferredoxin and high-potential iron protein (HIPIP) has been studied. The first-order rate constants (k(obs)) for dissolution of both the Fe(4)S(4)(S-Cys)(4) (2-) and Fe(4)S(4)(S-Cys)(4) (3-) clusters of the ferredoxin follow the same overall kinetic equation but with differing specific rate and equilibrium constants. The dependence of rate and equilibrium constants upon oxidation state may be rationalized on the basis of the accompanying change in electrostatic affinity of a cluster toward H(+) and HO(-). A more drastic change in the pH dependence of the kinetics of dissolution of the Fe(4)S(4) cluster of the HIPIP accompanies its change in oxidation state. Whereas the values of k(obs) for dissolution of HIPIP containing the Fe(4)S(4)(S-Cys)(4) (2-) cluster are strictly second order to [H(+)] and [HO(-)], the pH dependence for dissolution of the HIPIP Fe(4)S(4)(S-Cys)(4) (1-) cluster indicates a first-order dependence upon [H(+)], a second-order dependence upon [HO(-)], and a spontaneous or water rate. These reactivity differences may be related to changes in cluster charge density. Mechanisms of dissolution involve preequilibrium protonation at acidic pH and preequilibrium ligand exchange at basic pH.

Chromatium

The high potential iron-sulfur cluster of aconitase is a binuclear iron-sulfur cluster.

It has been reported (Ruzicka, F.J., and Beinert, H. (1978) J. Biol. Chem. 253, 2514-2517) that aconitase in the oxidized state, as isolated, shows an electron paramagnetic resonance signal centered at g = 2.01, typical of high potential iron-sulfur proteins. Since the magnetic state corresponding to this signal has thus far only been found in tetranuclear iron-sulfur clusters in model compounds and proteins, it could be expected that aconitase also contains a [4Fe-4S] cluster. We show here that core extrusion, in the presence of hexamethylphosphoramide and o-xylyl-alpha,alpha'-dithiol and subsequent ligand exchange with p-trifluoromethylbenzenethiol yield absorption spectra typical of binuclear iron-sulfur clusters. According to the absorbance measured, the concentration of the extruded [2Fe-2S] cluster quantitatively accounts for the iron-sulfur content of the preparations examined. Preliminary studies of the 19F nuclear magnetic resonance spectrum obtained on extrusion with p-trifluoromethylbenzenethiol confirm the presence of a binuclear cluster in aconitase.

Aconitate Hydratase

Macrophage-lymphocyte clusters in the immune response to soluble protein antigen in vitro. IX. Antigen-pulsed macrophages as a tool for specific absorption of cluster-initiating T cells.

T-cell populations from guinea-pigs sensitized to the protein antigens purified protein derivative of Mycobacterium tuberculosis, ovalbumin, or horseradish perioxidase can be selectively depleted of cells capable of initiating antigen-specific macrophage-lymphocyte clusters in vitro. The depletion is achieved by incubating the T cells on a monolayer of antigen-pulsed macrophages in a Petri dish for some hours and then gently aspirating the cells not adhering to the bottom of the dish. When subsequently assayed, the aspirated cells were found to be depleted of cluster-initiating lymphocytes committed to horseradish peroxidase, monolayers of macrophages pulsed with that antigen must be used. The optimum time for incubation on the absorbing monolayer appears to be 4 h, and two successive incubations are more effective than one. The cell density of the absorbing monolayer and the handling of the Petri dish may be critical for effective removal of the cluster-initating lymphocytes. With optimum procedure we have achieved up to 90% depletion of specific cells with no depletion of cells committed to a control antigen.

Absorption

A note on cluster analysis and depression: disparities in results produced by the application of different clustering methods.

Cluster analysis is the most logically suited method for establishing psychiatric classifications. Different mathematical methods of clustering do, however, produce disparate results when applied to the same set of data. This study attempted to quantify the extent of such disparities, and found them to be marked. It was concluded that until cluster analysis has undergone further mathematical and statistical development, it should be used with caution.

Adult

[Estimation of the distance between the iron-sulfur cluster of Fe-protein and the nearest iron-sulfur cluster of Mo-Fe-protein of nitrogenase on the basis of the inductive-resonance theory of energy transfer].

The distance between fluorescein mercuric acetate (FMA), attached to the HS-group of Fe- and Mo-Fe-protein, and the nearest iron-sulphur cluster (ISC) was determined. For Fe-protein the distance was 18--20 A and for Mo-Fe-protein 12--14 A. The distance between Fe-protein FMA and the nearest Mo-protein ISC determined by complementation of the labelled Fe-protein and native Mo-Fe-protein was 14--16 A. The distance between MO-OFe-protein ISC and complement Fe-protein ISC was 18--20 A. A te-protein ISC permitted to suppose that the electron was transfered from Fe-protein ISC to Mo-Fe-protein ISC by the contact of the ISC or with the help of ATP molecule.

Binding Sites

Phylogenetic inconsistency of pairwise SNP clustering for inferring tuberculosis transmission in a high-burden, endemic setting: a case study from Thailand.

Whole-genome sequence analysis is now widely used to delineate tuberculosis transmission clusters. A standard practice is to cluster bacterial isolates based on a fixed maximum genome-wide pairwise single nucleotide polymorphism (pwSNP) distance threshold. In this study, we evaluated the phylogenetic consistency of pwSNP-distance clustering with thresholds ranging between 1 and 25 single nucleotide polymorphisms (SNPs) using two contrasting data sets: (i) a data set from the UK (N = 390) published by T. M. Walker, C. L. C. Ip, R. H. Harrell, J. T. Evans, et al. (Lancet Infect Dis 13:137-146, 2013, https://doi.org/10.1016/S1473-3099(12)70277-3), which was foundational to the establishment of this method, and (ii) a data set from Thailand (N = 3,341), characterized by persistent transmission and sparse, non-systematic sampling. For the UK data set, the standard pwSNP-distance clustering using thresholds of &#x2265;12 SNPs yielded entirely monophyletic clusters and showed high concordance with a comparative monophyly constrained, tree-based method. In contrast, for the Thai data set, pwSNP-distance clustering often generated non-monophyletic clusters, even by the 25-SNP threshold. The pwSNP-distance and comparative tree-based clustering methods only showed large consistency at thresholds of &#x2265;22 SNPs. This suggests that SNP clusters defined by low distance thresholds (i.e., <12 SNPs for the UK data set, and <22 SNPs for the Thai data set) may lack robustness, and the problem is particularly severe for data sets characterized by persistent transmission, likely due to poorer cluster separation. Moreover, our findings indicate that large cluster sizes, high maximum intra-cluster genetic distances, and broad sample collection time spans may serve as useful indicators of potentially non-monophyletic clusters. We also demonstrate that mixed infections can produce spurious, phylogenetically long-range SNP linkages, underscoring the necessity of strict sequence quality control.IMPORTANCEFixed-threshold pairwise single nucleotide polymorphism (pwSNP)-distance clustering is commonly used to delineate tuberculosis transmission clusters. From an epidemiological perspective, a genuine transmission cluster must be monophyletic, originating from a single source. However, pwSNP-distance clustering is inherently simplistic and can therefore violate this principle, making the assessment of its phylogenetic consistency critical. Our results demonstrate that while this method effectively delineated complete transmission clusters for the data set from the UK, a low-burden and non-persistent transmission setting, it frequently generated non-monophyletic clusters when applied to the Thai data set, characterized by persistent transmission alongside sparse and non-systematic sampling. Furthermore, we found that clusters derived using low distance thresholds could notably vary between the pwSNP-distance and comparative tree-based clustering methods, suggesting limited reliability and robustness. To accurately delineate tuberculosis transmission clusters, especially for complex data from high-burden, endemic settings, we recommend transitioning from pwSNP-distance clustering toward more robust, phylogenetic clustering that respects evolutionary descent.

Mycobacterium tuberculosis

Polycystic Ovary Syndrome Physiologic Pathways Implicated Through Clustering of Genetic Loci.

CONTEXT: Polycystic ovary syndrome (PCOS) is a heterogeneous disorder, with disease loci identified from genome-wide association studies (GWAS) having largely unknown relationships to disease pathogenesis. OBJECTIVE: This work aimed to group PCOS GWAS loci into genetic clusters associated with disease pathophysiology. METHODS: Cluster analysis was performed for 60 PCOS-associated genetic variants and 49 traits using GWAS summary statistics. Cluster-specific PCOS partitioned polygenic scores (pPS) were generated and tested for association with clinical phenotypes in the Mass General Brigham Biobank (MGBB, N = 62 252). Associations with clinical outcomes (type 2 diabetes [T2D], coronary artery disease [CAD], and female reproductive traits) were assessed using both GWAS-based pPS (DIAMANTE, N = 898,130, CARDIOGRAM/UKBB, N = 547 261) and individual-level pPS in MGBB. RESULTS: Four PCOS genetic clusters were identified with top loci indicated as following: (i) cluster 1/obesity/insulin resistance (FTO); (ii) cluster 2/hormonal/menstrual cycle changes (FSHB); (iii) cluster 3/blood markers/inflammation (ATXN2/SH2B3); (iv) cluster 4/metabolic changes (MAF, SLC38A11). Cluster pPS were associated with distinct clinical traits: Cluster 1 with increased body mass index (P = 6.6 &#xd7; 10-29); cluster 2 with increased age of menarche (P = 1.5 &#xd7; 10-4); cluster 3 with multiple decreased blood markers, including mean platelet volume (P = 3.1 &#xd7;10-5); and cluster 4 with increased alkaline phosphatase (P = .007). PCOS genetic clusters GWAS-pPSs were also associated with disease outcomes: cluster 1 pPS with increased T2D (odds ratio [OR] 1.07; P = 7.3 &#xd7; 10-50), with replication in MGBB all participants (OR 1.09, P = 2.7 &#xd7; 10-7) and females only (OR 1.11, 4.8 &#xd7; 10-5). CONCLUSION: Distinct genetic backgrounds in individuals with PCOS may underlie clinical heterogeneity and disease outcomes.

Humans

High and low reduction potential 4Fe-4S clusters in Azotobacter vinelandii (4Fe-4S) 2ferredoxin I. Influence of the polypeptide on the reduction potentials.

Azotobacter vinelandii (4Fe-4S)2 ferredoxin I (Fd I) is an electron transfer protein with Mr equals 14,500 and Eo equals -420 mv. It exhibits and EPR signal of g equals 2.01 in its isolated form. This resonance is almost identical with the signal that originates from a "super-oxidized" state of the 4Fe-4S cluster of potassium ferricyanide-treated Clostridium ferredoxin. A cluster that exhibits this EPR signal at g equals 2.01 is in the same formal oxidation state as the cluster in oxidized Chromatium High-Potential-Iron-Protein (HiPIP). On photoreduction of Fd I with spinach chloroplast fragments, the resonance at g equals 2.01 vanishes and no EPR signal is observed. This EPR behavior is analogous to that of reduced HiPIP, which also fails to exhibit an EPR spectrum. These characteristics suggest that a cluster in A. vinelandii Fd I functions between the same pair of states on reduction as does the cluster in HiPIP, but with a midpoint reduction potential of -420 mv in contrast to the value of +350 mv characteristic of HiPIP. Quantitative EPR and stoichoimetry studies showed that only one 4Fe-4S cluster in this (4Fe-4S)2 ferredoxin is reduced. Oxidation of Fd I with potassium ferricyanide results in the uptake of 1 electron/mol as determined by quantitative EPR spectroscopy. This indicates that a cluster in Fd I shows no electron paramagnetic resonance in the isolated form of the protein accepts an electron on oxidation, as indicated by the EPR spectrum, and becomes paramagnetic. The EPR behavior of this oxidizable cluster indicates that it also functions between the same pair of oxidation states as does the Fe-S cluster in HiPIP. The midpoint reduction potential of this cluster is approximately +340 mv. A. vinelandii Fd I is the first example of an iron-sulfur protein which contains both a high potential cluster (approximately +340 mv) and a low potential cluster (-420 mv). Both Fe-S clusters appear to function between the same pair of oxidation states as the single Fe-S cluster in Chromatium HiPIP, although the midpoint reduction potentials of the two clusters are approximately 760 mv different.

Azotobacter

Application of cluster analysis for characterization of spatial distribution of particles by stereological methods.

A method for the detection and characterization of clusters of particles observed in section with the electron microscope is presented. Cluster analysis is performed by the division method described by Berthet et al. (1976). Starting from a single cluster, profiles from each electron micrograph are successively classified in sets containing an increasing number of clusters. The decrease in the mean free distance, lambda, between profiles in the clusters, is used for terminating the subdivision procedure. The function relating the mean free distance with the number of clusters is evaluated in each subdivision set. The actual number of clusters is selected on the basis of the slope of that function, at a point where lambda has a value close to the average profile diameter. The method assumes a convex shape for the clusters; the salient feature is that it provides a physical delineation of clusters in the section. Hence, an evaluation of some characteristics of clusters in the three-dimensional sample may be obtained by using standard stereological procedures. Characterization of the volume to which the individual particles of a population are eventually restricted can as a result be performed. Practical problems in the acquisition of the data needed for cluster analysis are discussed and a system using for that purpose a Quantimet 720 image analyser in a basic configuration, connected on line with a PDP 11/10 minicomputer, is presented. Application of the method is illustrated by the analysis of lysosomes in cultured hepatoma (HTC) cells, at the end of mitosis and during the S phase. Cluster analysis shows that in cells actively synthesizing DNA they are grouped in clusters representing 5.7% of the cellular volume. Moreover, the average number of particles per cluster falls from a minimum of thirteen at mitosis to only six at the S phase.

Cells, Cultured

Clustering patterns of behavioral and metabolic risk factors for noncommunicable diseases in Iran: findings from a national STEPS survey.

BACKGROUND: Noncommunicable diseases (NCDs) are the leading cause of mortality in Iran, driven by behavioral and metabolic risk factors that frequently co-occur. OBJECTIVE: To identify patterns of co-occurring behavioral and metabolic NCD risk factors among Iranian adults and characterize their demographic and socioeconomic correlates. METHODS: This cross-sectional study analyzed data from 16,618 adults aged &#x2265;25&#x2009;years who participated in Iran's 2021 nationally representative STEPS survey. Thirteen behavioral and metabolic variables, including physical activity, nutrition score, smoking frequency, alcohol intake, salt intake, body mass index, blood pressure, fasting plasma glucose, and lipid markers, were entered into a K-means clustering analysis. Clusters were characterized by their risk profiles and demographic/socioeconomic attributes. Multinomial logistic regression examined associations between cluster membership and sociodemographic factors. RESULTS: Five distinct behavioral-metabolic clusters emerged. The smokers-drinkers (SD) cluster (3.1%) comprised mostly older, less-educated men with high smoking and alcohol use. The healthy-low-risk (HLR) cluster (40.3%) showed favorable profiles and included younger, more educated individuals. The physically active (PA) cluster (6.6%) was characterized mainly by younger men with markedly high physical activity levels. The dyslipidemic (DLP) cluster (26.0%) exhibited high dyslipidemia and overweight prevalence, while the hypertensive-diabetic (HTD) cluster (24.0%) had the highest obesity, hypertension, and diabetes rates, common among older urban adults. CONCLUSION: Behavioral and metabolic NCD risk factors in Iran formed five distinct co-occurrence patterns. Nearly half of adults belonged to metabolically high-risk clusters, highlighting the need for targeted prevention strategies that combine lifestyle interventions with screening and management of obesity, hypertension, diabetes, and dyslipidemia.

Humans

Lineage-specific transmission and spatial clustering of Mycobacterium tuberculosis in Kaohsiung, Taiwan, in 2019-23: a population-based genomic study.

BACKGROUND: The epidemiology of tuberculosis in Taiwan has been influenced by the introduction of multiple Mycobacterium tuberculosis lineages and by the ageing of the population. We conducted a population-based study to investigate M tuberculosis transmission in Kaohsiung, a city in southern Taiwan. METHODS: In this study, we performed whole-genome sequencing (WGS) of M tuberculosis isolates from all culture-positive cases of tuberculosis notified in Kaohsiung between Jan 1, 2019 and Dec 31, 2023. We obtained routine epidemiological data for each case collected through the national tuberculosis control programme. We characterised the lineage composition of the isolate collection and evaluated genomic clustering of isolates, defined as a difference of 12 or fewer single-nucleotide polymorphisms. Univariable and multivariable logistic regression analyses were performed to estimate the odds of a case belonging to a genomic cluster based on host factors (age, sex, sputum smear status, and residential region) and pathogen factors (drug resistance status and strain lineage). Spatial aggregation of large genomic clusters (including greater than or equal to ten isolates) was assessed using a non-parametric statistical clustering method. We used a Bayesian transmission tree inference method to explore the patterns of age-dependent transmission. FINDINGS: During the study period, 5667 tuberculosis cases were notified in Kaohsiung, 4916 (86&#xb7;7%) of which were culture-positive. Of these 4916 cases, whole-genome sequencing was successfully performed for 4168 (84&#xb7;8%) isolates. 1219 (29&#xb7;2%) of 4168 individuals were female and 2947 (70&#xb7;7%) were male; the median age was 69&#xb7;7 years (IQR 57&#xb7;4-80&#xb7;7). The dominant lineages were lineage 1 (1749 [42&#xb7;0%] of 4168 isolates), lineage 2 (1510 [36&#xb7;2%]), and lineage 4 (905 [21&#xb7;7%]). 1069 (25&#xb7;6%) of 4168 were genomically linked and formed 287 clusters. Lineage 2 isolates had higher odds (aOR 2&#xb7;15 [95% CI 1&#xb7;80-2&#xb7;52]) than lineage 1 isolates of genomic clustering across all regions, whereas lineage 4 isolates had a significantly higher risk (2&#xb7;75 [1&#xb7;16-6&#xb7;89]) of genomic clustering than lineage 1 only in the rural northeast region, inhabited primarily by indigenous populations. Spatial clustering analysis corroborated these lineage-region interactions. Although younger adults (<35 years) had the highest individual-level odds (5&#xb7;64 [4&#xb7;16-7&#xb7;68]) of clustering in the logistic regression analysis compared with those aged 80 years or older, the transmission inference indicated that individuals aged 55-74 years were responsible for a greater proportion of inferred transmission events, contributing 50&#xb7;8% of all transmission events. INTERPRETATION: This sequencing study revealed that older adults (aged &#x2265;65 years) might have played a substantial and under-recognised role in the transmission of tuberculosis in Taiwan. The lineage-specific clustering and spatial patterns suggested that both pathogen characteristics and host demographics shaped tuberculosis transmission dynamics. These findings support the use of integrated genomic surveillance to guide precision tuberculosis control and motivate further research on age-specific transmission pathways and targeted interventions to advance tuberculosis elimination efforts. FUNDING: Taiwan National Health Research Institutes and Taiwan National Science and Technology Council.

Mycobacterium tuberculosis

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256&#x2009;Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms