PubMed HealthSearch

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

CeLLTra: aligning cell names with gene expression via a pathway-informed transformer.

MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

Association of MPO Expression with the Immune Microenvironment in Breast Cancer: Insights from Bioinformatics and Single-Cell Analyses.

Breast cancer remains a major cause of cancer-related mortality, and exploratory computational workflows can help prioritize immune-associated markers for further investigation. Here, we used the cancer genome atlas breast invasive carcinoma (TCGA-BRCA) bulk transcriptomic data and the public single-cell dataset GSE161529 to examine associations between myeloperoxidase (MPO) expression, clinical outcomes, immune infiltration, methylation, upstream-regulator annotations, single-cell expression patterns, virtual knockdown sensitivity outputs, drug-gene interaction retrieval, and absorption, distribution, metabolism, excretion, and toxicity (ADMET) annotation. MPO expression was lower in breast cancer tissues than in adjacent non-tumor tissues. Higher MPO expression was associated with a longer progression-free interval, whereas its associations with overall survival and disease-specific survival were not statistically significant. Receiver operating characteristic (ROC) analysis suggested tumor-normal separation within the analyzed public dataset, but this should not be interpreted as clinical diagnostic validation. Immune deconvolution and enrichment analyses indicated that MPO expression mainly tracked with immune- and myeloid-related transcriptional features, rather than establishing tumor-intrinsic regulation of the immune microenvironment. At single-cell resolution, the MPO signal was sparse, with only 85 MPO-positive cells detected before k-nearest neighbor (KNN)-based neighborhood expansion. Detectable MPO signal and MPO-associated scores were interpreted cautiously because they may be influenced by sparse expression, cell-type annotation uncertainty, dropout, doublets, or ambient RNA. In silico virtual knockdown suggested candidate immune- and inflammatory-related transcriptional changes, but these results were considered exploratory and require validation. Drug-gene interaction database (DGIdb)-based drug-gene retrieval and ADMET annotation were used only as preliminary chemical annotations and were not interpreted as therapeutic evidence. Overall, this study provides a reproducible in silico workflow for generating hypotheses about MPO-associated immune/myeloid features in breast cancer, which require external cohort validation and experimental confirmation.

Humans

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 × 50 DNA nanoballs and an approximate nominal footprint of 25 × 25 µm, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

Differential cell signaling testing for cell-cell communication inference from single-cell data by dominoSignal.

MOTIVATION: Algorithms for ligand-receptor network inference have emerged as commonly used tools to estimate cell-cell communication from reference single-cell data. Many studies employ these algorithms to compare signaling between conditions and lack methods to statistically identify signals that are significantly different. We previously developed the cell communication inference algorithm Domino, which considers ligand and receptor gene expression in association with downstream transcription factor activity scoring. We developed the dominoSignal software to innovate upon Domino and extend its functionality to test statistically differential cellular signaling. RESULTS: This new functionality includes the compilation of active signals as linkages from multiple subjects in a single-cell data set and testing condition-dependent signaling linkage. The software is applicable for analysis of single-cell data sets with multiple subjects as biological replicates as well as with bootstrapped replicates from data sets with few or pooled subjects. We use simulation studies to benchmark the number of subjects in compared groups and cells within an annotated cell type sufficient to accurately identify differential linkages. We demonstrate the application of the Differential Cell Signaling Test (DCST) in the dominoSignal software to investigate consequences of cancer cell phenotypes and immunotherapy on cell-cell communication in tumor microenvironments. These applications in cancer studies demonstrate the ability of differential cell signaling analysis to infer changes to cell communication networks from therapeutic or experimental perturbations, which is broadly applicable across biological systems. AVAILABILITY: dominoSignal is available through Bioconductor at https://www.bioconductor.org/packages/release/bioc/html/dominoSignal.html.

Cell Communication

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

dbscATAC: a resource of single-cell super-enhancers/enhancers and gene markers derived from scATAC-seq data.

MOTIVATION: scATAC-seq enables high-resolution mapping of cis-regulatory elements. It has been widely applied to uncover cell-type-specific regulatory networks and complement scRNA-seq analysis in numerous studies. However, a large number of datasets generated by scATAC-seq remain underutilized due to limited exploration of super-enhancers/typical enhancers and gene markers. A comprehensive resource enabling cell-type-specific annotation of cis-regulatory elements and their dynamic enhancer-gene linkages remains an urgent unmet need for scATAC-seq. RESULTS: We present dbscATAC, a specialized single-cell database for annotating super-enhancers, gene markers, and enhancer-gene interactions derived from scATAC-seq data. Using improved machine learning algorithms, we identified 213 835 super-enhancers across 520 tissue/cell types from three species, as well as 347 484 gene markers, 13 470 526 enhancers, and 10 402 346 enhancer-gene interactions derived from 1 668 076 single cells spanning 1028 tissue/cell types in 13 species. An easy-to-use online platform with multiple analytic modules and hierarchical query options was developed for searching, browsing and visualizing single-cell super-enhancers, enhancers, and gene markers. dbscATAC provides a comprehensive resource to facilitate the exploration of enhancer landscapes, gene regulation, and cell-type-specific characteristics in single-cell epigenomics. AVAILABILITY AND IMPLEMENTATION: The database with all the super-enhancer/enhancer annotation data is available at http://singlecelldb.com/dbscATAC/index.php. And the source code of dbscATAC for prediction of SEs, enhancers, and gene markers are available at https://github.com/EvansGao/dbscATAC. The source code, tissue/cell type description, and data summary can be downloaded at DOI: 10.6084/m9.figshare.28706414.scATAC-seq, Database, Super-enhancers/enhancers, Gene markers.

Enhancer Elements, Genetic

Coordinated inflammatory macrophage and vascular smooth muscle cell remodeling signatures in human atherosclerosis: An integrative single-cell and bulk transcriptomic analysis.

Atherosclerotic plaque progression is shaped by coordinated inflammatory and remodeling programs involving immune cells and vascular wall cells. Inflammatory macrophage activation and vascular smooth muscle cell (VSMC) phenotypic remodeling are central features of human atherosclerosis, but their transcriptomic relationships during plaque progression remain incompletely characterized. This study integrated single-cell and bulk transcriptomic datasets to examine highly inflammatory macrophage states, VSMC remodeling-related transcriptional programs, and candidate ligand-receptor expression patterns in human atherosclerotic plaques. Human atherosclerotic plaque single-cell RNA sequencing data from GSE260657 and bulk transcriptomic data from GSE28829 were analyzed. After quality control, 7628 cells were retained for single-cell analysis. Major cell types were annotated using canonical markers, followed by reclustering of macrophages and VSMC-related cells. Functional module scoring, differential expression analysis, Gene Ontology biological process enrichment, and Kyoto Encyclopedia of Genes and Genomes pathway analyses were performed to characterize macrophage transcriptional states. Slingshot was applied to infer VSMC pseudotime ordering. CellChat and NicheNet were used to prioritize candidate ligand-receptor expression patterns and ligand-associated VSMC target gene programs. External bulk transcriptomic analysis was performed to examine whether single-cell-derived inflammatory and remodeling signatures were represented at the tissue-transcriptome level during plaque progression. Macrophage reclustering identified a highly inflammatory macrophage state characterized by prominent inflammatory activation, cytokine-response, and stress-response features. Genes upregulated in this population were enriched in pathways related to tumor necrosis factor (TNF) response, nuclear factor kappa B signaling, leukocyte activation, cytokine signaling, lipid and atherosclerosis, toll-like receptor signaling, and inflammasome-associated inflammation. VSMC reclustering revealed contractile VSMCs, PTHLH+ synthetic VSMCs, KRT7+ VSMC-like cells, interferon-responsive VSMCs, pericyte-like mural cells, and osteogenic/modulated VSMCs. Pseudotime analysis showed a broad contractile-to-osteogenic/modulated transcriptional continuum accompanied by increased expression of remodeling-associated genes and selected inflammatory or remodeling-associated receptor genes. CellChat and NicheNet analyses prioritized candidate ligand-receptor and ligand-associated target gene expression patterns involving SPP1-CD44, TNF-TNFRSF1A, IL1B-IL1R1/IL1RAP, MIF-ACKR3, PDGFB-PDGFRB, and FN1-SDC1/ITGB1. In GSE28829, inflammatory macrophage-, osteogenic/modulated VSMC-, candidate ligand-receptor expression-, SPP1-CD44 candidate axis-, and NicheNet-prioritized target program-related signatures were more prominent in advanced plaques and were positively correlated with each other. This integrative transcriptomic analysis identified a highly inflammatory macrophage state and a VSMC remodeling continuum in human atherosclerotic plaques. Candidate ligand-receptor and ligand-associated target gene expression patterns linked inflammatory macrophage activation with osteogenic/modulated VSMC remodeling at the computational level. External bulk data further showed coordinated enrichment of inflammatory and remodeling signatures in advanced plaques. These findings provide a descriptive and hypothesis-generating transcriptomic framework for understanding inflammatory macrophage activation and VSMC remodeling in human atherosclerosis.

atherosclerosis

Genomic dimensions deconstruct the clinical heterogeneity of bipolar disorder.

Bipolar disorder's (BD) clinical heterogeneity has an unresolved genetic basis. We meta-analyzed genome-wide association studies (GWAS) of 16 BD subphenotypes in 226,032 individuals from 57 cohorts (38,022 cases); 10 advanced to multivariate and multi-trait analyses. Four factors (compulsive, psychotic, dysregulated, internalizing) explained 82.8% of shared genetic variance. BD1 and BD2 loaded on distinct factors despite a high genetic correlation; 87.0% of common-factor loci were significant in neither subtype. Unipolar mania aligned with psychosis over internalizing, and was distinguishable from BD1, and rapid cycling showed heritable cross-domain liability. We identified 356 risk loci, 158 novel, including the first univariate-GWAS associations for psychosis, unipolar mania, rapid cycling and schizoaffective disorder-and 249 credible genes (89 high-confidence), 12 with approved-drug or clinical-phase annotations. Cell-type association showed a midbrain dopaminergic-GABAergic gradient along the psychotic factor. BD's genetic architecture appears hierarchical-a general liability resolving into dimensions of course and comorbidity, beyond subtypes.

Journal Article

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

Integrated multi-omics profiling identifies aging-related molecular signatures and convergent interferon signaling in systemic lupus erythematosus.

BACKGROUND: Systemic lupus erythematosus (SLE) is characterized by chronic immune activation and molecular alterations that overlap with aging-related biological processes. However, how these alterations are organized across molecular layers and whether they converge on shared regulatory networks remain incompletely understood. METHODS: We performed an integrative multi-omics analysis combining in-house proteomic and phosphoproteomic data from 130 patients with SLE and 90 healthy controls (HCs) and publicly available transcriptomic datasets comprising 1,461 SLE patients. Proteins and phosphorylation sites were annotated using established aging-related gene resources. Differential protein abundance and phosphorylation changes were analyzed across disease-status and disease-activity comparisons. Nominal P-value thresholds were used for exploratory feature selection, whereas FDR-adjusted P values were used to assess robustness after multiple-testing correction. Kinase-substrate enrichment, transcription factor annotation, and cell-type-resolved transcriptomic comparison were used to explore potential regulatory programs. RESULTS: We identified 128 nominally altered proteins annotated to aging-related biological processes, including genomic instability, mitochondrial dysfunction, and epigenetic alterations. Phosphoproteomic analysis revealed 36 nominally altered phosphorylation sites, including previously unreported sites in IFI16 (S153, S780) and PKCδ (S507, S664). Clustering analysis demonstrated heterogeneous protein co-regulation patterns across disease states. Kinase activity inference suggested altered activity of TBK1 and IKKβ. TF analysis further highlighted STAT1, RELA, and PML as potential central nodes within the inferred regulatory network. Notably, these multi-omic alterations were not randomly distributed but showed convergence toward shared signaling pathways, particularly those related to interferon responses. CONCLUSIONS: This integrative multi-omics study identifies inflammatory and interferon-dominated molecular alterations in SLE PBMCs that overlap with aging-related biological processes and converge on shared regulatory networks. These findings provide a hypothesis-generating framework for investigating the intersection between chronic immune activation and aging-related molecular remodeling in SLE.

Humans

Putative function and prognostic molecular marker of mast cells in colorectal cancer.

BACKGROUND: The increased demand for markers for colorectal cancer (CRC) highlights the importance of investigating immune cells involved in CRC progression. This study aims to dissect the mast cells in CRC, characterize the role of mast cells in CRC development, coordinate molecular communication between mast cells and malignant cells, and construct and validate a prognostic classification model based on mast cell markers. METHODS: Single-cell transcriptome data of CRC patients were extracted from GSE146771 for cell classification and annotation. The malignant cells were identified by copykat and the communication between mast cells and malignant cells was analyzed by CellChat. Least absolute shrinkage and selection operator (LASSO) regression analysis and Cox regression analysis of mast cell markers were performed in the TCGA-COAD cohort to construct a prognostic classification model. qRT-PCR was performed to detect the mRNA expression of the molecules in the classification model in P815 and MC-9 cells. The co-culture experiment of MC38 and P815 cells were performed in 12-well transwell dish. Wound healing assay and Transwell assay were performed to detect cell migration and invasion. RESULTS: 10,186 high-quality cells in GSE146771 were annotated to 9 cell types. Six markers in mast cells (HDC, GATA2, ASAH1, BTBD19, TIMP1, FAM110A) were selected to construct a classification model. The high-risk score defined showed high infiltration of immunosuppressive cells, including endothelial cells, CAFs, Tregs and high angiogenesis and epithelial-mesenchymal transition (EMT) activities. In the model, HDC were abnormally low expressed in P815 cells, while BTBD19, FAM110A, GATA2, ASAH1 and TIMP1 showed excessive expression in P815 cells. Knockdown of GATA2 in the co-culture system of P815 and MC38 cells blocked cell migration and invasion. CONCLUSION: This study identified the cell types within CRC, elaborated the cellular functions of mast cells in CRC development and their molecular communication to coordinate malignant cells, and highlighted the molecular components and biological features that constitute promising prognostic classification model.

Mast Cells

Tribus: semi-automated discovery of cell identities and phenotypes from multiplexed imaging and proteomic data.

MOTIVATION: Multiplexed imaging and single-cell analysis are increasingly applied to investigate the tissue spatial ecosystems in cancer and other complex diseases. Accurate single-cell phenotyping based on marker combinations is a critical but challenging task due to (i) low reproducibility across experiments with manual thresholding, and, (ii) labor-intensive ground-truth expert annotation required for learning-based methods. RESULTS: We developed Tribus, an interactive knowledge-based classifier for multiplexed images and proteomic datasets that avoids hard-set thresholds and manual labeling. We demonstrated that Tribus recovers fine-grained cell types, matching the gold standard annotations by human experts. Additionally, Tribus can target ambiguous populations and discover phenotypically distinct cell subtypes. Through benchmarking against three similar methods in four public datasets with ground truth labels, we show that Tribus outperforms other methods in accuracy and computational efficiency, reducing runtime by an order of magnitude. Finally, we demonstrate the performance of Tribus in rapid and precise cell phenotyping with two large in-house whole-slide imaging datasets. AVAILABILITY AND IMPLEMENTATION: Tribus is available at https://github.com/farkkilab/tribus as an open-source Python package.

Proteomics