PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Count Data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Scale reliant mixed effects models enhance microbiome data analysis.

Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional - they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package.

Microbiota

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq

simPIC:flexible simulation of paired-insertion counts for single-cell ATAC sequencing data.

Single-cell Assay for Transposase Accessible Chromatin (scATAC-seq) is increasingly used at population scale to study how genetic variation shapes chromatin accessibility across diverse cell types. This widespread adoption of the assay has created a need for computational methods that can handle complex biological and technical variation. Yet method development is limited by the lack of flexible simulation tools with known ground truth. Here, we present simPIC, a simulation framework for generating realistic single-cell ATAC-seq data across individuals and cell types. simPIC supports both population-scale and single-individual simulations, with the ability to model cell groups, batch effects, and genotype-dependent variation in accessibility. These features enable realistic benchmarking for tasks such as chromatin accessibility quantitative trait locus (caQTL) mapping. simPIC generates data that closely match real datasets and better captures inter-individual and experimental variation compared to existing tools.

simulation

Whole genome sequencing reveals a specific microbiota in subglottic stenosis C. acnes may contribute to inflammation.

PURPOSE: Subglottic stenosis (SGS) progressively reduces the airway below the vocal folds. The cause is not known and there is a recurrent need of surgical treatment. Including all phenotypes, SGS affects 1/400 000/yr, with a female dominance. Previous studies have revealed a possible role of the Mycobacterium complex in SGS development. Our hypothesis is that microbiota is associated with the inflammation in SGS, if true it might affect the prevailing treatment options. METHODS: This prospective cross-sectional study included biopsies from 34 patients with subglottic stenosis, collected between 2020 and 2023. Nucleic acids were extracted from the tissue samples and analysed using whole genome sequencing. Microbial composition was characterized using taxonomic profiling of sequencing data. Species with sufficient read counts were selected for further validation using sequence alignment methods to ensure accuracy of identification. RESULTS: Using the most comprehensive form of genomic testing currently in clinical use, we present curated and stable data on the presence of Cutibacterium acnes in 28 out of the 34 cases. CONCLUSION: Cutibacterium acnes may serve as a driver of the inflammation characterizing SGS and should be considered in therapeutically oriented future studies.

Cutibacterium acnes

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

Proportionality-based association metrics in count compositional data.

Compositional data comprise vectors that describe the constituent parts of a whole. Data arising from various -omics platforms such as 16S and RNA sequencing are compositional in nature. In this kind of data, correlations between features on raw counts have no meaningful interpretation. Metrics of proportionality were formulated to address this problem. However, an inherent bias arises when these metrics are calculated empirically on count-based measures due to variability in read depths. We quantify the bias introduced by empirically calculating proportionality-based association metrics in count data. Additionally, we propose a means of estimating these metrics within a logit-normal multinomial model in pursuit of more accurate estimates. The model-based estimates are shown to outperform empirical estimates in simulated data and are applied to a mouse embryonic stem cell single-cell sequencing dataset, as well as a pediatric-onset multiple sclerosis metagenomic dataset.

Animals

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

Ribosomal DNA copy number variation associates with hematological profiles and renal function in the UK Biobank.

The phenotypic impact of genetic variation of repetitive features in the human genome is currently understudied. One such feature is the multi-copy 47S ribosomal DNA (rDNA) that codes for rRNA components of the ribosome. Here, we present an analysis of rDNA copy number (CN) variation in the UK Biobank (UKB). From the first release of UKB whole-genome sequencing (WGS) data, a discovery analysis in White British individuals reveals that rDNA CN associates with altered counts of specific blood cell subtypes, such as neutrophils, and with the estimated glomerular filtration rate, a marker of kidney function. Similar trends are observed in other ancestries. A range of analyses argue against reverse causality or common confounder effects, and all core results replicate in the second UKB WGS release. Our work demonstrates that rDNA CN is a genetic influence on trait variance in humans.

Humans

CountASAP: a lightweight, easy to use python package for processing ASAPseq data.

BACKGROUND: Declining sequencing costs coupled with the increasing availability of easy-to-use kits for the isolation of DNA and RNA transcripts from single cells have driven a rapid proliferation of studies centered around genomic and transcriptomic data. Simultaneously, a wealth of new techniques have been developed that utilize single cell technologies to interrogate a broad range of cell-biological processes. One recently developed technique, transposase-accessible chromatin with sequencing (ATAC) with select antigen profiling by sequencing (ASAPseq), provides a combination of chromatin accessibility assessments with measurements of cell-surface marker expression levels. While software exists for the characterization of these datasets, there currently exists no tool explicitly designed to reformat ASAP surface marker FASTQ data into a count matrix which can then be used for these downstream analyses. RESULTS: To address this lack of a dedicated tool for ASAPseq data processing, we created CountASAP, an easy-to-use Python package purposefully designed to transform FASTQ files from ASAP experiments into count matrices compatible with commonly-used downstream bioinformatic analysis packages. CountASAP takes advantage of the independence of the relevant data structures to perform fully parallelized matches of each sequenced read to user-supplied input ASAP oligos and unique cell-identifier sequences. We directly compare the performance and user-friendliness of CountASAP to existing tools using similarly-structured data from a more common sequencing experiment: cellular indexing of transcriptomes and epitopes by sequencing (CITEseq). Further benchmarking against existing tools helps to identify proper defaults for CountASAP and assess the agreement of outputs from all tested software. A final test using a novel ASAPseq dataset provides evidence that CountASAP can generate biologically meaningful results that correlate well with paired chromatin accessibility data. CONCLUSIONS: CountASAP shows good agreement with existing, well-tested data processing tools in the analysis of similarly-structured benchmarking data. CountASAP runs efficiently on a standard laptop, has user-friendly documentation, a one-step installation, and represents the first and only tool designed specifically for the processing of ASAPseq data.

Software

Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches.

Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Tools like the widely used Blacklist software facilitate this process, but their algorithmic details and parameter choices are not always clearly documented, affecting reproducibility and biological relevance. We examined the Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to variability in input data, aligner choice, and read length. We also identified and addressed a coding issue that led to over-annotation of high-signal regions. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets and suggest that sponge sequences offer a flexible, alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Journal Article

shinyDeepGxP: a user-friendly R shiny app for predicting surface protein abundance from scRNA-seq expression using deep learning in blood cells.

MOTIVATION: Understanding accurate immune cell heterogeneity and function in single-cell datasets requires access to protein-level information, which is often unavailable due to experimental limitations. RESULTS: We present shinyDeepGxP, an interactive web application featuring our deep learning model, DeepGxP, for predicting surface protein abundance from single-cell RNA-sequencing (scRNA-seq) data. This platform makes DeepGxP accessible to researchers without programming skills. Users can upload scRNA-seq count matrices and use "Predict Protein" to predict the abundance of 224 biologically relevant surface proteins. shinyDeepGxP provides visualizations to help identify distinct cell populations based on predicted protein profiles. Moreover, users can choose "Explore Model" to reveal key RNA predictors and their associated biological pathways for each protein. Overall, shinyDeepGxP is a user-friendly, freely available web tool that provides protein-level detail for RNA-only single-cell datasets, enabling multimodal discovery without additional experiments. AVAILABILITY AND IMPLEMENTATION: shinyDeepGxP can be launched on https://shiny.crc.pitt.edu/deepgxp/.

Journal Article

Detection of short tandem repeats in the cattle genome: a comparison of bioinformatic tools.

BACKGROUND: Short tandem repeats (STRs) are repetitive DNA sequences with 1–6 nucleotide repeat units, exhibiting high polymorphism due to varying repeat counts. STRs are more variable than SNPs and can cause genetic disorders. With population-scale cattle whole-genome sequencing data available, whole-genome STR identification has attracted new interest, but challenges remain due to the lack of standardized methods, sequencing data limitations, and the diversity of STR-calling tools. This study compared six STR-calling tools: HipSTR, GangSTR, and ExpansionHunter for short-read data, and Straglr, RepeatHMM, and LongTR for Oxford Nanopore (ONT) long-read data—using sequences from five Holstein cattle (two parent–offspring trios with a shared sire). This is the first cattle study to evaluate short- and long-read STR callers using both data types from the same animals. RESULTS: In short-read data, ExpansionHunter identified the highest number of polymorphic STRs (pSTRs) (327,690), followed by HipSTR (205,900) and GangSTR (110,680), with 93,023 loci detected by all three tools. In long-read data, LongTR detected 470,250 pSTRs, RepeatHMM 224,185, and Straglr 90,275, with only 33,253 loci shared among them. Mendelian consistency of STR genotypes in the trio offspring was high (> 0.8) for all short-read tools, with HipSTR and GangSTR highest at 0.98. LongTR was the only long-read tool with high consistency (0.88). Short-read tools also showed higher concordance in STR genotypes among themselves than was observed among long-read tools. However, long-read tools had a clear advantage in detecting large STRs. Relative to computational efficiency, HipSTR and GangSTR (short-reads), and LongTR (long-reads) required less memory and shorter runtimes than the other tools. CONCLUSIONS: Tool selection is critical for accurate whole-genome STR identification in cattle. For short-read data, HipSTR showed relatively high Mendelian consistency and concordance compared to the other tools, while ExpansionHunter was able to detect longer STRs but with lower Mendelian consistency. For long-read data, LongTR demonstrated higher consistency and computational efficiency relative to the other tools. Based on these results, HipSTR and LongTR are suggested as preferred options for short-read and ONT long-read datasets, respectively, in cattle STR analysis. These recommendations are based on the metrics observed in this study, and confirmatory analyses across additional breeds, larger sample sizes, and validated truth sets are encouraged.

Animals

Beyond blacklists: a critical assessment of exclusion set generation strategies and alternative approaches.

MOTIVATION: Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Alternatively, "sponge" or decoy sequences have been proposed to reduce alignment artifacts. RESULTS: We examined the widely used Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to sensitivity to input data, aligner choice, and read length. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. We additionally investigated the effect of the T2T-CHM13 genome assembly on improving biological signals. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets, and recommend the use of the T2T-CHM13 assembly or, for the hg38 genome assembly, "sponge" sequences as an alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Software

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

In vitro and in vivo studies on the impact of the familial adenomatous polyposis heterogeneous mutation MUC20-S671C on colorectal carcinogenesis and progression.

BACKGROUND: Familial adenomatous polyposis (FAP) is a hereditary colorectal cancer (CRC). We performed genetic testing on nine FAP patients and identified a recurrent mutation at the 671st site of the MUC20 gene-MUC20-S671C. This mutation has a detection frequency of zero in the 1000 Genomes Project database. Previous studies have demonstrated that MUC20 can promote CRC progression through epithelial-mesenchymal transition (EMT). We conducted a series of experiments to analyze the impact of this mutation on CRC cells, aiming to infer its potential role and significance in CRC patients. METHODS: We introduced the MUC20-S671C mutation into the CRC SW480 cell line using the CRISPR-Cas9 technique and established a stable cell line carrying this mutation. We then conducted various experiments to assess the effects of this mutation. The Transwell assay was used to evaluate cell invasion and migration. We also examined cell proliferation, cell cycle progression, and apoptosis rate. Furthermore, we tested the tumorigenic ability of these cells in NOD-scid IL2Rγ[null] (NSG) mice. Additionally, transcriptome sequencing was performed on both cell lines and mouse tumor tissues to obtain molecular regulatory network data, and key molecules were further validated. RESULTS: The results of Cell Counting Kit-8 (CCK-8), 5-ethynyl-2'-deoxyuridine (EdU), and colony formation assays indicated that the proliferation ability of mutant cells was significantly reduced. The Transwell assay demonstrated a marked decline in the invasion and migration capabilities of mutant cells. Flow cytometry analysis revealed that the mutation increased the apoptosis rate of CRC cells and might have caused S-phase arrest. The tumor formation assay in nude mice showed that the tumorigenic ability of mutant cells was weakened. Transcriptome sequencing of both the cells and tumor tissues suggested that the mutation altered the expression of apoptosis- and cell cycle-related molecules and also affected EMT. Further experiments confirmed that key molecules involved in the EMT process, such as E-cadherin, were upregulated, while Vimentin, MMP9, and MMP14 were significantly downregulated, indicating that the mutation weakened the EMT capability of CRC cells. CONCLUSIONS: We have identified a novel mutation, MUC20-S671C, in patients with FAP. Our study demonstrates that this mutation exerts its tumor-suppressive effect by reversing the EMT process.

MUC20-S671C

Non-structural maintenance of chromosome condensin I complex subunit H knockdown suppresses malignant progression of esophageal squamous cell carcinoma via the Wnt/β-catenin signaling pathway.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) remains a major cause of cancer-related mortality, and effective therapeutic targets are still limited. Non-structural maintenance of chromosome condensin I complex subunit H (NCAPH) has been implicated in tumorigenesis; however, its clinical relevance, functional roles, and underlying mechanisms in ESCC are not fully defined. We aimed to characterize the expression pattern, prognostic value, biological functions, and mechanistic basis of NCAPH in ESCC. METHODS: Public datasets from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) were analyzed to evaluate NCAPH expression and clinical associations. Single-cell RNA sequencing (scRNA-seq) data were used to map cell-type-specific distribution of NCAPH in tumor and adjacent tissues. NCAPH was silenced in KYSE150 and KYSE510 cells using lentiviral short hairpin RNAs (shRNAs), followed by Cell Counting Kit-8 (CCK-8), colony formation, wound-healing, and Transwell migration/invasion assays. A nude mouse xenograft model was established to assess the effect of NCAPH knockdown in vivo. RNA sequencing (RNA-seq), quantitative polymerase chain reaction (qPCR), western blotting, and enzyme-linked immunosorbent assay (ELISA) were performed to explore potential mechanisms. RESULTS: NCAPH was consistently upregulated in ESCC across multiple cohorts and was associated with unfavorable clinicopathological features and poorer survival. Functional assays demonstrated that NCAPH knockdown significantly inhibited ESCC cell proliferation, migration, invasion, and clonogenic growth. In vivo, NCAPH silencing suppressed xenograft tumor growth. Mechanistically, transcriptomic profiling and molecular validation indicated attenuation of Wnt/β-catenin signaling following NCAPH depletion, accompanied by reduced β-catenin and downstream targets. CONCLUSIONS: NCAPH promotes malignant progression of ESCC, at least in part through activation of the Wnt/β-catenin pathway, and may serve as a potential biomarker and therapeutic target.

Esophageal squamous cell carcinoma (ESCC)