PubMed HealthSearch

SEARCH · PubMed Health

Results for “batch effects”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Moderated designs can balance between batch-effect mitigation and cell loss due to hashtag-assisted pooling in single-cell experiments.

Minimizing experimental noise is integral to robust data generation in single-cell omics. The current standard for avoiding batch effects during sample processing is barcode- or hashtag-assisted combining of different experimental treatments into one pool, allowing all samples to be subject to the technical protocols uniformly. The final data points for each treatment group are then computationally separated based on the original hashtag labels. Clearly, whereas hashtagging all groups and pooling them in a single well is expected to minimize batch effects, the procedure can also lead to a loss of cells that cannot be confidently decoded during the computational demultiplexing step. Here, we examine four alternate experimental designs, namely, compound, reference, chain, and confounded, that could be used instead of a single-pool approach and quantify the batch effects as well as cell loss in each case. We find a linear relationship-the percentage of cells lost is double the number of hashtags used in the experiment. We use these analyses to identify experimental designs that can successfully mitigate batch effects while minimizing multiplexing, hence the cell loss, in each well. Although a reference design offers the best overall performance, this study can help individual investigators choose particular approaches that are best suited for their biological questions.

Journal Article

NPM: latent batch effects correction of omics data by nearest-pair matching.

MOTIVATION: Batch effects (BEs) are a predominant source of noise in omics data and often mask real biological signals. BEs remain common in existing datasets. Current methods for BE correction mostly rely on specific assumptions or complex models, and may not detect and adjust BEs adequately, impacting downstream analysis and discovery power. To address these challenges we developed NPM, a nearest-neighbor matching-based method that adjusts BEs and may outperform other methods in a wide range of datasets. RESULTS: We assessed distinct metrics and graphical readouts, and compared our method to commonly used BE correction methods. NPM demonstrates the ability in correcting for BEs, while preserving biological differences. It may outperform other methods based on multiple metrics. Altogether, NPM proves to be a valuable BE correction approach to maximize discovery in biomedical research, with applicability in clinical research where latent BEs are often dominant. AVAILABILITY AND IMPLEMENTATION: NPM is freely available on GitHub (https://github.com/bigomics/NPM) and on Omics Playground (https://bigomics.ch/omics-playground). Computer codes for analyses are available at (https://github.com/bigomics/NPM). The datasets underlying this article are the following: GSE120099, GSE82177, GSE162760, GSE171343, GSE153380, GSE163214, GSE182440, GSE163857, GSE117970, GSE173078, and GSE10846. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

Whole metagenome sequencing: not deep enough for complete microbial function recovery.

BACKGROUND: Whole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data was annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. RESULTS: Three findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in, e.g., case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this study's sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. We further showed that despite biomes differences in functional diversity and host contamination levels (e.g., soil, fecal), incomplete functional recovery at commonly used sequencing depths was consistently observed. CONCLUSIONS: Together, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights. Video Abstract.

Humans

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

OmicsTweezer: A distribution-independent cell deconvolution model for multi-omics Data.

Cell deconvolution estimates cell type proportions from bulk omics data, enabling insights into tissue microenvironments and disease. However, practical applications are often hindered by batch effects between bulk data and referenced single-cell data, a challenge that is frequently overlooked. To address this discrepancy, we developed OmicsTweezer, a distribution-independent cell deconvolution model. By integrating optimal transport with deep learning, OmicsTweezer aligns simulated and real data in a shared latent space, effectively mitigating data shifts and inter-omics distribution differences. OmicsTweezer is versatile, capable of deconvolving bulk RNA-seq, bulk proteomics, and spatial transcriptomics. Extensive evaluations on simulated and real-world datasets demonstrate its robustness and accuracy. Furthermore, applications in prostate and colon cancer showcase OmicsTweezer's ability to identify biologically meaningful cell types. As a unified deconvolution framework for multi-omics data, OmicsTweezer offers an efficient and powerful tool for studying disease microenvironments.

Humans

Development of methodology to support molecular endotype discovery from synovial fluid of individuals with knee osteoarthritis: The STEpUP OA consortium.

OBJECTIVES: To develop a protocol for largescale analysis of synovial fluid proteins, for the identification of biological networks associated with subtypes of osteoarthritis. METHODS: Synovial Fluid To detect molecular Endotypes by Unbiased Proteomics in Osteoarthritis (STEpUP OA) is an international consortium utilising clinical data (capturing pain, radiographic severity and demographic features) and knee synovial fluid from 17 participating cohorts. 1746 samples from 1650 individuals comprising OA, joint injury, healthy and inflammatory arthritis controls, divided into discovery (n = 1045) and replication (n = 701) datasets, were analysed by SomaScan Discovery Plex V4.1 (>7000 SOMAmers/proteins). An optimised approach to standardisation was developed. Technical confounders and batch-effects were identified and adjusted for. Poorly performing SOMAmers and samples were excluded. Variance in the data was determined by principal component (PC) analysis. RESULTS: A synovial fluid standardised protocol was optimised that had good reliability (<20% co-efficient of variation for >80% of SOMAmers in pooled samples) and overall good correlation with immunoassay. 1720 samples and >6290 SOMAmers met inclusion criteria. 48% of data variance (PC1) was strongly correlated with individual SOMAmer signal intensities, particularly with low abundance proteins (median correlation coefficient 0.70), and was enriched for nuclear and non-secreted proteins. We concluded that this component was predominantly intracellular proteins, and could be adjusted for using an 'intracellular protein score' (IPS). PC2 (7% variance) was attributable to processing batch and was batch-corrected by ComBat. Lesser effects were attributed to other technical confounders. Data visualisation revealed clustering of injury and OA cases in overlapping but distinguishable areas of high-dimensional proteomic space. CONCLUSIONS: We have developed a robust method for analysing synovial fluid protein, creating a molecular and clinical dataset of unprecedented scale to explore potential patient subtypes and the molecular pathogenesis of OA. Such methodology underpins the development of new approaches to tackle this disease which remains a huge societal challenge.

Humans

Diagnosing scientific replicability through probabilistic distinguishability.

MOTIVATION: Despite the widely recognized importance of replicability in biological research, computational methods to quantify irreplicability and identify irreplicable instances remain underdeveloped. This article presents an efficient and robust computational framework to address this gap. RESULTS: To tackle the challenge of defining an acceptable level of intrinsic heterogeneity among replicable studies, we introduce a distinguishability criterion, ensuring that replicable effects, while potentially heterogeneous, can be distinguished from zero effects and maintain consistent directions with high probability. We implement a Bayesian model criticism approach, reporting a Bayesian P-value to identify potential irreplicable instances. Through numerical experiments, we demonstrate the efficacy of the proposed methods in detecting batch effects in high-throughput experiments and identifying instances of the publication bias. Finally, we apply the framework to multi-tissue eQTL data from the GTEx consortium, uncovering tissue-specific eQTLs that represent biological heterogeneity across tissues. AVAILABILITY AND IMPLEMENTATION: An R package DiscRep implementing our method is available on GitHub (https://github.com/PengWang96/DiscRep).

Bayes Theorem

Meta-Merging the Transcriptomes of Gastric Tumors Redefines the Connections among Molecular and Clinical Subtypes.

INTRODUCTION: The availability of a large number of cancer expression profiles presents an excellent opportunity to re-investigate various biological and clinical questions. While several expression profiles have been established for different cancers, merging them may provide a more powerful platform for extensively extrapolating molecular and clinical features across multiple cohorts. MATERIALS AND METHODS: In this study, five gastric tumor expression profiles from the Gene Expression Omnibus [GEO] and one in-house cohort comprising a total of 1,060 samples were merged. The batch effect was removed using non-parametric ComBat analysis, and the seamless merging of datasets was confirmed through various parameters. RESULTS: Extrapolation of ACRG [Asian Cancer Research Group] and TCGA [The Cancer Genome Atlas] molecular subtypes in the merged cohort of 1,060 gastric tumors revealed nine distinct clusters. Notably, the following patterns were observed: [i] mutual exclusivity between Epithelial to Mesenchymal Transition [EMT] and Microsatellite Instability [MSI] subtypes in 90% of tumors; [ii] overlapping occurrence of EMT and MSI subtypes in the remaining tumors; [iii] overlap between MSI and Epstein-Barr Virus [EBV] subtype tumors; [iv] both commonalities and differences between EMT and Genomically Stable [GS] subtypes; and [v] an association between EBV positivity and PI3K mutation. CONCLUSION: The current study demonstrates that compiling a larger expression profile is valuable for revisiting the molecular features and epidemiology associated with molecular subtypes, thereby aiding in the development of novel diagnostics and targeted therapeutics.

Humans

Human Systems Immunology in the Omics Era: Challenges, Methods, and Emerging Directions.

The human immune system is a highly complex, dynamic, and heterogeneous network shaped by genetic, environmental, and temporal influences. Advances in high-throughput omics technologies have transformed our ability to study this complexity directly and comprehensively in human cohorts. These developments have positioned systems immunology as a powerful framework for investigating coordinated immune responses, identifying regulatory mechanisms, and linking molecular patterns to clinical phenotypes. However, the analytical challenges inherent to large-scale, multimodal datasets-including batch effects, small sample sizes, high dimensionality, and substantial interindividual heterogeneity-require rigorous study design, robust statistical modeling, and thoughtful data analysis strategies. In this review, we summarize key technological foundations enabling modern human systems immunology, outline common analytical pitfalls and effective mitigation approaches, discuss data integration concepts, and highlight emerging opportunities in the field. Together, these technological and analytical advances are redefining how immune function is measured and interpreted in real-world human biology and hold significant promise for enhancing mechanistic insight, biomarker discovery, and precision medicine across immunological diseases and interventions.

Humans

Bioinformatics Analysis and Experimental Validation of Key Genes Associated With Hypoxia and Ischemia in Myocardial Infarction.

BACKGROUND: This study aimed to screen and identify core hypoxia-ischemia-related genes associated with myocardial infarction (MI). METHOD: Two transcriptomic datasets, GSE97320 and GSE48060, were retrieved from the Gene Expression Omnibus (GEO) database. After data integration and batch effect elimination, differential expression analysis was performed to screen differentially expressed genes (DEGs), and the corresponding visualization analysis was conducted. Hypoxia-ischemia-related genes were acquired from the GeneCards database; hypoxia-ischemia related genes (HIRGs) were subsequently identified by intersecting the retrieved genes with screened DEGs. Gene Ontology (GO) functional enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were implemented to explore the biological functions and underlying signaling pathways of HIRGs. A combination of protein-protein interaction (PPI) network analysis and random forest (RF) algorithm was applied to screen hub genes from HIRGs. The external GEO dataset GSE66360 was utilized to validate the expression patterns of candidate hub genes. Furthermore, an acute myocardial infarction (AMI) mouse model was established, and quantitative real-time polymerase chain reaction (qPCR) was performed to detect the mRNA expression levels of hub genes in myocardial tissues for in&#xa0;vivo validation. RESULTS: A total of 633 DEGs and 308 hypoxia-ischemia-related genes were screened in the present study, among which 21 overlapping HIRGs were obtained. PLAUR and IL1B were finally identified as two hub genes from HIRGs based on PPI network and random forest algorithm. The qPCR results revealed that the expression levels of PLAUR and IL1B were significantly upregulated in the AMI group compared with the sham operation group (p&#x2009;<&#x2009;0.05). CONCLUSION: The present findings demonstrated that PLAUR and IL1B serve as pivotal genes involved in the pathological hypoxia-ischemia process of AMI. These two genes may act as novel biomarkers and promising therapeutic targets for the recognition and clinical intervention of hypoxia-ischemia injury following AMI.

Myocardial Infarction

Tahoe-100M: Mapping drug-induced molecular phenotypes at single-cell resolution.

We present Tahoe-100M, a giga-scale single-cell perturbation atlas comprising 100 million transcriptomes from 50 diverse cancer cell lines treated with 1,100 drug-dose conditions. This parallel profiling of thousands of perturbations at single-cell resolution with minimal batch effects is enabled by the Mosaic platform, which multiplexes genetically distinct cell models into balanced "cell villages." Beyond cataloging transcriptomic shifts, Tahoe-100M systematically quantifies cellular phenotypes, including proliferation, cytotoxicity, lineage-specific vulnerabilities, and cell-cycle changes. It captures population-level transcriptomic heterogeneity, characterizing whether drug responses drive cells toward divergent fates or convergent states. Pathway-based signatures define drug-induced expression programs, classify mechanisms of action, reveal off-target activities, and expose adaptive stress responses associated with resistance. By unifying cellular and molecular readouts, this broadly applicable perturbation atlas advances our ability to model gene regulation, drug response, and network dynamics. Its public release enables the training of AI frameworks to advance predictive models of cell behavior.

Humans

Stage-Independent Real-Time Subtype Classification and Comprehensive Biopsy Profiling of Urothelial Carcinomas by the Lund Taxonomy System.

Bladder cancer is a heterogeneous malignancy with diverse clinical outcomes, and conventional pathological assessment alone is insufficient to capture its underlying biology. Gene expression profiling can stratify tumors into molecular subtypes with prognostic and predictive potential, but the reliability of transcriptomic classification and its clinical utility remains to be established. The translational/observational UROSCANSEQ study (ISRCTN15459149) prospectively evaluates RNA-based Lund Taxonomy (LundTax) molecular subtype classification in a clinical setting. Among 784 consecutive biopsies collected between 2018 and 2022, RNA sequencing was successful for 90% of all biopsies, encompassing 662 bladder cancer patients with a stage distribution of 48% Ta, 27% T1, 24% &#x2265;T2, and 1% CIS. We demonstrate that the LundTax subtype classification algorithm, applied to individual samples, accurately identifies cancer cell phenotypes with characteristic gene and protein expression patterns in a manner robust to RNA quality, data preprocessing strategies, and batch effects, supporting its clinical feasibility across both non-muscle-invasive and muscle-invasive disease. We further extend the LundTax framework by incorporating single-sample molecular risk scores reflecting tumor grade, proliferation, and progression risk, as well as tumor microenvironment signatures. Both risk scores and overall immune and stromal content in biopsies were significantly associated with an increased risk of clinical progression in noninvasive disease. In a separate analysis of the relative cellular composition of the tumor microenvironment, however, only the fraction of natural killer cells remained significant. Together, the expanded LundTax system provides a comprehensive molecular portrait of individual tumor biopsies. By explicitly separating cancer cell-intrinsic phenotypes, prognostic indexes, and microenvironmental signals, the framework minimizes biological confounding and establishes a strong foundation for future studies evaluating clinical outcomes and treatment responses.

Humans

NoisyFlow: differentially private optimal transport using neural networks for secure biomedical data sharing across multiple institutions.

MOTIVATION: Biomedical models improve when trained on data pooled across institutions, but sensitive patient records (e.g. genomics, clinical data, and medical images) are difficult to share due to privacy constraints. Moreover, data collected at different sites often have shifted distributions because of covariate differences (including batch effects), so privacy-preserving sharing alone cannot simply resolve cross-site mismatch. Methods that protect individuals while explicitly aligning distributions are needed to enable reliable multi-institutional analyses. RESULTS: We present NoisyFlow, a three-stage differentially private framework for cross-institutional harmonization under distribution shift. In stage I, each site learns a differentially private flow-based generator of its local labeled distribution. In stage II, it learns a neural optimal transport map to a shared reference distribution. In stage III, a central server composes the released models to generate reference-aligned pseudo-data for downstream analysis without accessing raw records. Across four biomedical settings spanning single-cell genomics, histopathology, neurogenomics, and wearable sensing, NoisyFlow reduces distribution shift while preserving downstream utility under formal differential privacy guarantees. AVAILABILITY AND IMPLEMENTATION: The implementation of NoisyFlow is available at https://github.com/gersteinlab/NoisyFlow.

Information Dissemination

RAREsim2: flexible simulation of rare variant genetic data using real haplotypes.

MOTIVATION: Realistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., case-control status, technological or batch effects) and variant-level differences to represent a variety of causal models. RESULTS: We demonstrate RAREsim2's utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios, including various genetic ancestries, gene sizes, strengths of association, and proportions of risk variants. Type I Error was maintained and the test with the highest power matched previously known patterns. Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. AVAILABILITY AND IMPLEMENTATION: The RAREsim2 Python package is publicly available on Github (https://github.com/Hendricks-Research-Team/RAREsim2), PyPI (https://pypi.org/project/raresim/), and Zenodo (https://doi.org/10.5281/zenodo.19442523). Code for the example demonstration can be found at https://github.com/JessMurphy/RAREsim2-demo.

Software

Identification of key genes related to bone metastasis of breast cancer using bioinformatics methods and construction of a prognostic model.

Breast cancer (BC) ranks among the most prevalent cancers in females, with bone metastasis significantly compromising patients' quality of life and survival rates. Enhancing our comprehension of BC bone metastasis mechanisms at the molecular level holds promise for improving BC treatment and prognosis. Leveraging bioinformatics tools, we integrated multiple datasets, conducted comprehensive analyses across various databases, identified biomarkers associated with BC bone metastasis, and constructed a prognostic model. Firstly, 3 BC bone metastasis-related datasets were downloaded from gene expression omnibus, the data were merged, and batch effects were removed, followed by identification of differentially expressed genes (DEGs). Gene ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the DEGs. A protein-protein interaction network was constructed using the STRING database to screen hub genes. Then, survival analysis of hub genes was performed using the Cancer Genome Atlas (TCGA) database. A prognostic model was constructed using key genes with survival differences, and the model was evaluated. Two hundred ninety-two DEGs were identified. Gene ontology and KEGG pathway enrichment analysis yielded 769 biological processes (BPs), 78 cellular components, 43 molecular functions, and 50 KEGG pathways. Fifteen hub genes were selected from the protein-protein interaction network. Survival analysis revealed 6 genes related to BC survival. The prognostic model identified 4 genes with important predictive value for BC prognosis. Our study utilized bioinformatics analysis to identify a series of DEGs related to BC bone metastasis. Based on further selection of hub genes, we constructed a relatively ideal prognostic model for BC, and identified 4 genes (DLGAP5, TPX2, PLK1, and CENPN) with valuable predictive value for BC prognosis.

Humans

simPIC:flexible simulation of paired-insertion counts for single-cell ATAC sequencing data.

Single-cell Assay for Transposase Accessible Chromatin (scATAC-seq) is increasingly used at population scale to study how genetic variation shapes chromatin accessibility across diverse cell types. This widespread adoption of the assay has created a need for computational methods that can handle complex biological and technical variation. Yet method development is limited by the lack of flexible simulation tools with known ground truth. Here, we present simPIC, a simulation framework for generating realistic single-cell ATAC-seq data across individuals and cell types. simPIC supports both population-scale and single-individual simulations, with the ability to model cell groups, batch effects, and genotype-dependent variation in accessibility. These features enable realistic benchmarking for tasks such as chromatin accessibility quantitative trait locus (caQTL) mapping. simPIC generates data that closely match real datasets and better captures inter-individual and experimental variation compared to existing tools.

simulation