PubMed HealthSearch

SEARCH · PubMed Health

Results for “statistical inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Schizophrenia and bipolar disorder: a comparative analysis of genetic and brain network connectivity.

BACKGROUND: Schizophrenia (SCZ) and bipolar disorder (BD) are severe psychiatric conditions with overlapping clinical presentations, genetic risk factors, and brain network dysfunction. Whether alterations in large-scale intrinsic brain networks reflect shared or disorder-specific genetic influences remains poorly understood. Clarifying this distinction is essential for refining etiological models and improving diagnostic precision. METHODS: Genome-wide inferred statistics (GWIS) were applied to decompose the genetic architecture of SCZ and BD into shared and unique components. Using resting-state network (RSN) data from the UK Biobank, functional connectivity (FC) and structural connectivity (SC) were extracted as neuroimaging phenotypes. Causal inference approaches were subsequently employed to infer potential directional relationships between brain network connectivity and each disorder. RESULTS: Analyses revealed both common and distinct patterns of brain network connectivity associated with SCZ and BD. Notably, SC within the default mode network (DMN) exhibited opposing effects across the two disorders, suggesting divergent structural underpinnings despite clinical overlap. Additionally, SC within the limbic network (LN) and frontotemporal control network demonstrated potential causal relationships with both conditions, implicating these circuits astransdiagnostic neural substrates. CONCLUSION: These findings illuminate the shared and disorder-specific genetic and neural architecture underlying SCZ and BD. Integrating genome-wide genetic methods with large-scale neuroimaging data offers a powerful framework for disentangling psychiatric comorbidity and may inform more targeted diagnostic criteria and individualized treatment strategies.

Humans

Molecular dynamics simulations of positively selected codons in FcγRI reveal novel biochemical binding properties.

FcγRI is a high-affinity receptor for IgG, associated with autoimmune disease pathology and determines clinical responses to antibody-based immunotherapies. FcγRI has a complex evolutionary history that is not fully understood, and to address this we explored signatures of positive selection in the receptor's functional gene, FCGR1A, using codon-based selection tests on aligned 1-1 orthologous sequences from placental mammals (n = 32). Signatures of positive selection have occurred at several locations within the gene, with two sites (H148 (M2a ω 0.997 & M8 ω = 0.993)) and (W149 (M2a ω = 0.999 & M8 ω = 1.000)) exhibiting highest posterior probabilities, suggesting strong evidence of positive selection; these positions are known to form one of the FcγRI-IgG binding interfaces. We employed ancestral reconstruction to statistically infer prior codon sequences at these sites and identified ancestral H148P and W149R codons at different nodes in the phylogeny. Employing molecular dynamics simulations, we determined how evolutionary changes at these sites may have influenced the binding of FcγRI-IgG of modern-day Homo sapiens. Measuring RMSD, free energy, radius of gyration, hydrogen bond formation, and analyzing free energy landscapes, we demonstrate that structural instability between mutant structures vs the WT counterpart; however, overall binding potential increases at position 148, yet decreases at 149 in potential. H148P protonation at physiological pH remains similar, yet during acidotic calculations, protonation is likely reduced, with predicted reduction in affinity for IgG. While ancestral W149R substitutions demonstrate an implication for electron conjugation. Examining key sites at this binding FcγRI-IgG interface, our data demonstrate that these two codons have evolved in humans to be relatively insensitive to shifts in pH promoting a more stable interaction with the Fc portion of IgG during diseases that promote acidosis.

Receptors, IgG

S-GMAS: Genome-Wide Mediation Analysis With Brain Subcortical Shape Mediators.

Mediation analysis is widely utilized in neuroscience to investigate the role of brain image phenotypes in the neurological pathways from genetic exposures to clinical outcomes. However, it is still difficult to conduct mediation analyses with whole genome-wide exposures and brain subcortical shape mediators due to several challenges including (i) large-scale genetic exposures, that is, millions of single-nucleotide polymorphisms (SNPs); (ii) nonlinear Hilbert space for shape mediators; and (iii) statistical inference on the direct and indirect effects. To tackle these challenges, this paper proposes a genome-wide mediation analysis framework with brain subcortical shape mediators. First, to address the issue caused by the high dimensionality in genetic exposures, a fast genome-wide association analysis is conducted to discover potential genetic variants with significant genetic effects on the clinical outcome. Second, the square-root velocity function representations are extracted from the brain subcortical shapes, which fall in an unconstrained linear Hilbert subspace. Third, to identify the underlying causal pathways from the detected SNPs to the clinical outcome implicitly through the shape mediators, we utilize a shape mediation analysis framework consisting of a shape-on-scalar model and a scalar-on-shape model. Furthermore, the bootstrap resampling approach is adopted to investigate both global and spatial significant mediation effects. Finally, our framework is applied to the corpus callosum shape data from the Alzheimer's Disease Neuroimaging Initiative.

Humans

Causal Mediation Analysis for Integrating Exposure, Genomic, and Phenotype Data.

Causal mediation analysis provides an attractive framework for integrating diverse types of exposure, genomic, and phenotype data. Recently, this field has seen a surge of interest, largely driven by the increasing need for causal mediation analyses in health and social sciences. This article aims to provide a review of recent developments in mediation analysis, encompassing mediation analysis of a single mediator and a large number of mediators, as well as mediation analysis with multiple exposures and mediators. Our review focuses on the recent advancements in statistical inference for causal mediation analysis, especially in the context of high-dimensional mediation analysis. We delve into the complexities of testing mediation effects, especially addressing the challenge of testing a large number of composite null hypotheses. Through extensive simulation studies, we compare the existing methods across a range of scenarios. We also include an analysis of data from the Normative Aging Study, which examines DNA methylation CpG sites as potential mediators of the effect of smoking status on lung function. We discuss the pros and cons of these methods and future research directions.

causal inference

Dissecting the relationship between haplotypes around ATXN2 CAG repeats and the number of CAA interruptions by long-read sequencing.

BACKGROUND: CAG repeat expansions in ATXN2 are implicated as risk factors for several neurological diseases, including spinocerebellar ataxia type 2 (SCA2) when >=33 CAG repeats are present, and amyotrophic lateral sclerosis (ALS) when 27-33 CAG repeats are present. However, how haplotypes around the repeats and CAA interruptions within the repeats are associated with disease phenotypes remains poorly understood. Previous studies on haplotypes around ATXN2 were limited to SNPs very close to the repeats (<5kb) or were based on statistical inference only. METHODS: Here, we used long-read sequencing on the Oxford Nanopore Technologies (ONT) platform to simultaneously infer haplotypes around ATXN2, the number of CAG repeats, and the number of CAA interruptions, along with NYGC ALS Consortium NGS dataset. We further sequenced 41 individuals (EUR = 39) with neurological diseases with intermediate repeats by ONT. RESULTS: We found that haplotypes around ATXN2 and the number of interruptions show ethnicity-specific and ALS-specific distribution. Three CAA interruptions are present at low prevalence (~1%) in control populations in multiple ancestry groups, but high prevalence (~55%) in ALS individuals with intermediate repeats. Furthermore, we examined 159 individuals with ALS (~90% European ancestry) with intermediate ATXN2 repeats and found a unique haplotype in ALS individuals with three CAA interruptions, which can be tagged by an SNV, rs148019457. We also validated that the rs148019457-G allele is only present in haplotypes with three CAA interruptions. CONCLUSIONS: In summary, our study shows that 3 CAA interruptions are rarely seen in healthy controls but are common in those with expanded ATXN2 CAG repeats who have neurological disorders, and that rs148019457 tags a specific haplotype with 3 CAA interruptions within expanded ATXN2 CAG repeats in individuals of European ancestry. These results have implications for the development of precision genomic medicine for neurological disorders, and the tag SNP may help identify those with interruptions from existing population genotyping data.

ATXN2

Reconstructing the early spatial spread of pandemic respiratory viruses in the United States.

Understanding the geographic spread of emerging respiratory viruses is critical for pandemic preparedness, yet the early spatiotemporal dynamics of the 2009 H1N1 pandemic influenza and severe acute respiratory syndrome coronavirus 2 in the United States remain unclear. While mobility and genomic data have revealed important aspects of pandemic spatial spread, several key questions remain: Did the two pandemics follow similar spatial transmission routes? How rapidly did they spread across the United States? What role did stochastic processes play in early spatial transmission? To address these questions, we integrated high-resolution disease data with a robust, data-efficient inference framework combining air travel, commuting flows, and pathogen superspreading potentials to reconstruct their spatial spread across US metropolitan areas. The two pandemics exhibited distinct transmission pathways across locations; however, both pandemics established local circulation in most metropolitan areas within weeks, driven by several shared transmission hubs. Early spatial spread was more strongly associated with air travel than with commuting, though stochastic dynamics introduced substantial uncertainty in transmission routes, creating challenges for timely detection and control. Simulations indicate that broad wastewater surveillance coverage beyond top transmission hubs coupled with effective infection control may slow initial spatial expansion. Our findings highlight the rapid, stochastic spread of pandemic respiratory pathogens and the difficulties of early outbreak containment.

Humans

Molecular epidemiology and phylogeographic architecture of oncogenic intracellular bacteria in cervical cancer patients across Northern China.

BACKGROUND: Oncogenic intracellular bacteria, including Chlamydia trachomatis, Mycoplasma genitalium, and Fusobacterium nucleatum, have emerged as significant contributors to cervical carcinogenesis. Despite growing interest in microbial oncology, the molecular epidemiological landscape and phylogeographic distribution of these pathogens in Northern China remain poorly characterized. This study aimed to determine the prevalence, co-infection patterns, genotypic diversity, and spatial phylogeographic clustering of oncogenic intracellular bacteria among cervical cancer patients across five provinces of Northern China. METHODS: A cross-sectional, multi-center study was conducted between March 2022 and November 2024 across Shaanxi, Heilongjiang, Beijing, Shandong, and Inner Mongolia. Cervical swab specimens were collected from 1247 confirmed cervical cancer patients. Pathogen detection was performed using multiplex real-time polymerase chain reaction, 16S rRNA gene amplicon sequencing, and whole-genome sequencing. Phylogeographic analyses employed maximum likelihood and Bayesian evolutionary inference frameworks. Statistical analyses included multivariate logistic regression and geographic information system-based spatial clustering. RESULTS: The overall prevalence of at least one oncogenic intracellular bacterium was 68.3% (n&#xa0;=&#xa0;852). Chlamydia trachomatis was the most prevalent pathogen detected in 41.2% of participants. Co-infection with two or more bacteria was identified in 29.7% of cases and was independently associated with advanced-stage cervical cancer (adjusted odds ratio&#xa0;=&#xa0;2.87; 95% confidence interval: 1.94 to 4.23; p&#xa0;<&#xa0;0.001). Phylogeographic analysis revealed three distinct molecular clades with evidence of bidirectional gene flow between Shaanxi and Heilongjiang. Whole-genome sequencing identified 14 novel virulence gene variants not previously characterized in Chinese clinical isolates. CONCLUSIONS: Oncogenic intracellular bacteria are highly prevalent and genotypically diverse among cervical cancer patients in Northern China. The identified phylogeographic clustering and novel virulence variants have direct implications for regional screening programs, targeted antimicrobial strategies, and the development of region-specific molecular diagnostic panels.

Cervical cancer

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models

Simultaneous inference for generalized linear models with unmeasured confounders.

Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may be substantially biased. This paper investigates the large-scale hypothesis testing problem for multivariate generalized linear models in the presence of confounding effects. Under arbitrary confounding mechanisms, we propose a unified statistical estimation and inference framework that harnesses orthogonal structures and integrates linear projections into three key stages. It begins by disentangling marginal and uncorrelated confounding effects to recover the latent coefficients. Subsequently, latent factors and primary effects are jointly estimated through lasso-type optimization. Finally, we incorporate projected and weighted bias-correction steps for hypothesis testing. Theoretically, we establish the identification conditions of various effects and non-asymptotic error bounds. We show effective Type-I error control of asymptotic-tests as sample and response sizes approach infinity. Numerical experiments demonstrate that the proposed method controls the false discovery rate by the Benjamini-Hochberg procedure and is more powerful than alternative methods. By comparing single-cell RNA-seq counts from two groups of samples, we demonstrate the suitability of adjusting confounding effects when significant covariates are absent from the model.

Hidden variables

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Popper's philosophy for epidemiologists.

This paper discusses the application of Popper's philosophy to epidemiological research, examining in particular the problems of replication without risk of refutation, of mistaking statistical sophistication for deductive inference, and of dealing with causality at a general level. An example is given of a Popperian approach to the test of a causal hypothesis concerning cancer of the cervix.

Adolescent

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer

TL-HDMR: a transfer learning framework for advancing equitable causal inference reveals metabolic signatures of stroke across multiple ancestries.

The limited genetic diversity in genome-wide association studies (GWAS) poses a significant challenge to the generalizability and equity of biomedical discoveries. Most causal inferences, particularly from high-dimensional phenomes (e.g. metabolomics), are primarily based on European populations, and their applicability to other ancestries remains uncertain. Traditional multivariable Mendelian randomization (MVMR) methods further struggle in high-dimensional and correlated settings due to collinearity and model instability. To bridge this gap, we present a two-step transfer learning framework for high-dimensional MR (TL-HDMR), designed to enhance causal exposure detection in understudied populations. Our approach leverages the Minimax Concave Penalty for asymptotically unbiased estimation amidst exposure correlations. Crucially, we introduce two novel pre-transfer procedures-HDMR.TSD for sourcing beneficial data and HDMR.PRESSO for filtering pleiotropic instruments-to ensure robust knowledge transfer. Extensive simulations demonstrated TL-HDMR's superior performance in ROC curves and mean absolute error over alternative methods. When applied to identify causal metabolites for stroke across multi-ancestry cohorts (European, East Asian, South Asian, and African), TL-HDMR successfully pinpointed both shared and ethnic-specific causal biomarkers, showcasing its unique capability for equitable causal inference. This work provides a powerful statistical tool that not only addresses critical methodological challenges but also promotes inclusivity and fairness in human health research.

Humans

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3&#xb7;8&#x2009;&#xd7;&#x2009;10-10 to 2&#xb7;4&#x2009;&#xd7;&#x2009;10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7&#xb7;87 million to $287&#x2009;000) and 29-fold (from $1&#xb7;98 million to $69&#x2009;100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0&#xb7;01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans

Absolute copy number aware CNV calling of sub-megabase segments in ultra-low coverage single-cell DNA sequencing data.

Recent advances in ultra-low coverage whole-genome sequencing (WGS) of single cells have enabled detailed analysis of copy number variation at a throughput approaching that of single-cell RNA sequencing. However, downstream computational methods have not seen comparable advances and are largely adaptations of deep sequencing methodology with reduced precision. Here, we present ASCENT, a computational method built to take full advantage of modern direct tagmentation-based WGS at ultra-low depth. Using joint segmentation with high-resolution bins, we accurately detect small segments, achieving accurate copy number profiles even at 100 000 reads per cell. ASCENT implements true absolute copy state inference for single cells, based on statistical modeling of coverage rather than comparison to a reference, while taking variable segment copy state into account. Further, ASCENT implements per-segment copy-neutral loss of heterozygosity (LOH) calling without the need for non-tumor or bulk WGS reference. When applied to a pediatric B-ALL sample, ASCENT finds copy-neutral LOH in a small segment and a minor subclone defined by breakpoints missed in bulk WGS. Thus, by applying appropriate computational methods, single-cell WGS provides clear advantages over bulk, even at a relatively low cell number and sequencing depth.

DNA Copy Number Variations