PubMed HealthSearch

Biomedical subjects

Valentina Boeva

Publications and source records attributed to Valentina Boeva.

2 recordsLinked to original sources

Chiron3D: an interpretable deep learning framework for understanding the DNA code of chromatin looping.

MOTIVATION: Three-dimensional folding of the genome into structures such as chromatin loops is essential for gene regulation. Current experimental methods for mapping these structures, like Hi-C and HiChIP, are labor-intensive and require repeated assays to test hypothesized mutation effects. This motivates the need for predictive approaches that reveal the sequence determinants of chromatin loops. RESULTS: In this work, we present a novel and interpretable computational pipeline for predicting CTCF-mediated chromatin loops. We propose Chiron3D, a DNA-only model trained in a cell-type specific manner to predict CTCF HiChIP contact maps. By leveraging pre-trained embeddings from a foundation model, our approach is competitive with baselines that take CTCF ChIP-seq as additional input, while enabling nucleotide-level attribution to the input DNA sequence. Using our framework, we provide likely mechanistic insights into the physical control of loop dynamics. Specifically, we find that the strength of the loop extrusion anchorage site is largely governed by the amount and binding affinity of CTCF sites at the boundaries. Furthermore, we reveal that loop stability is regulated by the amount of intra-loop CTCF binding sites, where fewer intra-loop sites are associated with greater loop stability. Using targeted, single-nucleotide edit simulations with Chiron3D, we show that both loop strength and stability can be precisely controlled. Together, these results provide novel mechanistic insights into the physical control of genome organization and highlight the potential of decoding the DNA sequence logic in silico. AVAILABILITY: The Chiron3D pipeline is made available at https://github.com/BoevaLab/Chiron3D.

Chromatin

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2 µm × 2 µm resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts