PubMed HealthSearch

SEARCH · PubMed Health

Results for “replication compartments”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Proviral functions of HMGB1 in HAdV-C5 replication compartments.

UNLABELLED: Human adenoviruses (HAdVs) induce significant reorganization of the nuclear environment, leading to the formation of virus-induced subnuclear structures known as replication compartments (RCs). Within these RCs, viral genome replication, gene expression, and modulation of cellular antiviral responses are tightly coordinated, making them valuable models for studying virus-host interactions. In a recent study, we analyzed the protein composition of HAdV type 5 (HAdV-C5) RCs isolated from infected primary cells at different time points during infection using quantitative proteomics. We identified several chromatin modifiers, including the high-mobility group box 1 protein (HMGB1) as components associated with RCs and demonstrated that HMGB1 can be relocalized to RCs from different HAdV species, thereby modulating viral replication in a species-specific manner. In the present work, using click-chemistry and proximity ligation assays, we discovered that HMGB1 localizes to sites of DNA replication within RCs and that its interaction with DBP in RCs is dependent on both DNA replication and RC assembly. HMGB1-knockdown experiments demonstrated that HMGB1 is required for efficient viral gene expression. However, despite its proviral role in viral replication, we found that HMGB1 levels decreased in late stages of infection due to transcriptional downregulation. Furthermore, by overexpressing HMGB1, we showed that this regulation of HMGB1 levels during infection is critical for optimal HAdV-C5 replication. These results highlight the complex regulatory relationship between HMGB1 and HAdV-C5 infection. IMPORTANCE: In an extensive proteomics analysis, we found that HMGB1, an important cellular chromatin protein, was enriched in adenovirus replication compartments. In this study, we aimed to better understand the role of HMGB1 in the infection process of a human DNA virus, HAdV-C5. We tested different virus types, including some with specific gene deletions and mutations. Our results showed that during infection, HMGB1 levels decreased because the virus suppressed its production. Despite this, even at lower levels, HMGB1 still helped the virus replicate by interacting with key viral proteins and DNA at sites where the virus is actively replicating. Overall, our findings highlight how HMGB1 plays a crucial role in facilitating efficient virus replication, making it an important factor in the infection process.

HMGB1 Protein

An essential and highly selective protein import pathway encoded by nucleus-forming phage.

UNLABELLED: Targeting proteins to specific subcellular destinations is essential in prokaryotes, eukaryotes, and the viruses that infect them. Chimalliviridae phages encapsulate their genomes in a nucleus-like replication compartment composed of the protein chimallin (ChmA) that excludes ribosomes and decouples transcription from translation. These phages selectively partition proteins between the phage nucleus and the bacterial cytoplasm. Currently, the genes and signals that govern selective protein import into the phage nucleus are unknown. Here we identify two components of this novel protein import pathway: a species-specific surface-exposed region of a phage intranuclear protein required for nuclear entry and a conserved protein, PicA, that facilitates cargo protein trafficking across the phage nuclear shell. We also identify a defective cargo protein that is targeted to PicA on the nuclear periphery but fails to enter the nucleus, providing insight into the mechanism of nuclear protein trafficking. Using CRISPRi-ART protein expression knockdown of PicA, we show that PicA is essential early in the chimallivirus replication cycle. Together our results allow us to propose a multistep model for the Protein Import Chimallivirus (PIC) pathway, where proteins are targeted to PicA by amino acids on their surface, and then licensed by PicA for nuclear entry. The divergence in the selectivity of this pathway between closely-related chimalliviruses implicates its role as a key player in the evolutionary arms race between competing phages and their hosts. SIGNIFICANCE STATEMENT: The phage nucleus is an enclosed replication compartment built by Chimalliviridae phages that, similar to the eukaryotic nucleus, separates transcription from translation and selectively imports certain proteins. This allows the phage to concentrate proteins required for DNA replication and transcription while excluding DNA-targeting host defense proteins. However, the mechanism of selective trafficking into the phage nucleus is currently unknown. Here we determine the region of a phage nuclear protein that targets it for nuclear import and identify a conserved, essential nuclear shell-associated protein that plays a key role in this process. This work provides the first mechanistic model of selective import into the phage nucleus.

Preprint

Viral replication through phase separation: Cytosolic and nuclear condensates.

Replication of many RNA and DNA viruses occurs within specialized intracellular hubs organized as membraneless biomolecular condensates (BCs) driven by liquid-liquid phase separation. As obligate intracellular parasites, viruses depend on the host cell machinery to complete their replication cycles and therefore actively remodel the intracellular environment to favor viral genome replication, transcription, and assembly. Cytosolic and nuclear phase-separated replication compartments (RC) provide concentrated and dynamic platforms that promote efficient interactions between viral genomes and viral or host proteins essential for infection. The formation of viral replication BCs is typically facilitated by viral proteins enriched in intrinsically disordered regions and low-complexity domains, which enable multivalent interactions with viral nucleic acids and cellular factors. These interactions are mediated by diverse biophysical forces, including hydrophobic and π interactions, hydrogen bonding, molecular crowding, and osmotic effects. Throughout infection, viral BCs remain highly dynamic, allowing continuous exchange of components and functional maturation of replication hubs. Their properties and activities are further regulated by post-translational modifications of viral and host proteins, such as phosphorylation, acetylation, and methylation. In this review, we summarize current evidence supporting liquid-liquid phase separation as a central organizing principle of viral RCs. We focus on representative RNA and DNA viruses that replicate in the cytosol or nucleus, highlighting virus-specific strategies, conserved mechanisms, and the consequences of BC formation for viral replication efficiency, host antiviral responses, and therapeutic intervention.

Phase Separation

An essential and highly selective protein import pathway encoded by nucleus-forming phage.

Targeting proteins to specific subcellular destinations is essential in prokaryotes, eukaryotes, and the viruses that infect them. Chimalliviridae phages encapsulate their genomes in a nucleus-like replication compartment composed of the protein chimallin (ChmA) that excludes ribosomes and decouples transcription from translation. These phages selectively partition proteins between the phage nucleus and the bacterial cytoplasm. Currently, the genes and signals that govern selective protein import into the phage nucleus are unknown. Here, we identify two components of this protein import pathway: a species-specific surface-exposed region of a phage intranuclear protein required for nuclear entry and a conserved protein, PicA (Protein importer of chimalliviruses A), that facilitates cargo protein trafficking across the phage nuclear shell. We also identify a defective cargo protein that is targeted to PicA on the nuclear periphery but fails to enter the nucleus, providing insight into the mechanism of nuclear protein trafficking. Using CRISPRi-ART protein expression knockdown of PicA, we show that PicA is essential early in the chimallivirus replication cycle. Together, our results allow us to propose a multistep model for the Protein Import Chimallivirus pathway, where proteins are targeted to PicA by amino acids on their surface and then licensed by PicA for nuclear entry. The divergence in the selectivity of this pathway between closely related chimalliviruses implicates its role as a key player in the evolutionary arms race between competing phages and their hosts.

Viral Proteins

Spatial Mapping and Interactome Profiling of m6A-Modified R-Loops via Chemically Inducible Split-APEX2 Proximity Labeling.

m6A-Modified R-loops (m6A-R-loops) play crucial roles in epigenetic regulation and genome stability, yet resolving their spatial distribution and protein interactomes in live cells remains challenging. To address this, we developed m6A-R-loop proximity labeling (m6A-RLPL), a chemically inducible split-APEX2 proximity labeling technology integrating dual-target recognition using the RNA-DNA hybrid binding domain of RNase H1 for R-loop targeting and m6A reader protein's YTH domain for m6A recognition, coupled with an abscisic acid (ABA)-inducible dimerization system for signal amplification. This technology revealed host m6A-R-loops enriched with nucleoli under normal conditions. When applied to herpes simplex virus (HSV) infection, it further demonstrated viral m6A-R-loops undergoing dramatic accumulation within phase-separated granules in replication compartments during late-stage infection. Proximity proteomics identified ZC3H4 and CCDC124 as essential regulators maintaining these structures, which serve as transcription sites for HSV late genes, with disruption selectively impairing viral transcription. m6A-RLPL establishes a generalizable approach for spatially resolved profiling of m6A-R-loop interactomes and organizational dynamics in living systems.

Humans

PARTAGE: Parallel analysis of replication timing and gene expression.

The human genome is partitioned into functional compartments that replicate at specific times during the S-phase. This temporal program, referred to as replication timing (RT), is co-regulated with the 3D genome organization, is cell type-specific, and changes during development in coordination with gene expression. Moreover, RT alterations are linked to abnormal gene expression, genome instability, and structural variation in multiple diseases, including cancer. However, mechanistic links between RT, large-scale 3D genome architecture, and transcriptional regulation remain poorly understood. A major limitation is that current approaches require the separate profiling of RT and transcriptomes from independent batches of samples, obscuring the complex co-regulation between the epigenome and transcriptome. Here, we developed PARTAGE, a multiomics approach that enables joint profiling of copy number variation (CNV), RT, and gene expression from the same sample, providing a more accurate integrative view of the complex relationships between RT and gene regulation.

Journal Article

DciA, the Bacterial Replicative Helicase Loader, Promotes LLPS in the Presence of ssDNA.

The loading of the bacterial replicative helicase DnaB is an essential step for genome replication and depends on the assistance of accessory proteins. Several of these proteins have been identified across the bacterial phyla. DciA is the most common loading protein in bacteria, yet the one whose mechanism is the least understood. We have previously shown that DciA from Vibrio cholerae is composed of a globular domain followed by an unfolded extension and demonstrated its strong affinity for DNA. Here, we characterize the condensates formed by VcDciA upon interaction with a short single-stranded DNA substrate. We demonstrate the fluidity of these condensates using light microscopy and address their network organization through electron microscopy, thereby bridging events to conclude on a liquid-liquid phase separation behavior. Additionally, we observe the recruitment of DnaB in the droplets, concomitant with the release of DciA. We show that the well-known helicase loader DnaC from Escherichia coli is also competent to form these phase-separated condensates in the presence of ssDNA. Our phenomenological data are still preliminary as regards the existence of these condensates in vivo, but open the way for exploring the potential involvement of DciA in the formation of non-membrane compartments within the bacterium to facilitate the assembly of replication players on chromosomal DNA.

DNA, Single-Stranded

Maintenance of nucleosome organization through replication and transcription counteracts aberrant coalescence of active chromatin.

Nucleosomes with their associated modifications organize and regulate the genome. It is unclear how this is integrated with the requirement of replication and transcription to access the DNA template without jeopardizing chromatin function. Here, we reveal a unified requirement for the histone chaperone FACT in mediating nucleosome disruption and reassembly during mammalian replication and transcription. Upon acute FACT depletion, replisome and RNA polymerase progression is halted genome wide, and chromatin structure in their wake collapses, with reduced nucleosome occupancy, irregular spacing, and intermediate assemblies. Chromatin states deteriorate as modified histones are lost due to a lack of histone recycling. Chromatin fiber disorder further manifests in the 3D genome, triggering active genes to coalesce in aberrant microcompartments. Similarly, aberrant compartments form in cells failing to maintain chromatin fiber structure through replication. Nucleosome organization therefore dynamically regulates genome architecture, guarding against spurious chromatin aggregation.

Nucleosomes

A reproducible computational transcriptomic framework for cell-type-resolved fibroinflammatory-AKT remodeling in human heart failure.

BACKGROUND: Human heart failure involves multicellular transcriptional remodeling, but public transcriptomic studies often remain disconnected from cell-type localization and perturbational interpretation. METHODS: We developed a reproducible computational workflow integrating human left-ventricular bulk transcriptomes, donor-level cell-type pseudobulk results from a human heart-failure single-cell/single-nucleus atlas, external snRNA-seq support, curated module scoring, focused ligand-receptor prioritization and LINCS/L1000 perturbational matching. RESULTS: Cross-cohort analysis identified 14,358 same-direction HF-associated genes, including 1633 replicated HF-up and 785 replicated HF-down genes. Donor-level pseudobulk analysis localized disease remodeling to cardiomyocyte, fibroblast and myeloid compartments. Activated fibroblast and inflammatory myeloid programs defined a fibroinflammatory remodeling axis connected to context-dependent AKT-associated transcriptional shifts. External snRNA-seq support was strongest for fibroblast activation and AKT-associated remodeling, with etiology-dependent heterogeneity across validation resources. L1000FWD screening prioritized safety-aware perturbational hypotheses, including glimepiride and simvastatin as interpretable candidates requiring experimental validation. CONCLUSIONS: This study provides a computational transcriptomic framework linking reproducible human HF signatures, cell-type-resolved fibroinflammatory remodeling and perturbational genomic prioritization without claiming drug efficacy or AKT causality.

Humans

SARS-CoV-2 Orf3a protein interaction mapping using unnatural amino acid incorporation.

Mapping transient protein-protein interactions remain a major challenge in studying viral host-pathogen interfaces. While some virus-host interactions are stable and readily captured, the majority are highly dynamic, reflecting the need for viral proteins to engage distinct host factors at different stages of the life cycle. Here, we employ a protein engineering strategy based on the site-specific incorporation of the unnatural acid p-azido-L-phenylalanine (AzF) to enable photo-crosslinking proteomic analysis of the SARS-CoV-2 accessory protein Orf3a in live cells. Genetic installation of AzF at residue K198 of Orf3a permitted UV-induced covalent capture of proximal host interacting proteins, overcoming challenges associated with membrane localization and limited protein abundance. A total of 248 high-confidence Orf3a-interacting proteins were reproducibly identified and subjected to gene ontology analysis, revealing enrichment in innate immune signaling, antiviral defense, RNA processing, and viral replication-associated pathways. Orf3a is an accessory protein that functions as a viroporin and traffics across multiple cellular compartments, and was found to interact with host RNA helicases, RNA-binding proteins, immune regulators, and metabolic enzymes implicated in SARS-CoV-2 infection. Together, these results demonstrate that genetically encoded, site-specific photo-crosslinking enables selective capture of transient interactions that are often missed by nonspecific 254 nm UV crosslinking approaches and highlights Orf3a as a multifunctional protein that engages diverse host pathways. More broadly, this study establishes a generalizable framework for leveraging unnatural amino acid-based protein engineering approaches to interrogate dynamic host-pathogen interactions.

Humans

The brain as an HIV reservoir: Recent findings using autopsy tissues from people with HIV.

HIV persistence within anatomical reservoirs remains the primary barrier to achieving an HIV cure. While antiretroviral therapy effectively suppresses plasma viremia, it does not eliminate integrated proviral genomes that persist in long-lived cellular compartments. The central nervous system (CNS) is a clinically important HIV reservoir, characterized by immune privilege and the persistence of tissue-resident infection despite effective antiretroviral therapy (ART). Evidence from postmortem studies reveals that HIV DNA, RNA, and even intact replication-competent proviruses remain detectable in brain tissue from virally suppressed people with HIV. Evidence derived primarily from in situ approaches and viable-cell studies supports myeloid-lineage reservoirs, particularly microglia and CNS-associated macrophages, as key cellular sources of persistence, while the extent and biological relevance of astrocyte infection remains debated. These reservoirs exhibit transcriptional activity and are associated with chronic neuroinflammation, which may contribute to HIV-associated neurocognitive disorders, despite systemic viral suppression. Here, we synthesize recent findings from autopsy brain studies, including work enabled by major biorepositories, such as the National NeuroHIV Tissue Consortium and rapid-autopsy programs, including the Last Gift, both of which are essential for studying HIV reservoirs in the CNS. We summarize methodologies for detecting and characterizing HIV in brain tissue, highlight heterogeneous patterns of regional distribution and compartmentalization, and review emerging links between CNS persistence and neuroinflammation. We conclude with priorities for harmonized tissue processing, multi-modal single-cell and spatial profiling, and coordinated cross-cohort analyses to clarify the contribution of CNS reservoirs to neuroHIV pathogenesis and systemic rebound.

Humans

Architectural logic of the 3D genome: mechanisms of dysregulation and emerging cancer therapeutics.

The three-dimensional (3D) genome provides an essential layer of organization that shapes genome function in space and time. Chromatin compartments and topologically associating domains (TADs) arise from the interplay between intrinsic properties of chromatin and architectural factors, including cohesin and CTCF. Despite substantial progress in defining these structural features, whether 3D genome architecture plays a causal role in regulating processes such as transcription, DNA replication, and DNA repair, or instead reflects underlying regulatory activity, remains unresolved. Here, we use the distinction between chromatin-intrinsic features and architectural factors as a framework to evaluate evidence for causality in genome structure-function relationships. We extend this framework to cancer, where both intrinsic alterations (including noncoding mutations, structural variants, and changes in chromatin state) and architectural factor perturbations (such as mutations in architectural proteins and dysregulation of transcriptional machinery) disrupt genome organization and contribute to disease progression. These findings suggest that alterations in genome structure can, in some contexts, actively reshape oncogenic programs. A major limitation in applying 3D genome insights to cancer biology is the cost and complexity of omics assays. Recent advances in artificial intelligence (AI) and machine learning (ML) enable inference and prediction of 3D genome organization from sequence and epigenomic features, providing insight into the extent to which genome folding is encoded intrinsically versus dynamically regulated in architectural factors. This perspective provides a unified view of how genome structure is established, how it relates to function, and how its disruption contributes to tumorigenesis.

3D genome

Epigenetic alterations in rheumatoid arthritis: multilayer mechanisms and translational opportunities.

Rheumatoid arthritis (RA) is a chronic inflammatory disease driven by immune dysregulation, in which genetic susceptibility and environmental exposures promote persistent synovitis, progressive joint damage, and systemic comorbidities. Recent epigenomic studies show several recurring abnormalities. Many RA susceptibility variants lie outside protein-coding sequence and map to immune-cell and synovial fibroblast regulatory elements, linking inherited risk to enhancer activity, methylation quantitative trait effects, and distal gene control. Blood-based epigenome-wide association studies identify disease-associated DNA methylation signatures, but these signals require careful control for leukocyte composition, smoking, treatment exposure, and disease stage. RA fibroblast-like synoviocytes also display stable methylome remodeling, including relative hypomethylation at loci involved in inflammation, migration, matrix degradation, and apoptosis resistance, while TET3-associated 5-hydroxymethylcytosine has emerged as a functional contributor to chemokine production and invasive stromal behavior. Histone modifications, chromatin accessibility, and 3D genome organization define pathogenic regulatory states and connect non-coding risk loci to effector genes in immune and stromal compartments. Finally, miRNAs, lncRNAs, circRNAs, snoRNAs, extracellular RNAs, and m6A-related pathways add post-transcriptional and chromatin-linked layers with potential biomarker value. We synthesize these findings and discuss translational opportunities for diagnosis, stratification, flare monitoring, and therapeutic targeting, while emphasizing incomplete replication, uneven evidence across epigenetic layers, biospecimen variability, and the need for causal, longitudinal, cell-type-resolved validation.

Humans

Rab10 coordinates SADS-CoV non-lytic egress through the ERGIC-TGN-lysosome trafficking pathway.

Swine acute diarrhea syndrome coronavirus (SADS-CoV) is a bat-originated alphacoronavirus that causes devastating enteric disease in neonatal piglets and possesses significant potential for cross-species transmission. While the early stages of the coronavirus life cycle have been extensively characterized, the host factors indispensable for virion assembly and subsequent export remain largely enigmatic. Here, by performing a genome-wide CRISPR-Cas9 knockout screen using a recombinant icSADS-CoV-GFP reporter virus, we identified the small GTPase Rab10 as a critical host dependency factor for SADS-CoV infection. Viral life cycle analysis revealed that Rab10 is not required for viral attachment, entry, or initial genome replication, but is essential for the virion transport and non-lytic egress. Rab10 deficiency markedly reduced the extracellular release of viral RNA, viral proteins, and infectious progeny, as well as the secretion of SADS-CoV virus-like particles. Confocal imaging showed that Rab10 and viral protein-positive intracellular structures were associated with LMAN1, TGN46, and LAMP1 positive compartments. These findings support a model in which Rab10 coordinates a virus-containing vesicles trafficking pathway associated with ERGIC-TGN-lysosome compartments. Mechanistically, Rab10 facilitates the loading of the viral envelope (E) protein into transport vesicles derived from the ERGIC. Rab10 associates with the SADS-CoV E protein, and mapping analyses implicated the C-terminal PDZ-binding motif, particularly residue V75, in efficient Rab10 association and viral release. Collectively, our findings identify Rab10 as a host regulator of SADS-CoV non-lytic egress and highlight the E-Rab10 interaction and the vesicular trafficking machinery as a potential target for developing antiviral strategies.

Animals

MIA-Jet: Multi-scale Identification Algorithm of Chromatin Jets.

The mammalian genome is organized into large-scale chromosome territories, compartments, domains, and at the smallest scale, chromatin loops and stripes. The newest element is a chromatin jet, a diffused line perpendicular to the main diagonal in the Hi-C contact map, which was reported in quiescent mammalian lymphocytes supporting a two-sided symmetric cohesin loop extrusion model. A similar structure is observed in Repli-HiC data, where relatively thin and straight chromatin fountains indicate coupling of DNA replication forks. However, the precise biological implications of these jet-like structures are unknown due to the limitations in computational methods. We developed MIA-Jet, a multi-scale ridge detection algorithm that can accurately detect jets of variable lengths, widths, and angles. When tested on Hi-C, Repli-HiC, ChIA-PET, ChIA-Drop, and Micro-C data in mouse, human, roundworm, and zebrafish cells, MIA-Jet outperformed existing methods. In human cells, jets were enriched in cohesin loading sites and early replication initiation zones. Applying MIA-Jet to Hi-C data generated from protein-degraded cells revealed that jets are dependent on cohesin but not YY1, and jet signals are strengthened after depleting WAPL. We envision MIA-Jet to be broadly applicable to any 3D genome mapping data, thereby providing new insights into the functional roles of chromatin jets.

3D genome mapping

Fetal signatures in the 3D genome of iPSC-derived neurons and their implications for disease modeling.

Induced pluripotent stem cells (iPSCs) have revolutionized neuroscience, providing an approach to generate patient-specific neurons for modeling of neurological diseases. However, it remains unclear how closely iPSC-derived neurons replicate the chromatin architecture of authentic brain neurons. Here, we uniformly processed newly generated Hi-C data from iPSC-derived neurons and neurons isolated from the human postmortem brain, together with previously published data sets comprising 228 human and 89 mouse Hi-C and snm3C-seq samples from different cell subtypes. These data were merged into 96 high-coverage contact maps used to examine chromatin features ranging from chromatin compartments and topologically associating domains (TADs) to chromatin loops, Polycomb-mediated contacts, and frequently interacting regions (FIREs). We find that iPSC-derived neurons largely retain the chromatin state of undifferentiated cells and resemble fetal rather than mature neurons. iPSC-derived neurons exhibit unusually strong compartmentalization, an enrichment of developmental genes at TAD borders, and a marked reduction of long-range repressive Polycomb-mediated contacts that typically silence early fetal programs. Although immature, iPSC-derived neurons offer advantages for modeling interactions between disease-associated SNPs and target genes, as many psychiatric disorders have neurodevelopmental origins. Integrating iPSC-derived and postmortem neuronal data sets therefore provides complementary insights into the chromatin landscape underlying disease-associated interactions. Our study offers a valuable Hi-C resource for the community and provides a detailed comparison of chromatin architecture throughout neuronal maturation, underscoring its importance for validating neuronal models and providing a robust framework for future studies.

Journal Article

Deciphering the ghost proteome in ovarian cancer cells by deep proteogenomic characterization.

Proteogenomics is becoming a powerful tool in personalized medicine by linking genomics, transcriptomics and mass spectrometry (MS)-based proteomics. Due to increasing evidence of alternative open reading frame-encoded proteins (AltProts), proteogenomics has a high potential to unravel the characteristics, variants, expression levels of the alternative proteome, in addition to already annotated proteins (RefProts). To obtain a broader view of the proteome of ovarian cancer cells compared to ovarian epithelial cells, cell-specific total RNA-sequencing profiles and customized protein databases were generated. In total, 128 RefProts and 30 AltProts were identified exclusively in SKOV-3 and PEO-4 cells. Among them, an AltProt variant of IP_715944, translated from DHX8, was found mutated (p.Leu44Pro). We show high variation in protein expression levels of RefProts and AltProts in different subcellular compartments. The presence of 117 RefProt and two AltProt variants was described, along with their possible implications in the different physiological/pathological characteristics. To identify the possible involvement of AltProts in cellular processes, cross-linking-MS (XL-MS) was performed in each cell line to identify AltProt-RefProt interactions. This approach revealed an interaction between POLD3 and the AltProt IP_183088, which after molecular docking, was placed between POLD3-POLD2 binding sites, highlighting its possibility of the involvement in DNA replication and repair.

Humans

Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC.

Three-dimensional nuclear DNA architecture comprises well-studied intra-chromosomal (cis) folding and less characterized inter-chromosomal (trans) interfaces. Current predictive models of 3D genome folding can effectively infer pairwise cis-chromatin interactions from the primary DNA sequence but generally ignore trans contacts. There is an unmet need for robust models of trans-genome organization that provide insights into their underlying principles and functional relevance. We present TwinC, an interpretable convolutional neural network model that reliably predicts trans contacts measurable through proximity ligation-dependent (in situ and intact Hi-C) and independent (DNA SPRITE) genome-wide chromatin conformation assays. . TwinC uses a paired sequence design from replicate Hi-C experiments to learn single base pair relevance in trans interactions across two stretches of DNA. The method achieves high predictive accuracy (AUROC=0.80) on a cross-chromosomal test set from in situ and intact Hi-C experiments in heart tissue. Furthermore, we train TwinC using in situ Hi-C data from the widely used GM12878 cell line and validate its performance with orthogonal DNA SPRITE assay in the same cell type. Mechanistically, the neural network learns the importance of compartments, chromatin accessibility, clustered transcription factor binding and G-quadruplexes in forming trans contacts. In summary, TwinC models and interprets trans genome architecture, shedding light on this poorly understood aspect of gene regulation.

Journal Article