PubMed HealthSearch

Biomedical subjects

Ferhat Ay

Publications and source records attributed to Ferhat Ay.

6 recordsLinked to original sources

Soffritto: a deep learning model for predicting high-resolution replication timing.

MOTIVATION: Replication timing (RT) refers to the order in which DNA loci are replicated during S phase. RT is cell-type specific and implicated in cellular processes including transcription, differentiation, and disease. RT is typically quantified genome-wide using two-fraction assays (e.g. Repli-Seq) which sort cells into early and late S phase fractions followed by DNA sequencing, yielding a ratio as the RT signal. While two-fraction RT data are widely available in multiple cell lines, it is limited in its ability to capture high-resolution RT features. To address this, high-resolution Repli-Seq, which quantifies RT across 16 fractions, was developed, but it is costly and technically challenging with very limited data generated to date. RESULTS: Here, we developed Soffritto, a deep learning model that predicts high-resolution RT data using two-fraction RT data, histone ChIP-seq data, GC content, and gene density as input. Soffritto is composed of a Long Short-Term Memory (LSTM) module and a prediction module. The LSTM module learns long- and short-range interactions between genomic bins, while the prediction module is composed of a fully connected layer that outputs a 16-fraction probability vector for each bin using the LSTM module's embeddings as input. By performing both within cell line and cross-cell line training and testing for five human and mouse cell lines, we show that Soffritto is able to capture experimental 16-fraction RT signals with high accuracy, and the predicted signals allow detection of high-resolution RT patterns. AVAILABILITY AND IMPLEMENTATION: Soffritto is available at https://github.com/ay-lab/Soffritto.

Deep Learning

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

3D genome mapping identifies subgroup-specific chromosome conformations and tumor-dependency genes in ependymoma.

Ependymoma is a tumor of the brain or spinal cord. The two most common and aggressive molecular groups of ependymoma are the supratentorial ZFTA-fusion associated and the posterior fossa ependymoma group A. In both groups, tumors occur mainly in young children and frequently recur after treatment. Although molecular mechanisms underlying these diseases have recently been uncovered, they remain difficult to target and innovative therapeutic approaches are urgently needed. Here, we use genome-wide chromosome conformation capture (Hi-C), complemented with CTCF and H3K27ac ChIP-seq, as well as gene expression and DNA methylation analysis in primary and relapsed ependymoma tumors, to identify chromosomal conformations and regulatory mechanisms associated with aberrant gene expression. In particular, we observe the formation of new topologically associating domains ('neo-TADs') caused by structural variants, group-specific 3D chromatin loops, and the replacement of CTCF insulators by DNA hyper-methylation. Through inhibition experiments, we validate that genes implicated by these 3D genome conformations are essential for the survival of patient-derived ependymoma models in a group-specific manner. Thus, this study extends our ability to reveal tumor-dependency genes by 3D genome conformations even in tumors that lack targetable genetic alterations.

Child

Replication timing networks reveal a link between transcription regulatory circuits and replication timing control.

DNA replication occurs in a defined temporal order known as the replication timing (RT) program and is regulated during development, coordinated with 3D genome organization and transcriptional activity. However, transcription and RT are not sufficiently coordinated to predict each other, suggesting an indirect relationship. Here, we exploit genome-wide RT profiles from 15 human cell types and intermediate differentiation stages derived from human embryonic stem cells to construct different types of RT regulatory networks. First, we constructed networks based on the coordinated RT changes during cell fate commitment to create highly complex RT networks composed of thousands of interactions that form specific functional subnetwork communities. We also constructed directional regulatory networks based on the order of RT changes within cell lineages, and identified master regulators of differentiation pathways. Finally, we explored relationships between RT networks and transcriptional regulatory networks (TRNs) by combining them into more complex circuitries of composite and bipartite networks. Results identified novel trans interactions linking transcription factors that are core to the regulatory circuitry of each cell type to RT changes occurring in those cell types. These core transcription factors were found to bind cooperatively to sites in the affected replication domains, providing provocative evidence that they constitute biologically significant directional interactions. Our findings suggest a regulatory link between the establishment of cell-type-specific TRNs and RT control during lineage specification.

Cell Differentiation

Profiling the long noncoding RNA interaction network in the regulatory elements of target genes by chromatin in situ reverse transcription sequencing.

Long noncoding RNAs (lncRNAs) can regulate the activity of target genes by participating in the organization of chromatin architecture. We have devised a "chromatin-RNA in situ reverse transcription sequencing" (CRIST-seq) approach to profile the lncRNA interaction network in gene regulatory elements by combining the simplicity of RNA biotin labeling with the specificity of the CRISPR/Cas9 system. Using gene-specific gRNAs, we describe a pluripotency-specific lncRNA interacting network in the promoters of Sox2 and Pou5f1, two critical stem cell factors that are required for the maintenance of pluripotency. The promoter-interacting lncRNAs were specifically activated during reprogramming into pluripotency. Knockdown of these lncRNAs caused the stem cells to exit from pluripotency. In contrast, overexpression of the pluripotency-associated lncRNA activated the promoters of core stem cell factor genes and enhanced fibroblast reprogramming into pluripotency. These CRIST-seq data suggest that the Sox2 and Pou5f1 promoters are organized within a unique lncRNA interaction network that determines the fate of pluripotency during reprogramming. This CRIST approach may be broadly used to map lncRNA interaction networks at target loci across the genome.

Animals

17q21 asthma-risk variants switch CTCF binding and regulate IL-2 production by T cells.

Asthma and autoimmune disease susceptibility has been strongly linked to genetic variants in the 17q21 haploblock that alter the expression of ORMDL3; however, the molecular mechanisms by which these variants perturb gene expression and the cell types in which this effect is most prominent are unclear. We found several 17q21 variants overlapped enhancers present mainly in primary immune cell types. CD4+ T cells showed the greatest increase (threefold) in ORMDL3 expression in individuals carrying the asthma-risk alleles, where ORMDL3 negatively regulated interleukin-2 production. The asthma-risk variants rs4065275 and rs12936231 switched CTCF-binding sites in the 17q21 locus, and 4C-Seq assays showed that several distal cis-regulatory elements upstream of the disrupted ZPBP2 CTCF-binding site interacted with the ORMDL3 promoter region in CD4+ T cells exclusively from subjects carrying asthma-risk alleles. Overall, our results suggested that T cells are one of the most prominent cell types affected by 17q21 variants.

Asthma