PubMed Health⌕ Search

Biomedical subjects

Pavel Sumazin

Publications and source records attributed to Pavel Sumazin.

9 recordsLinked to original sources

Asynchronous transitions from high-risk hepatoblastoma to carcinoma.

BACKGROUND & AIMS: Most pediatric hepatocellular tumors are classified as hepatoblastoma (HB) or hepatocellular carcinoma (HCC), yet a subset exhibits mixed histological and molecular features. These hepatoblastomas with carcinoma features (HBCs) include cases provisionally designated as hepatocellular neoplasm-not otherwise specified (HCN-NOS). Their biology remains poorly understood, with unresolved questions about their cellular composition and outcomes. It is unclear whether HBCs comprise hybrid cells with combined HB and HCC characteristics (HBC cells) or admixtures of distinct HB and HCC cells. We characterized the biology, etiology, cellular composition, and evolutionary dynamics of HBCs. METHODS: We performed multi-omics profiling - including single-nucleus RNA sequencing, single-nucleus DNA sequencing, and multi-region longitudinal bulk RNA and DNA sequencing - to characterize HBC composition, evolution, and treatment response. Two-thirds of our samples were post-chemotherapy resections. RESULTS: HBCs comprise heterogeneous mixtures of HB-like, HBC-like, and HCC-like molecular cell types. Outcomes in HBC are significantly worse than in HB, and HBC cells are more chemoresistant than HB cells, with resistance shaped by their cell identity, genetic alterations, and embryonic differentiation stage. HBC cells originate from HB cells that were arrested at early hepatic stem cell development stages because of aberrant WNT signaling activation. Inhibition of WNT signaling promoted differentiation and enhanced sensitivity to chemotherapy. Furthermore, each analyzed HBC reflected a dynamic process of multiple HB-to-HBC and HBC-to-HCC transitions, underscoring their evolutionary complexity. A limitation of our study is our inability to pinpoint the role of chemotherapy-induced genome modifications. CONCLUSIONS: Multi-omics profiling of HBCs revealed key insights into their biology and composition, demonstrating that they originate from HB precursors at early hepatic stem cell development stages and that their differentiation arrest depends on sustained aberrant WNT signaling activity. IMPACT AND IMPLICATIONS: Hepatoblastomas with carcinoma features (HBCs) represent a poorly understood subset of pediatric liver tumors with mixed characteristics of hepatoblastoma (HB) and hepatocellular carcinoma (HCC). Using multi-omics profiling, we show that HBCs comprise heterogeneous mixtures of HB-like, intermediate HBC-like, and HCC-like cell populations that arise from HB precursors arrested at early hepatic stem cell developmental stages due to aberrant WNT signaling. This differentiation arrest contributes to chemoresistance and poorer clinical outcomes compared with HB. Importantly, pharmacologic inhibition of WNT signaling promoted differentiation and increased chemotherapy sensitivity, suggesting a potential therapeutic strategy. These findings refine the biological classification of HBCs and highlight differentiation-based treatment approaches for this aggressive tumor subtype.

Multiomics↗

Tissue-specific regulatory elements in mammalian promoters.

Transcription factor-binding sites and the cis-regulatory modules they compose are central determinants of gene expression. We previously showed that binding site motifs and modules in proximal promoters can be used to predict a significant portion of mammalian tissue-specific transcription. Here, we report on a systematic analysis of promoters controlling tissue-specific expression in heart, kidney, liver, pancreas, skeletal muscle, testis and CD4 T cells, for both human and mouse. We integrated multiple sources of expression data to compile sets of transcripts with strong evidence for tissue-specific regulation. The analysis of the promoters corresponding to these sets produced a catalog of predicted tissue-specific motifs and modules, and cis-regulatory elements. Predicted regulatory interactions are supported by statistical evidence, and provide a foundation for targeted experiments that will improve our understanding of tissue-specific regulatory networks. In a broader context, methods used to construct the catalog provide a model for the analysis of genomic regions that regulate differentially expressed genes.

Algorithms↗

DNA motifs in human and mouse proximal promoters predict tissue-specific expression.

Comprehensive identification of cis-regulatory elements is necessary for accurately reconstructing gene regulatory networks. We studied proximal promoters of human and mouse genes with differential expression across 56 terminally differentiated tissues. Using in silico techniques to discover, evaluate, and model interactions among sequence elements, we systematically identified regulatory modules that distinguish elevated from inhibited expression in the corresponding transcripts. We used these putative regulatory modules to construct a single predictive model for each of the 56 tissues. These predictors distinguish tissue-specific elevated from inhibited expression with statistical significance in 80% of the tissues (45 of 56). The predictors also reveal synergy between cis-regulatory modules and explain large-scale tissue-specific differential expression. For testis and liver, the predictors include computationally predicted motifs. For most other tissues, the predictors reveal synergy between experimentally verified motifs and indicate genes that are regulated by similar tissue-specific machinery. The identification in proximal promoters of cis-regulatory modules with tissue-specific activity lays the groundwork for complete characterization and deciphering of cis-regulatory DNA code in mammalian genomes.

Animals↗

Identifying tissue-selective transcription factor binding sites in vertebrate promoters.

We present a computational method aimed at systematically identifying tissue-selective transcription factor binding sites. Our method focuses on the differences between sets of promoters that are associated with differentially expressed genes, and it is effective at identifying the highly degenerate motifs that characterize vertebrate transcription factor binding sites. Results on simulated data indicate that our method detects motifs with greater accuracy than the leading methods, and its detection of strongly overrepresented motifs is nearly perfect. We present motifs identified by our method as the most overrepresented in promoters of liver- and muscle-selective genes, demonstrating that our method accurately identifies known transcription factor binding sites and previously uncharacterized motifs.

Animals↗

Mining ChIP-chip data for transcription factor and cofactor binding sites.

MOTIVATION: Identification of single motifs and motif pairs that can be used to predict transcription factor localization in ChIP-chip data, and gene expression in tissue-specific microarray data. RESULTS: We describe methodology to identify de novo individual and interacting pairs of binding site motifs from ChIP-chip data, using an algorithm that integrates localization data directly into the motif discovery process. We combine matrix-enumeration based motif discovery with multivariate regression to evaluate candidate motifs and identify motif interactions. When applied to the HNF localization data in liver and pancreatic islets, our methods produce motifs that are either novel or improved known motifs. All motif pairs identified to predict localization are further evaluated according to how well they predict expression in liver and islets and according to how conserved are the relative positions of their occurrences. We find that interaction models of HNF1 and CDP motifs provide excellent prediction of both HNF1 localization and gene expression in liver. Our results demonstrate that ChIP-chip data can be used to identify interacting binding site motifs. AVAILABILITY: Motif discovery programs and analysis tools are available on request from the authors.

Algorithms↗

DWE: discriminating word enumerator.

MOTIVATION: Tissue-specific transcription factor binding sites give insight into tissue-specific transcription regulation. RESULTS: We describe a word-counting-based tool for de novo tissue-specific transcription factor binding site discovery using expression information in addition to sequence information. We incorporate tissue-specific gene expression through gene classification to positive expression and repressed expression. We present a direct statistical approach to find overrepresented transcription factor binding sites in a foreground promoter sequence set against a background promoter sequence set. Our approach naturally extends to synergistic transcription factor binding site search. We find putative transcription factor binding sites that are overrepresented in the proximal promoters of liver-specific genes relative to proximal promoters of liver-independent genes. Our results indicate that binding sites for hepatocyte nuclear factors (especially HNF-1 and HNF-4) and CCAAT/enhancer-binding protein (C/EBPbeta) are the most overrepresented in proximal promoters of liver-specific genes. Our results suggest that HNF-4 has strong synergistic relationships with HNF-1, HNF-4 and HNF-3beta and with C/EBPbeta. AVAILABILITY: Programs are available for use over the Web at http://rulai.cshl.edu/tools/dwe.

Algorithms↗

Similarity of position frequency matrices for transcription factor binding sites.

MOTIVATION: Transcription-factor binding sites (TFBS) in promoter sequences of higher eukaryotes are commonly modeled using position frequency matrices (PFM). The ability to compare PFMs representing binding sites is especially important for de novo sequence motif discovery, where it is desirable to compare putative matrices to one another and to known matrices. RESULTS: We describe a PFM similarity quantification method based on product multinomial distributions, demonstrate its ability to identify PFM similarity and show that it has a better false positive to false negative ratio compared to existing methods. We grouped TFBS frequency matrices from two libraries into matrix families and identified the matrices that are common and unique to these libraries. We identified similarities and differences between the skeletal-muscle-specific and non-muscle-specific frequency matrices for the binding sites of Mef-2, Myf, Sp-1, SRF and TEF of Wasserman and Fickett. We further identified known frequency matrices and matrix families that were strongly similar to the matrices given by Wasserman and Fickett. We provide methodology and tools to compare and query libraries of frequency matrices for TFBSs. AVAILABILITY: Software is available to use over the Web at http://rulai.cshl.edu/MatCompare SUPPLEMENTARY INFORMATION: Database and clustering statistics, matrix families and representatives are available at http://rulai.cshl.edu/MatCompare/Supplementary.

Algorithms↗

Genome-wide prediction and analysis of function-specific transcription factor binding sites.

DNA-binding transcription factors play a central role in transcription regulation, and the annotation of transcription-factor binding sites in upstream regions of human genes is essential for building a genome-wide regulatory network. We describe methodology to accurately predict the transcription-factor binding sites in the proximal-promoter region of function-specific genes. In order to increase the accuracy of transcription factor binding-site prediction, we rely on recent genome sequence data, known transcription factor binding-site matrices, and Gene Ontology biological-function-based gene classification. Using TRANSFAC position-frequency matrices, we detected individual and cooperating transcription-factor binding sites in proximal promoters of ENSEMBL annotated human genes. We used the over representation of detected binding sites in the proximal promoters as compared to the second exons to control specificity. We confirmed the majority of transcription-factor binding sites predicted in proximal promoters of immune-response genes with evidence from existing literature. We validated the predicted cooperation between transcription factors NF-kappa B and IRF in the regulation of gene expression with microarray transcript profiling data and literature-derived protein-protein interaction network. We also identified over-represented individual and pairs of transcription-factor binding sites in the proximal promoters of each Gene Ontology biological-process gene group. Our tools and analysis provide a new resource for deciphering transcription regulation in different biological paradigms.

Computational Biology↗

Deconvolving sequence variation in mixed DNA populations.

We present an original approach to identifying sequence variants in a mixed DNA population from sequence trace data. The heart of the method is based on parsimony: given a wildtype DNA sequence, a set of observed variations at each position collected from sequencing data, and a complete catalog of all possible mutations, determine the smallest set of mutations from the catalog that could fully explain the observed variations. The algorithmic complexity of the problem is analyzed for several classes of mutations, including block substitutions, single-range deletions, and single-range insertions. The reconstruction problem is shown to be NP-complete for single-range insertions and deletions, while for block substitutions, single character insertion, and single character deletion mutations, polynomial time algorithms are provided. Once a minimum set of mutations compatible with the observed sequence is found, the relative frequency of those mutations is recovered by solving a system of linear equations. Simulation results show the algorithm successfully deconvolving mutations in p53 known to cause cancer. An extension of the algorithm is proposed as a new method of high throughput screening for single nucleotide polymorphisms by multiplexing DNA.

Algorithms↗