PubMed Health⌕ Search

Biomedical subjects

Chaolin Zhang

Publications and source records attributed to Chaolin Zhang.

6 recordsLinked to original sources

Exon inclusion signatures enable accurate estimation of splicing factor activity.

Splicing factors control exon inclusion in messenger RNAs, shaping transcriptome and proteome diversity. Their catalytic activity is regulated by multiple layers, making single-omic measurements on their own fall short in identifying which splicing factors underlie a phenotype. Here, we posit that splicing factor activity can be estimated from changes in exon inclusion. To test this hypothesis, we benchmarked methods for constructing splicing factor→exon networks and estimating splicing factor activity. We found that combining RNA-seq perturbation-based networks with VIPER (Virtual Inference of Protein Activity by Enriched Regulon analysis) accurately captures splicing factor activation as modulated by multiple regulatory layers. This approach integrates splicing factor regulation into a single score derived solely from exon inclusion signatures, allowing functional interpretation of heterogeneous conditions. As a proof of concept, we identify recurrent cancer splicing programs, revealing oncogenic- and tumor suppressor-like splicing factors missed by conventional methods. These programs correlate with patient survival and key cancer hallmarks: initiation, proliferation, and immune evasion. Altogether, we show splicing factor activity can be accurately estimated from exon inclusion changes, enabling comprehensive analyses of splicing regulation with minimal data requirements.

VIPER↗

Lineage-specific splicing regulation of MAPT gene in the primate brain.

Divergence of precursor messenger RNA (pre-mRNA) alternative splicing (AS) is widespread in mammals, including primates, but the underlying mechanisms and functional impact are poorly understood. Here, we modeled cassette exon inclusion in primate brains as a quantitative trait and identified 1,170 (∼3%) exons with lineage-specific splicing shifts under stabilizing selection. Among them, microtubule-associated protein tau (MAPT) exons 2 and 10 underwent anticorrelated, two-step evolutionary shifts in the catarrhine and hominoid lineages, leading to their present inclusion levels in humans. The developmental-stage-specific divergence of exon 10 splicing, whose dysregulation can cause frontotemporal lobar degeneration (FTLD), is mediated by divergent distal intronic MBNL-binding sites. Competitive binding of these sites by CRISPR-dCas13d/gRNAs effectively reduces exon 10 inclusion, potentially providing a therapeutically compatible approach to modulate tau isoform expression. Our data suggest adaptation of MAPT function and, more generally, a role for AS in the evolutionary expansion of the primate brain.

tau Proteins↗

An increased specificity score matrix for the prediction of SF2/ASF-specific exonic splicing enhancers.

Numerous disease-associated point mutations exert their effects by disrupting the activity of exonic splicing enhancers (ESEs). We previously derived position weight matrices to predict putative ESEs specific for four human SR proteins. The score matrices are part of ESEfinder, an online resource to identify ESEs in query sequences. We have now carried out a refined functional SELEX screen for motifs that can act as ESEs in response to the human SR protein SF2/ASF. The test BRCA1 exon under selection was internal, rather than the 3'-terminal IGHM exon used in our earlier studies. A naturally occurring heptameric ESE in BRCA1 exon 18 was replaced with two libraries of random sequences, one seven nucleotides in length, the other 14. Following three rounds of selection for in vitro splicing via internal exon inclusion, new consensus motifs and score matrices were derived. Many winner sequences were demonstrated to be functional ESEs in S100-extract-complementation assays with recombinant SF2/ASF. Motif-score threshold values were derived from both experimental and statistical analyses. Motif scores were shown to correlate with levels of exon inclusion, both in vitro and in vivo. Our results confirm and extend our earlier data, as many of the same motifs are recognized as ESEs by both the original and our new score matrix, despite the different context used for selection. Finally, we have derived an increased specificity score matrix that incorporates information from both of our SF2/ASF-specific matrices and that accurately predicts the exon-skipping phenotypes of deleterious point mutations.

Alternative Splicing↗

A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.

Transcription factor binding sites (TFBSs) are short DNA sequences interacting with transcription factors (TFs), which regulate gene expression. Due to the relatively short length of such binding sites, it is largely unclear how the specificity of protein-DNA interaction is achieved. Here, we have performed a genome-wide analysis of TFBS-like sequences for the transcriptional repressor, RE1 Silencing Transcription Factor (REST), as well as for several other representative mammalian TFs (c-myc, p53, HNF-1 and CREB). We find a nonrandom distribution of inexact sites for these TFs, referred to as highly-degenerate TFBSs, that are enriched around the cognate binding sites. Comparisons among human, mouse and rat orthologous promoters reveal that these highly-degenerate sites are conserved significantly more than expected by random chance, suggesting their positive selection during evolution. We propose that this arrangement provides a favorable genomic landscape for functional target site selection.

Animals↗

Profiling alternatively spliced mRNA isoforms for prostate cancer classification.

BACKGROUND: Prostate cancer is one of the leading causes of cancer illness and death among men in the United States and world wide. There is an urgent need to discover good biomarkers for early clinical diagnosis and treatment. Previously, we developed an exon-junction microarray-based assay and profiled 1532 mRNA splice isoforms from 364 potential prostate cancer related genes in 38 prostate tissues. Here, we investigate the advantage of using splice isoforms, which couple transcriptional and splicing regulation, for cancer classification. RESULTS: As many as 464 splice isoforms from more than 200 genes are differentially regulated in tumors at a false discovery rate (FDR) of 0.05. Remarkably, about 30% of genes have isoforms that are called significant but do not exhibit differential expression at the overall mRNA level. A support vector machine (SVM) classifier trained on 128 signature isoforms can correctly predict 92% of the cases, which outperforms the classifier using overall mRNA abundance by about 5%. It is also observed that the classification performance can be improved using multivariate variable selection methods, which take correlation among variables into account. CONCLUSION: These results demonstrate that profiling of splice isoforms is able to provide unique and important information which cannot be detected by conventional microarrays.

Algorithms↗

Significance of gene ranking for classification of microarray samples.

Many methods for classification and gene selection with microarray data have been developed. These methods usually give a ranking of genes. Evaluating the statistical significance of the gene ranking is important for understanding the results and for further biological investigations, but this question has not been well addressed for machine learning methods in existing works. Here, we address this problem by formulating it in the framework of hypothesis testing and propose a solution based on resampling. The proposed r-test methods convert gene ranking results into position p-values to evaluate the significance of genes. The methods are tested on three real microarray data sets and three simulation data sets with support vector machines as the method of classification and gene selection. The obtained position p-values help to determine the number of genes to be selected and enable scientists to analyze selection results by sophisticated multivariate methods under the same statistical inference paradigm as for simple hypothesis testing methods.

Algorithms↗