PubMed Health⌕ Search

Biomedical subjects

Rotem Sorek

Publications and source records attributed to Rotem Sorek.

At least 19 recordsLinked to original sources

Assessing the number of ancestral alternatively spliced exons in the human genome.

BACKGROUND: It is estimated that between 35% and 74% of all human genes undergo alternative splicing. However, as a gene that undergoes alternative splicing can have between one and dozens of alternative exons, the number of alternatively spliced genes by itself is not informative enough. An additional parameter, which was not addressed so far, is therefore the number of human exons that undergo alternative splicing. We have previously described an accurate machine-learning method allowing the detection of conserved alternatively spliced exons without using ESTs, which relies on specific features of the exon and its genomic vicinity that distinguish alternatively spliced exons from constitutive ones. RESULTS: In this study we use the above-described approach to calculate that 7.2% (+/- 1.1%) of all human exons that are conserved in mouse are alternatively spliced in both species. CONCLUSION: This number is the first estimation for the extent of ancestral alternatively spliced exons in the human genome.

Algorithms↗

Genomic fossils as a snapshot of the human transcriptome.

Processed pseudogenes (PPGs) are cDNA sequences that were generated through reverse transcription of mature, spliced mRNAs and have subsequently been reinserted at a new genomic location. These cDNA sequences are usually no longer transcribed and are considered "dead on arrival." Here we show that PPGs can be used to generate a map of the transcriptome. By analyzing thousands of human PPGs, we were able to discover hundreds of transcript variants so far unidentified. An experimental verification of a subset of these variants by RT-PCR indicates that most of them are still active in the human transcriptome. Furthermore, we demonstrate that PPGs can enable the identification of ancient splice variants that were expressed ancestrally but are now extinct. Our results show that the genome itself carries a "virtual cDNA library" that can readily be used to analyze both present and ancestral transcripts. Our approach can be applied to sequenced metazoan genomes to computationally annotate splicing variation even when expressed sequences are unavailable.

Alternative Splicing↗

Transcription-mediated gene fusion in the human genome.

Transcription of a gene usually ends at a regulated termination point, preventing the RNA-polymerase from reading through the next gene. However, sporadic reports suggest that chimeric transcripts, formed by transcription of two consecutive genes into one RNA, can occur in human. The splicing and translation of such RNAs can lead to a new, fused protein, having domains from both original proteins. Here, we systematically identified over 200 cases of intergenic splicing in the human genome (involving 421 genes), and experimentally demonstrated that at least half of these fusions exist in human tissues. We showed that unique splicing patterns dominate the functional and regulatory nature of the resulting transcripts, and found intergenic distance bias in fused compared with nonfused genes. We demonstrate that the hundreds of fused genes we identified are only a subset of the actual number of fused genes in human. We describe a novel evolutionary mechanism where transcription-induced chimerism followed by retroposition results in a new, active fused gene. Finally, we provide evidence that transcription-induced chimerism can be a mechanism contributing to the evolution of protein complexes.

Evolution, Molecular↗

Naturally occurring antisense: transcriptional leakage or real overlap?

Naturally occurring antisense transcription is associated with the regulation of gene expression through a variety of biological mechanisms. Several recent genome-wide studies reported the identification of potential antisense transcripts for thousands of mammalian genes, many of them resulting from alternatively polyadenylated transcripts or heterogeneous transcription start sites. However, it is not clear whether this transcriptional plasticity is intentional, leading to regulated overlap between the transcripts, or, alternatively, represents a "leakage" of the RNA transcription machinery. To address this question through an evolutionary approach, we compared the genomic organization of genes, with or without antisense, between human, mouse, and the pufferfish Fugu rubripes. Our hypothesis was that if two neighboring genes overlap and have a sense-antisense relationship, we would expect negative selection acting on the evolutionary separation between them. We found that antisense gene pairs are twice as likely to preserve their genomic organization throughout vertebrates' evolution compared to nonantisense pairs, implying an overlap existence in the ancestral genome. In addition, we show that increasing the genomic distance between pairs of genes having a sense-antisense relationship is selected against. These findings indicate that, at least in part, the abundance of antisense transcripts observed in expressed data represents real overlap rather than transcriptional leakage. Moreover, our results imply that natural antisense transcription has considerably affected vertebrate genome evolution.

Animals↗

Is abundant A-to-I RNA editing primate-specific?

A-to-I RNA editing is common in all eukaryotes, and is associated with various neurological functions. Recently, A-to-I editing was found to occur frequently in the human transcriptome. In this article, we show that the frequency of A-to-I editing in humans is at least an order of magnitude higher than in the mouse, rat, chicken or fly genomes. The extraordinary frequency of RNA editing in human is explained by the dominance of the primate-specific Alu element in the human transcriptome, which increases the number of double-stranded RNA substrates.

3' Untranslated Regions↗

Is there any sense in antisense editing?

Several recent studies have hypothesized that sense-antisense RNA-transcript pairs create dsRNA duplexes that undergo extensive A-to-I RNA editing. We studied human and mouse genomic antisense regions and found that the editing level in these areas is negligible. This observation questions the scope of sense-antisense duplexes formation in-vivo, which is the basis for several proposed regulatory mechanisms.

Alu Elements↗

Accurate identification of alternatively spliced exons using support vector machine.

MOTIVATION: Alternative splicing is a major component of the regulatory action on mammalian transcriptomes. It is estimated that over half of all human genes have more than one splice variant. Previous studies have shown that alternatively spliced exons possess several features that distinguish them from constitutively spliced ones. Recently, we have demonstrated that such features can be used to distinguish alternative from constitutive exons. In the current study, we used advanced machine learning methods to generate robust classifier of alternative exons. RESULTS: We extracted several hundred local sequence features of constitutive as well as alternative exons. Using feature selection methods we find seven attributes that are dominant for the task of classification. Several less informative features help to slightly increase the performance of the classifier. The classifier achieves a true positive rate of 50% for a false positive rate of 0.5%. This result enables one to reliably identify alternatively spliced exons in exon databases that are believed to be dominated by constitutive exons.

Algorithms↗

Minimal conditions for exonization of intronic sequences: 5' splice site formation in alu exons.

Alu exonization, which is an evolutionary pathway that creates primate-specific transcriptomic diversity, is a powerful tool for studying alternative-splicing regulation. Through bioinformatic analyses combined with experimental methodology, we identified the mutational changes needed to create functional 5' splice sites in Alu. We revealed a complex mechanism by which the sequence composition of the 5' splice site and its base pairing with the small nuclear RNA U1 govern alternative splicing. We show that in Alu-derived GC introns the strength of the base pairing between U1 snRNA and the 5' splice site controls the skipping/inclusion ratio of alternative splicing. Based on these findings, we identified 7810 Alus within the human genome that are prone to exonization. Mutations in these Alus may cause genetic disorders or contribute to human-specific protein diversity.

5' Untranslated Regions↗

AluGene: a database of Alu elements incorporated within protein-coding genes.

Alu elements are short interspersed elements (SINEs) approximately 300 nucleotides in length. More than 1 million Alus are found in the human genome. Despite their being genetically functionless, recent findings suggest that Alu elements may have a broad evolutionary impact by affecting gene structures, protein sequences, splicing motifs and expression patterns. Because of these effects, compiling a genomic database of Alu sequences that reside within protein-coding genes seemed a useful enterprise. Presently, such data are limited since the structural and positional information on genes and Alu sequences are scattered throughout incompatible and unconnected databases. AluGene (http://Alugene.tau.ac.il/) provides easy access to a complete Alu map of the human genome, as well as Alu-associated information. The Alu elements are annotated with respect to coding region and exon/intron location. This design facilitates queries on Alu sequences, locations, as well as motifs and compositional properties via a one-stop search page.

Alu Elements↗

In search of antisense.

In recent years, natural antisense transcripts (NATs) have been implicated in many aspects of eukaryotic gene expression including genomic imprinting, RNA interference, translational regulation, alternative splicing, X-inactivation and RNA editing. Moreover, there is growing evidence to suggest that antisense transcription might have a key role in a range of human diseases. Consequently, there have been several recent attempts to identify novel NATs. To date, approximately 2500 mammalian NATs have been found, indicating that antisense transcription might be a common mechanism of regulating gene expression in human cells. There are increasingly diverse ways in which antisense transcription can regulate gene expression and evidence for the involvement of NATs in human disease is emerging. A range of bioinformatic resources could be used to assist future antisense research.

Alternative Splicing↗

How prevalent is functional alternative splicing in the human genome?

Comparative analyses of ESTs and cDNAs with genomic DNA predict a high frequency of alternative splicing in human genes. However, there is an ongoing debate as to how many of these predicted splice variants are functional and how many are the result of aberrant splicing (or 'noise'). To address this question, we compared alternatively spliced cassette exons that are conserved between human and mouse with EST-predicted cassette exons that are not conserved in the mouse genome. Presumably, conserved exon-skipping events represent functional alternative splicing. We show that conserved (functional) cassette exons possess unique characteristics in size, repeat content and in their influence on the protein. By contrast, most non-conserved cassette exons do not share these characteristics. We conclude that a significant portion of cassette exons evident in EST databases is not functional, and might result from aberrant rather than regulated splicing.

Alternative Splicing↗

A non-EST-based method for exon-skipping prediction.

It is estimated that between 35% and 74% of all human genes can undergo alternative splicing. Currently, the most efficient methods for large-scale detection of alternative splicing use expressed sequence tags (ESTs) or microarray analysis. As these methods merely sample the transcriptome, splice variants that do not appear in deeply sampled tissues have a low probability of being detected. We present a new method by which we can predict that an internal exon is skipped (namely whether it is a cassette-exon) merely based on its naked genomic sequence and on the sequence of its mouse ortholog. No other data, such as ESTs, are required for the prediction. Using our method, which was experimentally validated, we detected hundreds of novel splice variants that were not detectable using ESTs. We show that a substantial fraction of the splice variants in the human genome could not be identified through current human EST or cDNA data.

Alternative Splicing↗

The birth of an alternatively spliced exon: 3' splice-site selection in Alu exons.

Alu repetitive elements can be inserted into mature messenger RNAs via a splicing-mediated process termed exonization. To understand the molecular basis and the regulation of the process of turning intronic Alus into new exons, we compiled and analyzed a data set of human exonized Alus. We revealed a mechanism that governs 3' splice-site selection in these exons during alternative splicing. On the basis of these findings, we identified mutations that activated the exonization of a silent intronic Alu.

Adenosine Deaminase↗

Widespread occurrence of antisense transcription in the human genome.

An increasing number of eukaryotic genes are being found to have naturally occurring antisense transcripts. Here we study the extent of antisense transcription in the human genome by analyzing the public databases of expressed sequences using a set of computational tools designed to identify sense-antisense transcriptional units on opposite DNA strands of the same genomic locus. The resulting data set of 2,667 sense-antisense pairs was evaluated by microarrays containing strand-specific oligonucleotide probes derived from the region of overlap. Verification of specific cases by northern blot analysis with strand-specific riboprobes proved transcription from both DNA strands. We conclude that > or =60% of this data set, or approximately 1,600 predicted sense-antisense transcriptional units, are transcribed from both DNA strands. This indicates that the occurrence of antisense transcription, usually regarded as infrequent, is a very common phenomenon in the human genome. Therefore, antisense modulation of gene expression in human cells may be a common regulatory mechanism.

Algorithms↗

A novel algorithm for computational identification of contaminated EST libraries.

A key goal of the Human Genome Project was to understand the complete set of human proteins, the proteome. Since the genome sequence by itself is not sufficient for predicting new genes and alternative splicing events that lead to new proteins, expressed sequence tags (ESTs) are used as the primary tool for these purposes. The high prevalence of artifacts in dbEST, however, often leads to invalid predictions. Here we describe a novel method for recognizing genomic DNA contamination and other artifacts that cannot be identified using current EST cleaning techniques. Our method uses the alignment of the entire set of ESTs to the human genome to identify highly contaminated EST libraries. We discovered 53 highly contaminated libraries and a subset of 24 766 ESTs from these libraries that probably represent contamination with genomic DNA, pre-mRNA, and ESTs that span non-canonical introns. Although this is only a small fraction of the entire EST dataset, each contaminating sequence could create a spurious transcript prediction. Indeed, in the clustering and assembly tool that we used, these sequences would have caused incorrect inference of 9575 new splice variants and 6370 new genes. Conclusions based on EST analysis, including prediction of alternative splicing, should be re-evaluated in light of these results. Our method, along with the identified set of contaminated sequences, will be essential for applications that depend on large EST datasets.

Algorithms↗

Evolutionary dynamics of large numts in the human genome: rarity of independent insertions and abundance of post-insertion duplications.

We determined the phylogenetic positions of 82 large nuclear pseudogenes of mitochondrial origin (numts) within the human genome. For each numt, two possibilities pertaining to its origin were considered: (1) independent insertion from the mitochondria into the nucleus, or (2) genomic duplication subsequent to the insertion. A significant increase in the rate of numt accumulation is seen after the divergence of Platyrrhini (New World monkeys) from the Catarrhini (Old World monkeys, apes and humans). By using pairwise phylogenetic analyses, we were able to demonstrate that this peak in numt accumulation is mostly the result of duplication of preexisting nuclear numts rather than the result of an increase in mitochondrial-sequence insertion. In fact, only about a third of all the numt repertoire in the human nuclear genome is due to insertions of mitochondrial sequences, the rest originated as duplications of preexisting numts. Hence, we conclude that numt insertion occurs at a much lower rate than previously reported. As expected under the assumption that genomic duplications occur at rates that are uninfluenced by content, older numts were found to be duplicated more times than recently inserted ones.

DNA Transposable Elements↗

Intronic sequences flanking alternatively spliced exons are conserved between human and mouse.

Comparison of the sequences of mouse and human genomes revealed a surprising number of nonexonic, nonexpressed conserved sequences, for which no function could be assigned. To study the possible correlation between these conserved intronic sequences and alternative splicing regulation, we developed a method to identify exons that are alternatively spliced in both human and mouse. We compiled two exon sets: one of alternatively spliced conserved exons and another of constitutively spliced conserved exons. We found that 77% of the conserved alternatively spliced exons were flanked on both sides by long conserved intronic sequences. In comparison, only 17% of the conserved constitutively spliced exons were flanked by such conserved intronic sequences. The average length of the conserved intronic sequences was 103 bases in the upstream intron and 94 bases in the downstream intron. The average identity levels in the immediately flanking intronic sequences were 88% and 80% for the upstream and downstream introns, respectively, higher than the conservation levels of 77% that were measured in promoter regions. Our results suggest that the function of many of the intronic sequence blocks that are conserved between human and mouse is the regulation of alternative splicing.

3' Flanking Region↗