PubMed Health⌕ Search

PubMed · 8786136

Evaluation of gene structure prediction programs.

Abstract

We evaluate a number of computer programs designed to predict the structure of protein coding genes in genomic DNA sequences. Computational gene identification is set to play an increasingly important role in the development of the genome projects, as emphasis turns from mapping to large-scale sequencing. The evaluation presented here serves both to assess the current status of the problem and to identify the most promising approaches to ensure further progress. The programs analyzed were uniformly tested on a large set of vertebrate sequences with simple gene structure, and several measures of predictive accuracy were computed at the nucleotide, exon, and protein product levels. The results indicated that the predictive accuracy of the programs analyzed was lower than originally found. The accuracy was even lower when considering only those sequences that had recently been entered and that did not show any similarity to previously entered sequences. This indicates that the programs are overly dependent on the particularities of the examples they learn from. For most of the programs, accuracy in this test set ranged from 0.60 to 0.70 as measured by the Correlation Coefficient (where 1.0 corresponds to a perfect prediction and 0.0 is the value expected for a random prediction), and the average percentage of exons exactly identified was less than 50%. Only those programs including protein sequence database searches showed substantially greater accuracy. The accuracy of the programs was severely affected by relatively high rates of sequence errors. Since the set on which the programs were tested included only relatively short sequences with simple gene structure, the accuracy of the programs is likely to be even lower when used for large uncharacterized genomic sequences with complex structure. While in such cases, programs currently available may still be of great use in pinpointing the regions likely to contain exons, they are far from being powerful enough to elucidate its genomic structure completely.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M Burset, R Guigó. 1996-06-15. Evaluation of gene structure prediction programs.. https://doi.org/10.1006/geno.1996.0298

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Translational regulation of Sf1 integrates alternative splicing and hematopoietic stem cell fate.

The transition of hematopoietic stem cells (HSCs) from quiescence to lineage commitment requires precise posttranscriptional control, yet the contribution of mRNA isoform regulation remains poorly defined. Here, we identify a translationally controlled splicing program that contributes to HSC fate decisions. Using activity-based signatures of 305 splicing regulators, we uncover widespread posttranscriptional modulation of the spliceosome in stem and progenitor cells. The branch point recognition factor Sf1 emerges as a key node, regulated by a conserved structured 5' untranslated region (UTR) that cooperates with the RNA-binding protein Igf2bp2 to control its translation. Disrupting this cis-trans module reduces Sf1 protein synthesis and skews differentiation toward stem and erythroid programs. Mechanistically, Sf1-dependent alternative splicing remodels 5' UTRs of hematopoietic and DNA damage response genes, altering their translation and modulating DNA damage resolution. Together, these findings reveal an unrecognized translational layer controlling spliceosome activity and link RNA regulons, alternative splicing, and HSC fate determination.

Alternative Splicing↗

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗