PubMed HealthSearch

PubMed · 40279271

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

Abstract

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alison Kyrouz, Lian Liu, Lixin Qin, Brian Tjaden. 2025-05-06. Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.. https://doi.org/10.1093/bioinformatics%2Fbtaf250

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Hidden proteins encoded by non-canonical open reading frames: A review.

There is increasing evidence that translation is not limited to annotated protein-coding genes. Ribosome profiling sequencing, mass spectrometry-based proteomics, and immunopeptidomics have identified the productive translation of non-canonical open reading frames (ORFs). This suggests that the functional proteome includes not only conserved proteins but also proteins hidden in non-coding RNAs and de novo proteins. Some of these translated products are functional peptides, while others may be non-functional, potentially arising from evolutionary events. Several non-canonical ORF-encoded peptides have been found to regulate multiple physiological and pathological functions, particularly in cancer, immunity, and inflammation, indicating that they have potential as biomarkers and novel therapeutic targets. To better understand the diversity of functional peptides and translated non-canonical ORFs based on existing data, we summarize their classification according to transcriptional features and supporting evidence, including non-canonical ORFs located in ncRNAs and canonical mRNAs. This review provides a concise summary of the origin, discovery methods, and classification of non-canonical ORFs. It offers insights into the origins and functions of non-canonical ORF-encoded peptides from an evolutionary perspective, while also exploring the biological functions and regulatory mechanisms of these non-canonical ORF-encoded hidden proteins in tumorigenesis and progression.

Open Reading Frames

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

A positive-sense single-stranded RNA virus acquired a negative-sense open reading frame through recombination.

Although positive- and negative-sense single-stranded RNA viruses are ubiquitous in nature, there is currently no evidence of recombination or reassortment between viruses with these two major forms of genome organization. Here, we describe the discovery of brine shrimp virga-like virus 1 (BSVV1), a novel positive-sense single-stranded RNA virus with a recombinant genome structure derived from two viral phyla with differing genome organizations. The genome of BSVV1 comprises three open reading frames (ORFs). ORF1 resembles the RNA-dependent RNA polymerase of Ips virga-like virus 1 (a positive-sense RNA virus), while ORF2, transcribed in the positive orientation, is related to the glycoprotein of Hubei bunya-like virus 10 and other negative-sense RNA viruses. The predicted ORF3 was unique to BSVV1 without known homologs identified. The presence of the three protein products was verified by mass spectrometry. Notably, our analysis also revealed that BSVV1 is geographically widespread and found in brine shrimp from at least eight countries on four continents. In addition, BSVV1 was successfully cultured and proliferated to high viral loads during brine shrimp development. In sum, we provide compelling evidence of an ancient recombination event between negative- and positive-sense single-stranded RNA viruses, enriching our understanding of the evolution of genome structures in RNA viruses.

Open Reading Frames