PubMed HealthSearch

PubMed · 39419446

Mapping Start Codons of Small Open Reading Frames by N-Terminomics Approach.

Abstract

sORF-encoded peptides (SEPs) refer to proteins encoded by small open reading frames (sORFs) with a length of less than 100 amino acids, which play an important role in various life activities. Analysis of known SEPs showed that using non-canonical initiation codons of SEPs was more common. However, the current analysis of SEP sequences mainly relies on bioinformatics prediction, and most of them use AUG as the start site, which may not be completely correct for SEPs. Chemical labeling was used to systematically analyze the N-terminal sequences of SEPs to accurately define the start sites of SEPs. By comparison, we found that dimethylation and guanidinylation are more efficient than acetylation. The ACN precipitation and heating precipitation performed better in SEP enrichment. As an N-terminal peptide enrichment material, Hexadhexaldehyde was superior to CNBr-activated agarose and NHS-activated agarose. Combining these methods, we identified 128 SEPs with 131 N-terminal sequences. Among them, two-thirds are novel N-terminal sequences, and most of them start from the 11-31st amino acids of the original sequence. Partial novel N-termini were produced by proteolysis or signal peptide removal. Some SEPs' transcription start sites were corrected to be non-AUG start codons. One novel start codon was validated using GFP-tag vectors. These results demonstrated that the chemical labeling approaches would be beneficial for identifying the start codons of sORFs and the real N-terminal of their encoded peptides, which helps better understand the characterization of SEPs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mingbo Peng, Tianjing Wang, Yujie Li, Zheng Zhang, Cuihong Wan. 2024-10-16. Mapping Start Codons of Small Open Reading Frames by N-Terminomics Approach.. https://doi.org/10.1016/j.mcpro.2024.100860

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Hidden proteins encoded by non-canonical open reading frames: A review.

There is increasing evidence that translation is not limited to annotated protein-coding genes. Ribosome profiling sequencing, mass spectrometry-based proteomics, and immunopeptidomics have identified the productive translation of non-canonical open reading frames (ORFs). This suggests that the functional proteome includes not only conserved proteins but also proteins hidden in non-coding RNAs and de novo proteins. Some of these translated products are functional peptides, while others may be non-functional, potentially arising from evolutionary events. Several non-canonical ORF-encoded peptides have been found to regulate multiple physiological and pathological functions, particularly in cancer, immunity, and inflammation, indicating that they have potential as biomarkers and novel therapeutic targets. To better understand the diversity of functional peptides and translated non-canonical ORFs based on existing data, we summarize their classification according to transcriptional features and supporting evidence, including non-canonical ORFs located in ncRNAs and canonical mRNAs. This review provides a concise summary of the origin, discovery methods, and classification of non-canonical ORFs. It offers insights into the origins and functions of non-canonical ORF-encoded peptides from an evolutionary perspective, while also exploring the biological functions and regulatory mechanisms of these non-canonical ORF-encoded hidden proteins in tumorigenesis and progression.

Open Reading Frames

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames