PubMed HealthSearch

SEARCH · PubMed Health

Results for “Open Reading Frames”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Auditing bacterial dark-gene screens for superimposed open reading frame artefacts: A multi-layer analysis of Rv2438A in Mycobacterium tuberculosis.

Essentiality and knockdown-vulnerability screens can promote spurious bacterial open reading frames when those frames overlap essential genes, because such a frame inherits its neighbour's signals undiluted and therefore satisfies the screen's criteria better than a genuine small gene. We present a multi-layer audit that tests this failure mode across genome annotation, transposon mutagenesis, CRISPR interference, homology, transcript mapping, proteomics, and population variation. We apply it to Rv2438A, a 92-codon conserved hypothetical open reading frame of Mycobacterium tuberculosis ranked first by our own dark-gene target screen. Rv2438A is superimposed on the essential NAD synthetase locus nadE: 44% lies within its coding sequence on the opposite strand, and the remainder covers its promoter and transcription start site. Consequently, three of five Himar1 sites lie within nadE, no CRISPRi guide can target Rv2438A without binding nadE, and the cross-species hit maps to the same nadE start junction. Rv2438A lacks its own transcription start site and is absent from every proteomic dataset that detects nadE. A genome-wide scan identifies six short, overlapping, uncharacterised loci among 3907 annotated genes, but only Rv2438A combines overlap and essentiality with non-detection across all proteomic datasets; rare genome-wide, it ranked first among screen hits. We provide an implementable audit workflow and a codon-position control, but measure the control's sensitivity as only two of five genes with attested protein, limiting it to confirmatory use. Overlap coordinates and neighbour-specific experimental resolution should therefore be reported before bacterial dark genes are prioritised.

CRISPR interference

Distinct types of short open reading frames are translated in plant cells.

Genomes contain millions of short (<100 codons) open reading frames (sORFs), which are usually dismissed during gene annotation. Nevertheless, peptides encoded by such sORFs can play important biological roles, and their impact on cellular processes has long been underestimated. Here, we analyzed approximately 70,000 transcribed sORFs in the model plant Physcomitrella patens (moss). Several distinct classes of sORFs that differ in terms of their position on transcripts and the level of evolutionary conservation are present in the moss genome. Over 5000 sORFs were conserved in at least one of 10 plant species examined. Mass spectrometry analysis of proteomic and peptidomic data sets suggested that tens of sORFs located on distinct parts of mRNAs and long noncoding RNAs (lncRNAs) are translated, including conserved sORFs. Translational analysis of the sORFs and main ORFs at a single locus suggested the existence of genes that code for multiple proteins and peptides with tissue-specific expression. Functional analysis of four lncRNA-encoded peptides showed that sORFs-encoded peptides are involved in regulation of growth and differentiation in moss. Knocking out lncRNA-encoded peptides resulted in a decrease of moss growth. In contrast, the overexpression of these peptides resulted in a diverse range of phenotypic effects. Our results thus open new avenues for discovering novel, biologically active peptides in the plant kingdom.

Bryopsida

Mapping Start Codons of Small Open Reading Frames by N-Terminomics Approach.

sORF-encoded peptides (SEPs) refer to proteins encoded by small open reading frames (sORFs) with a length of less than 100 amino acids, which play an important role in various life activities. Analysis of known SEPs showed that using non-canonical initiation codons of SEPs was more common. However, the current analysis of SEP sequences mainly relies on bioinformatics prediction, and most of them use AUG as the start site, which may not be completely correct for SEPs. Chemical labeling was used to systematically analyze the N-terminal sequences of SEPs to accurately define the start sites of SEPs. By comparison, we found that dimethylation and guanidinylation are more efficient than acetylation. The ACN precipitation and heating precipitation performed better in SEP enrichment. As an N-terminal peptide enrichment material, Hexadhexaldehyde was superior to CNBr-activated agarose and NHS-activated agarose. Combining these methods, we identified 128 SEPs with 131 N-terminal sequences. Among them, two-thirds are novel N-terminal sequences, and most of them start from the 11-31st amino acids of the original sequence. Partial novel N-termini were produced by proteolysis or signal peptide removal. Some SEPs' transcription start sites were corrected to be non-AUG start codons. One novel start codon was validated using GFP-tag vectors. These results demonstrated that the chemical labeling approaches would be beneficial for identifying the start codons of sORFs and the real N-terminal of their encoded peptides, which helps better understand the characterization of SEPs.

Open Reading Frames

The Herpesvirus Saimiri open reading frame 73 gene product interacts with the cellular protein p32.

The role of the gamma-2 herpesvirus open reading frame (ORF) 73 gene product has become the focus of considerable interest. It has recently been shown that the Kaposi's sarcoma-associated herpesvirus (KSHV) latency-associated nuclear antigen (LANA) is expressed during a latent infection and can modulate both viral and cellular gene expression. The herpesvirus saimiri (HVS) ORF 73 gene product has some sequence homology to LANA; however, the role of HVS ORF 73 is unknown. We have previously demonstrated that HVS ORF73 is expressed in a stably transduced human carcinoma cell line, where HVS genomes persist as nonintegrated circular episomes. This implies that there may be some functional homology between these proteins. To further investigate the role of the HVS ORF 73 protein, the yeast two-hybrid system was employed to identify interacting cellular proteins. We demonstrate that ORF 73 interacts with the cellular protein p32 and triggers the accumulation of p32 in the nucleus. Using reporter gene-based transient-transfection assays, we demonstrate that ORF 73 can transactivate a number of heterologous promoter constructs and also upregulate its own promoter. Moreover, ORF 73 and p32 act synergistically to transactivate these promoters. The binding of ORF 73 to p32 is mediated by an amino-terminal arginine-rich domain, which contains two functionally distinct nuclear localization signals. The p32 binding domains are required for ORF 73 transactivating abilities and for ORF 73 to induce nuclear accumulation of p32. These results suggest that ORF 73 can function as a regulator of gene expression and that p32 is involved in ORF 73-dependent transcriptional activation.

Amino Acid Sequence

Hidden proteins encoded by non-canonical open reading frames: A review.

There is increasing evidence that translation is not limited to annotated protein-coding genes. Ribosome profiling sequencing, mass spectrometry-based proteomics, and immunopeptidomics have identified the productive translation of non-canonical open reading frames (ORFs). This suggests that the functional proteome includes not only conserved proteins but also proteins hidden in non-coding RNAs and de novo proteins. Some of these translated products are functional peptides, while others may be non-functional, potentially arising from evolutionary events. Several non-canonical ORF-encoded peptides have been found to regulate multiple physiological and pathological functions, particularly in cancer, immunity, and inflammation, indicating that they have potential as biomarkers and novel therapeutic targets. To better understand the diversity of functional peptides and translated non-canonical ORFs based on existing data, we summarize their classification according to transcriptional features and supporting evidence, including non-canonical ORFs located in ncRNAs and canonical mRNAs. This review provides a concise summary of the origin, discovery methods, and classification of non-canonical ORFs. It offers insights into the origins and functions of non-canonical ORF-encoded peptides from an evolutionary perspective, while also exploring the biological functions and regulatory mechanisms of these non-canonical ORF-encoded hidden proteins in tumorigenesis and progression.

Open Reading Frames

PKD1 upstream open reading frames affect Polycystin-1 expression and polycystic kidney disease phenotypes.

Autosomal dominant polycystic kidney disease (ADPKD) accounts for 5%-10% of prevalent end-stage kidney failure (ESKD). ADPKD cysts result from a loss of sufficient functional expression of PKD1/Polycystin-1 (PC1) in approximately 80% of families. Kidney disease severity correlates with the extent to which PC1 dosage is reduced below a critical level, and evidence suggests therapeutic benefit from increasing PC1 expression in these conditions. Upstream open reading frame (uORF) translation can reduce translation of a protein's coding sequence. Ribosome profiling data and bioinformatic predictions suggested the presence of conserved PKD1 uORFs, so we sought to explore their biological role. We generated luciferase reporters and two humanized PKD1 5' UTR mouse models with or without single nucleotide edits removing uORF start codons (&#x394;uORF) to define active uORFs and test their impact on PC1 translation. PKD1 uORF start codons can robustly initiate translation, and &#x394;uORF conveys a 2-4 fold increase in PC1 protein expression and resultant prevention of kidney cysts in Dnajb11 as well as in Pkd1 missense models. PKD1 uORF1-blocking steric antisense oligonucleotides (ASOs) substantially increase PC1 expression in vitro. PKD1 uORFs play an important role in the low basal expression of WT PKD1, and their inhibition represents an opportunity to therapeutically increase PC1 translation in polycystic kidney and liver disease resulting from reduced dosage of PC1.

Animals

Expression of De Novo Open Reading Frames in Natural Populations of Drosophila melanogaster.

De novo genes, which originate from noncoding DNA, are known to have a high rate of turnover over short evolutionary timescales, such as within a species. Thus, their expression is often lineage- or genetic background-specific. However, little is known about their levels and breadth of expression as populations of a species diverge. In this study, we utilized publicly available RNA-seq data to examine the expression of newly evolved open reading frames (neORFs) in comparison to non- and protein-coding genes in Drosophila melanogaster populations from the derived species range in Europe and the ancestral range in sub-Saharan Africa. Our datasets included two adult tissue types as well as whole bodies at two temperatures for both sexes and three larval/prepupal developmental stages in a single tissue and sex, which allowed us to examine neORF expression and divergence across multiple sample types as well as sex and population. We detected a relatively large proportion (approximately 50%) of annotated neORFs as expressed in the population samples, with neORFs often showing greater expression divergence between populations than non- or protein-coding genes. However, differential expression of neORFs between populations tended to occur in a sample type-specific manner. On the other hand, neORFs displayed less sex-biased expression than the other two gene classes, with the majority of sex-biased neORFs detected in whole bodies, which may be attributable to the presence of the gonads. We also found that neORFs shared among multiple lines in the original set of inbred lines in which they were first detected were more likely to be both expressed and differentially expressed in the new population samples, suggesting that neORFs at a higher frequency (i.e. present in more individuals) within a species are more likely to be functional.

Animals

A positive-sense single-stranded RNA virus acquired a negative-sense open reading frame through recombination.

Although positive- and negative-sense single-stranded RNA viruses are ubiquitous in nature, there is currently no evidence of recombination or reassortment between viruses with these two major forms of genome organization. Here, we describe the discovery of brine shrimp virga-like virus 1 (BSVV1), a novel positive-sense single-stranded RNA virus with a recombinant genome structure derived from two viral phyla with differing genome organizations. The genome of BSVV1 comprises three open reading frames (ORFs). ORF1 resembles the RNA-dependent RNA polymerase of Ips virga-like virus 1 (a positive-sense RNA virus), while ORF2, transcribed in the positive orientation, is related to the glycoprotein of Hubei bunya-like virus 10 and other negative-sense RNA viruses. The predicted ORF3 was unique to BSVV1 without known homologs identified. The presence of the three protein products was verified by mass spectrometry. Notably, our analysis also revealed that BSVV1 is geographically widespread and found in brine shrimp from at least eight countries on four continents. In addition, BSVV1 was successfully cultured and proliferated to high viral loads during brine shrimp development. In sum, we provide compelling evidence of an ancient recombination event between negative- and positive-sense single-stranded RNA viruses, enriching our understanding of the evolution of genome structures in RNA viruses.

Open Reading Frames

Analysis of nested alternate open reading frames and their encoded proteins.

Transcriptional and post-transcriptional mechanisms diversify the proteome beyond gene number, while maintaining a sequence relationship between original and altered proteins. A new mechanism breaks this paradigm, generating novel proteins by translating alternative open reading frames (Alt-ORFs) within canonical host mRNAs. Uniquely, 'alt-proteins' lack sequence homology with host ORF-derived proteins. We show global amino acid frequencies, and consequent biochemical characteristics of Alt-ORFs nested within host ORFs (nAlt-ORFs), are genetically-driven, and predicted by summation of frequencies of hundreds of encompassing host codon-pairs. Analysis of 101 human nAlt-ORFs of length &#x2265;150 codons confirms the theoretical predictions, revealing an extraordinarily high median isoelectric point (pI) of 11.68, due to anomalous charged amino acid levels. Also, nAlt-ORF proteins exhibit a >2-fold preference for reading frame 2 versus 3, predicted mitochondrial and nuclear localization, and elevated codon adaptation index indicative of natural selection. Our results provide a theoretical and conceptual framework for exploration of these largely unannotated, but potentially significant, alternative ORFs and their encoded proteins.

Journal Article

High-quality peptide evidence for annotating non-canonical open reading frames as human proteins.

A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.

GENCODE

The mighty microproteins: from versatile cellular regulators to precision medicine therapeutics.

Microproteins, are tiny proteins encoded by small open reading frame (sORF), translation of these non-canonical open reading frames (ncORFs) has been implicated in diverse biological processes and diseases. This review summarizes recent developments in the discovery, biogenesis, and functional characterization of microproteins, and their involvement in various disease, with special focus on their roles in cancer, cardiovascular, metabolic, neurodegenerative and immune-related disorders. We emphasize the regulation of key cellular pathways by microproteins, including mitochondrial homeostasis, apoptosis, metabolic reprogramming, and immune signaling, all of which affect disease initiation and progression. Emerging evidence also supports their potential as disease biomarkers and therapeutic candidates for precision medicine. Finally, the review critically discusses the current challenges including discrepancies in microprotein annotation, the limitations of ribosome profiling and proteogenomic approaches, the gap between computationally predicted and experimentally validated microproteins, and the need for rigorous orthogonal validation by means of CRISPR-based genome editing, ribosome release assays, mutational analysis, high-resolution mass spectrometry, and functional studies. Finally, we review recent development of AI-assisted ORF prediction, single-cell translatomics, spatial proteomics, and integrated multi-omics as emerging technologies reshaping. Microprotein discovery and functional annotation. Finally, we discuss the translational potential of microproteins and highlight the remaining challenges to clinical application, including peptide stability, pharmacokinetics, tissue-specific delivery, immunogenicity, and the need for rigorous preclinical and clinical validation. Together, this review provides an updated and critical overview of the rapidly evolving microprotein field and highlights future research priorities for translating these molecules into clinically useful biomarkers and precision therapeutics.

Microproteins

MRI roadmap-guided transendocardial delivery of exon-skipping recombinant adeno-associated virus restores dystrophin expression in a canine model of Duchenne muscular dystrophy.

Duchenne muscular dystrophy (DMD) cardiomyopathy patients currently have no therapeutic options. We evaluated catheter-based transendocardial delivery of a recombinant adeno-associated virus (rAAV) expressing a small nuclear U7 RNA (U7smOPT) complementary to specific cis-acting splicing signals. Eliminating specific exons restores the open reading frame resulting in translation of truncated dystrophin protein. To test this approach in a clinically relevant DMD model, golden retriever muscular dystrophy (GRMD) dogs received serotype 6 rAAV-U7smOPT via the intracoronary or transendocardial route. Transendocardial injections were administered with an injection-tipped catheter and fluoroscopic guidance using X-ray fused with magnetic resonance imaging (XFM) roadmaps. Three months after treatment, tissues were analyzed for DNA, RNA, dystrophin protein, and histology. Whereas intracoronary delivery did not result in effective transduction, transendocardial injections, XFM guidance, enabled 30&#xb1;10 non-overlapping injections per animal. Vector DNA was detectable in all samples tested and ranged from <1 to >3000 vector genome copies per cell. RNA analysis, western blot analysis, and immunohistology demonstrated extensive expression of skipped RNA and dystrophin protein in the treated myocardium. Left ventricular function remained unchanged over a 3-month follow-up. These results demonstrated that effective transendocardial delivery of rAAV-U7smOPT was achieved using XFM. This approach restores an open reading frame for dystrophin in affected dogs and has potential clinical utility.

Animals

Alternative genetic codes in bacteria and archaea identified with a fast k-mer-based algorithm.

The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.

Genetic Code

Isolation and characterization of a novel Schitoviridae phage VipHU7 that infects Vibrio parahaemolyticus.

In recent years, various bacteriophages that infect Vibrio spp. have been isolated and characterized. However, many characteristics concerning their infection mechanisms remain unknown. Here, we isolated and characterized a novel phage, VipHU7, that infects Vibrio parahaemolyticus. The morphology of VipHU7 was examined using transmission electron microscopy, which demonstrated that it has a short, noncontractile tail characteristic of podoviruses. VipHU7 formed clear plaques with halo zones on a bacterial lawn of V. parahaemolyticus MFS 1101, and host range analysis revealed that it had a limited host range. Analysis of the propagation and one-step growth curve of VipHU7 in liquid medium indicated that its replication rate increases in the presence of divalent cations, which did not affect its adsorption. Genome sequencing revealed that the VipHU7 genome is 76,454 bp long, with a 38.48% GC content and 112 predicted open reading frames. VipHU7 has a genomic structure similar to that of other Varunavirus phages that infect Vibrio spp. VIRIDIC analysis showed that the intergenomic similarity between VipHU7 and Vibrio phage BUCT194 was 83.7%, indicating that VipHU7 is a novel species belonging to the family Schitoviridae and genus Varunavirus.

Vibrio parahaemolyticus

Lytic properties and genomic analysis of bacteriophage Brt_Psa3, targeting Pseudomonas syringae pv. actinidiae.

Pseudomonas syringae pv. actinidiae (Psa) is the causative agent of bacterial canker in kiwifruit (Actinidia spp.). Psa biovar 3 is the most prevalent and virulent, causing frequent and severe outbreaks worldwide. While current treatments have low efficacy, bacteriophages emerge as possible environmentally safe alternative biocontrol agents. In this study, bacteriophage Brt_Psa3 was isolated from the soil of a kiwifruit orchard in Portugal. Morphologically, Brt_Psa3 forms clear plaques and has a Podoviral morphotype. The bacteriophage exhibited broad lytic activity against several plant-pathogenic Pseudomonas strains, including Psa isolates. The isolated bacteriophage has a latent period of 100&#xa0;min, a burst size of 143 particles/cell, and demonstrates stability at different temperatures and pH values found in kiwifruit orchards. In addition, Brt_Psa3 exhibited tolerance to UVA irradiation during 120&#xa0;min of incubation. Brt_Psa3 belongs to the Autographiviridae family and Ghunavirus genus, based on full-genome nucleotide alignment and supported by phylogenetic analysis of structural proteins. The phage contains 51 open reading frames with no antibiotic resistance genes identified, within a genome of 40.509 base pairs. In vitro experiments with kiwifruit leaves demonstrated significant reduction of Psa levels (40%) on leaf surfaces, highlighting the bacteriophage's therapeutic potential in managing bacterial canker in kiwifruits.

Pseudomonas syringae

Characterization of Novel Luteoviruses in Canadian Highbush Blueberries Using High-Throughput Sequencing.

The Fraser Valley of British Columbia, Canada is among the top ten blueberry producing regions globally. Viral diseases are established in the region and significantly reduce average yields. While testing for two viruses is routine, characterization of all the viruses present in the region is incomplete. We used high-throughput sequencing to obtain an unbiased overview of RNA viruses present in 97 plants collected across the region. In addition to known viruses, we identified four luteoviruses previously unidentified in the region. Two of them matched the blueberry virus L (BlVL) and blueberry virus M (BlVM). recently found in the USA, while the third constitutes a new major variant of BlVM (BlVM-2), and the fourth a new luteovirus, which we named blueberry virus N (BlVN). The genome sequences were ~5 kbp long and contained four open-reading frames similar to other luteoviruses. PCR screening revealed that these luteoviruses are widespread in the region, and that plants typically harbour more than one of these luteoviruses. While luteoviruses are typically vectored by aphids, they were also present in nursery stock, indicating that spread also occurs via vegetative propagation.

High-Throughput Nucleotide Sequencing

Complete Genome Sequence of Atractylodes virus A, a Novel Carlavirus Infecting Atractylodes lancea.

Atractylodes lancea is an important medicinal crop, but viral diseases have become increasingly severe in recent years. Viruses in the genus Carlavirus are common plant pathogens that are primarily transmitted by aphids and can cause stunted growth and reduced yields in their host plants. In this study, a novel carlavirus, tentatively named Atractylodes virus A (AVA), was identified from A. lancea plants exhibiting virus-like symptoms. The full-length genome of AVA is 8,817 nt in length and contains six open reading frames, displaying the typical genomic organization of the genus Carlavirus. The replicase and CP exhibit 53.43% and 56.72% amino acid identity, respectively, with those of previously characterized carlaviruses, indicating that AVA represents a novel species in the genus Carlavirus. This study expands our understanding of viral diversity in A. lancea and provides a scientific foundation for the prevention and control of viral diseases as well as for resistance breeding.

Genome, Viral

Molecular characterization of a novel partitivirus harboring an additional third dsRNA segment from Trichoderma harzianum.

We report the complete genome sequence of a novel partitivirus identified from Trichoderma harzianum NFCF092 strain, designated Trichoderma harzianum partitivirus 4 (ThPV4). Unlike canonical members of the family Partitiviridae, which possess a bipartite genome consisting of two double-stranded RNA (dsRNA) segments encoding an RNA-dependent RNA polymerase (RdRP) and a capsid protein (CP), ThPV4 harbors a third dsRNA segment encoding a protein of unknown function. The complete genome consists of dsRNA1 (1,950&#xa0;bp; encoding the RdRP), dsRNA2 (1,772&#xa0;bp; encoding the CP), and dsRNA3 (1,629&#xa0;bp; encoding a protein with unknown function). Sequence analysis shows that each segment possesses a single open reading frame (ORF). The deduced amino acid sequence of the RdRP shows the highest similarity (90.5% identity) to that of Trichoderma gamsii alphapartitivirus 1. Phylogenetic analyses based on the RdRP indicate that ThPV4 clusters within the genus Alphapartitivirus of the family Partitiviridae. To our knowledge. ThPV4 is the first member of the genus Alphapartitivirus identified from T. harzianum to possess an additional, conserved third dsRNA segment.

Phylogeny