PubMed HealthSearch

SEARCH · PubMed Health

Results for “Non-Coding RNAs”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Molecular cloning and nucleotide sequence of the cDNA for sperm-specific lactate dehydrogenase-C from mouse.

Mouse sperm-specific lactate dehydrogenase-C (LDH-C) cDNA was cloned and sequenced from lambda gt11 expression library. The LDH-C cDNA insert of 1236 bp consists of the protein-coding sequence (999 bp), the 5' (54 bp) and 3' (113 bp) non-coding regions, and the poly(A) tail (70 bp). The Northern blot analysis of poly(A)-containing RNAs from mouse testes and liver indicates that the LDH-C gene is expressed in testes but not in liver, and that its mRNA is approx. 1400 nucleotides in length. The nucleotide and amino acid sequences of the mouse LDH-C cDNA show 73% and 72% homologies, respectively, with those of the mouse LDH-A. The Southern blot analysis of genomic DNAs from mouse liver and human placenta indicates the presence of multiple LDH-C gene-related sequences.

Amino Acid Sequence

Evolutionary relationships in the cucumoviruses: nucleotide sequence of tomato aspermy virus RNA 1.

RNA 1 of the V strain of tomato aspermy virus (TAV) consists of 3410 nucleotides and contains one open reading frame (ORF) of 2982 nucleotides, resembling RNA 1 of cucumber mosaic virus (CMV) strains Q and Fny (68% and 66% identical, respectively) and of brome mosaic virus (BMV) (41% identical). In comparisons between amino acid sequences, three conserved regions (N-terminal, C-terminal and central) between TAV and each CMV were found. The N- and C-terminal regions were also conserved with BMV, and contained, respectively, consensus motifs for methyltransferases and for nucleic acid helicases. The 5' and 3' non-coding sequences were highly similar to those of TAV RNA 2. When the sequences for the genomic RNAs of the V and C strains of TAV, and of their encoded products, are compared with those reported for CMV strains representing either subgroup I (Fny-CMV) or subgroup II (Q-CMV) of CMV, it was found that the different virus-encoded proteins are conserved differently between these three viruses. Also, the divergence between TAV and both CMV subgroups has proceeded at different rates for the different ORFs. On the whole, the divergence between TAV and CMV is of the same order as that found between CMV subgroups I and II, which suggests that TAV, Q-CMV and Fny-CMV could be considered as representing three equivalent subgroups of a taxonomic entity.

Amino Acid Sequence

Present status of the sugarcane mosaic subgroup of potyviruses.

Until recently, sugarcane mosaic virus (SCMV) was believed to be a single potyvirus consisting of a large number of strains, differing from each other in certain biological and antigenic properties. The use of affinity-purified polyclonal antibodies directed towards the surface-located, virus-specific amino termini of the coat proteins showed that 17 strains from Australia and the United States represented four distinct potyviruses, namely johnsongrass mosaic virus (JGMV), maize dwarf mosaic virus (MDMV), sorghum mosaic virus (SrMV) and SCMV. Comparisons of strains from each of these four viruses on the basis of reactions on differential sorghum and oat cultivars, cell-free translation of RNAs, morphology and serology of cytoplasmic cylindrical inclusions, amino acid sequence and peptide profiling of coat proteins, 3' non-coding nucleotide sequences, and molecular hybridization with probes corresponding to the 3' non-coding regions, resulted in exactly the same taxonomic assignments as obtained using amino-terminal serology. These results further confirm that the former sugarcane mosaic virus actually consists of four distinct viruses and show that MDMV, SrMV, and SCMV are more closely related to each other than they are to JGMV. Because these four viruses are closely related but distinct, formation of a sugarcane mosaic subgroup in the genus Potyvirus would be appropriate.

Edible Grain

Whole transcriptome sequencing analyses of islets reveal ncRNA regulatory networks underlying impaired insulin secretion and increased β-cell mass in high fat diet-induced diabetes mellitus.

AIM: Our study aims to identify novel non-coding RNA-mRNA regulatory networks associated with β-cell dysfunction and compensatory responses in obesity-related diabetes. METHODS: Glucose metabolism, islet architecture and secretion, and insulin sensitivity were characterized in C57BL/6J mice fed on a 60% high-fat diet (HFD) or control for 24 weeks. Islets were isolated for whole transcriptome sequencing to identify differentially expressed (DE) mRNAs, miRNAs, IncRNAs, and circRNAs. Regulatory networks involving miRNA-mRNA, lncRNA-mRNA, and lncRNA-miRNA-mRNA were constructed and functions were assessed through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. RESULTS: Despite compensatory hyperinsulinemia and a significant increase in β-cell mass with a slow rate of proliferation, HFD mice exhibited impaired glucose tolerance. In isolated islets, insulin secretion in response to glucose and palmitic acid deteriorated after 24 weeks of HFD. Whole transcriptomic sequencing identified a total of 1324 DE mRNAs, 14 DE miRNAs, 179 DE lncRNAs, and 680 DE circRNAs. Our transcriptomic dataset unveiled several core regulatory axes involved in the impaired insulin secretion in HFD mice, such as miR-6948-5p/Cacna1c, miR-6964-3p/Cacna1b, miR-3572-5p/Hk2, miR-3572-5p/Cckar and miR-677-5p/Camk2d. Additionally, proliferative and apoptotic targets, including miR-216a-3p/FKBP5, miR-670-3p/Foxo3, miR-677-5p/RIPK1, miR-802-3p/Smad2 and ENSMUST00000176781/Caspase9 possibly contribute to the increased β-cell mass in HFD islets. Furthermore, competing endogenous RNAs (ceRNA) regulatory network involving 7 DE miRNAs, 15 DE lncRNAs and 38 DE mRNAs might also participate in the development of HFD-induced diabetes. CONCLUSIONS: The comprehensive whole transcriptomic sequencing revealed novel non-coding RNA-mRNA regulatory networks associated with impaired insulin secretion and increased β-cell mass in obesity-related diabetes.

Mice

AUUUA motifs are dispensable for rapid degradation of the mouse c-myc RNA.

Sequence determinants responsible for c-myc RNA rapid turn-over are localized within the 3' non-coding region which is mainly characterized by the presence of two polyadenylation signals and a high content in A and U. Although the AUUUA/UUAUUUA motif is commonly thought to specify a whole class of unstable RNAs coding for various onco-proteins and cytokines, site-directed mutagenesis showed that both of the two such sequences found in the mouse c-myc RNA are dispensable for rapid RNA degradation. Although less efficient than the whole 3' non-coding region, the last 50 nucleotides of c-myc RNA, mainly made up of U and A and devoid of AUUUA/UUAUUUA motif, are sufficient to confer instability to the coding sequence.

Animals

Mapping and characterization of Paracentrotus lividus mitochondrial transcripts: multiple and overlapping transcription units.

This paper reports the mapping of both mature and precursor Paracentrotus lividus mitochondrial transcripts. Several mtRNAs were found to have 5' and 3' termini which differ from those inferred through DNA sequencing (Cantatore et al. 1989). The 3' ends of the two rRNAs (12S and 16S) overlap with the downstream transcripts (tRNAGlu and CoI mRNA) by 5 and 10 nt respectively. The 132 nt non-coding region is extensively transcribed: in particular it contains a 124 nt RNA and the 5' end of a possible precursor of 13 clustered tRNAs. This latter overlaps by 7 nt with the 3' end of the 124 nt RNA. In addition to the mature RNAs, 32 high molecular weight RNAs, which are probably the precursors of the smaller more abundant mature species, were detected by Northern blotting. The mapping of these transcripts indicates that they are processed at the level of tRNA or tRNA-like sequences and suggests the existence of two transcription initiation sites upstream of the ND1 and the cytochrome b genes respectively. In the light of these results it appears that P. lividus mitochondrial DNA transcription takes place via multiple and probably overlapping transcription units. Moreover, the wide variation in the steady-state levels of the mature mRNAs indicates that sea urchin mitochondrial DNA expression is also regulated at the level of RNA decay.

Animals

The human embryonic myosin alkali light chain gene: use of alternative promoters and 3' non-coding regions.

Recently we have found evidence that the human embryonic myosin alkali light chain (MLC1 emb) gene has two functional promoters and that its mRNAs exhibit heterogeneity in their 3'untranslated regions (UTR). To study this more in detail we have isolated and characterized the human MLC1emb gene. We focussed in particular on 2 kilobases of 5'flanking region and the alternative 3'UTRs. RNA primer extension and S1 mapping analyses revealed that the MLC1emb gene can indeed be driven either by a proximal or a distal promoter, both in fetal and adult cardiac tissue. These MLC1emb RNAs can contain either the proximal or distal 3'UTR. In contrast to this, in fetal as well as adult masseter muscle MLC1emb mRNA is predominantly transcribed from the proximal promoter and contains mainly the distal 3'UTR. These results explain the known heterogeneity of MLC1emb mRNAs. Finally, we present evidence that the murine MLC1emb gene also contains a functional distal promoter element which has hitherto been undetected.

Animals

Role of RNA structures in c-myc and c-fos gene regulations.

Proto-oncogenes c-myc and c-fos are subjected to a complex set of controls operating both at the transcriptional and post-transcriptional levels. We report here that: (i) antisense transcription occurs at the murine c-myc locus. However, its biological significance remains to be established; (ii) transcription of both genes is regulated in various situations by a block to elongation of nascent RNA chains. In the case of c-myc, the blockade involves a RNA structure whose nature remains unknown; (iii) elements responsible for the high degree of instability of c-myc and c-fos mRNAs reside in their 3' non-coding regions. A U-rich region, reminiscent of that present in the granulocyte-monocyte colony-stimulating factor mRNA destabilizer, is likely to be involved in the rapid degradation of c-fos mRNA; (iv) exon 1 substitution by intron 1-derived sequences lessens or negates the effect of the 3' destabilizer in abnormal c-myc RNAs from Burkitt's lymphomas and mouse plasmacytomas.

Animals

Identification and external validation of a prognostic signature based on myeloid-derived suppressor cells-related LncRNAs to evaluate survival prognosis and treatment efficacy in invasive breast carcinoma.

BACKGROUND: Originating in the hematopoietic tissue, myeloid-derived suppressor cells (MDSCs) significantly contribute to tumor-related immunological processes. However, their relationship with long noncoding RNAs (lncRNAs) and breast cancer remains incompletely understood. In this study, we introduced MDSCs-associated lncRNAs as novel prognostic biomarkers to assess outcomes in patients with invasive breast carcinoma (BRCA). METHODS: Information regarding BRCA cases, including clinical and genomic details, was obtained from the TCGA repository. Predictive indicators were discovered, and their reliability underwent thorough verification. A clinically useful nomogram was developed following application-based validation. Additional investigations encompassed functional analysis, TMB assessment, TME profiling, immunotherapy efficacy forecasting, and drug sensitivity testing along with target identification. Long non-coding RNA expression was measured using reverse transcription quantitative PCR. RESULTS: A risk stratification model incorporating eight MDSCs-related lncRNAs effectively predicted patient outcomes. Kaplan-Meier (K-M) survival analysis clearly indicated a much worse prognosis among patients classified as high-risk (p&#xa0;<&#xa0;0.001). The nomogram accurately forecasted overall survival (OS). Analysis of functional enrichment revealed that pathways associated with epithelial cells showed activity among patients at higher risk. Characterization of the tumor microenvironment showed increased immune cell presence in those classified as low-risk. Conversely, individuals with greater risk displayed higher tumor mutational burden. TIDE and IPS analyses indicated superior immunotherapy responsiveness in the low-risk BRCA subgroup. Among 47 drugs with notable IC50 variations, Ribociclib, PD173074, KU-55933, NU7441, and nutlin-3a exhibited lower IC50 values within the low-risk group, whereas Lapatinib demonstrated greater efficacy among the high-risk group. Moreover, 10 potential therapeutic agents and their targets were predicted for high-risk patients. RT-qPCR validation confirmed the robustness of the model. CONCLUSIONS: We successfully verified a new model of molecular markers of MDSCs-related lncRNAs, offering critical insights for predicting outcomes and guiding therapeutic decisions in BRCA cases.

Bioinformatics

RNase III cleavages in non-coding leaders of Escherichia coli transcripts control mRNA stability and genetic expression.

The primary transcripts of the rpsO-pnp, rnc-era-recO and metY-nusA-infB operons of E coli are each processed by RNase III, upstream of the first translated gene, in hair-pin structures formed by the 5' non-coding leader. The mRNAs of the 3 operons, of which the 5' terminal motifs have been removed by RNase III, decay significantly more rapidly than the uncut transcripts which accumulate in the RNase III deficient strain. The rapid decay of a primary transcript of the metY-nusA-infB operon, initiated at a secondary promoter in the vicinity of the RNase III sites, suggests that the 5' features upstream of the RNase III cutting sites are responsible for the stability of the uncut RNAs. RNase III autocontrols its own expression by removing the 5' motif which stabilizes its mRNA. Similarly, the synthesis of polynucleotide phosphorylase and of protein Era are also controlled by RNase III cleavages which trigger the degradation of their messengers. The role of RNase III in the regulation of gene expression and the possible mechanisms of mRNA stabilization and of 5' to 3' decay initiated by RNase III processing are discussed.

Base Sequence

Beyond Canonical Neoantigens: Emerging Technologies for Identification of Noncanonical Antigens and Implications for Personalized Cancer Vaccines.

Over the past decade, advances in sequencing technologies and computational pipelines enabled the development of personalized cancer vaccines (PCVs). Current PCV strategies primarily target cancer neoantigens generated by non-synonymous DNA mutations, which can result in altered amino acid sequences capable of eliciting tumor-specific immune responses. More recently, a distinct class of tumor-specific antigens (TSA), termed noncanonical or cryptic antigens, has emerged as an additional source of immunogenic targets. Unlike canonical neoantigens, noncanonical antigens typically cannot be identified by tumor/normal whole-exome sequencing, as they do not arise from classical DNA mutations. Instead, they are often associated with less well recognized and/or aberrant processes in the pathways from DNA to human leukocyte antigen (HLA)-presented peptides. Examples include transposable elements, circular RNA, translation of alternative open reading frames and/or long non-coding RNA, among others. Emerging evidence suggests that noncanonical antigens represent a substantial portion of the tumor-specific immunopeptidome and, similar to canonical neoantigens, are absent during thymic selection and can evade central tolerance and elicit T cell responses. Technological advances have increasingly facilitated the identification of noncanonical antigens. Long-read RNA sequencing reveals noncanonical transcripts by improving transcriptome assembly, while ribosome profiling provides genome-wide maps of actively translated regions, facilitating the discovery of peptides from aberrant translation events. Specialized molecular approaches enable enrichment and sequencing of circular RNAs, and immunopeptidomics using mass spectrometry allows for direct characterization of HLA-presented peptides. Together, these technological advances have led to an increasing interest in prioritizing and targeting noncanonical antigens in the next generation of PCVs. This review provides an overview of the diverse origins of TSAs beyond classical neoantigens and discusses emerging approaches that may enable the integration of these antigens in future clinical trials.

circular RNA

Genome-wide epigenomic atlas and multi-omics responses of Eriocheir sinensis to natural extreme heat.

BACKGROUND: Global climate warming has led to increasingly frequent and prolonged extreme summer heat events, posing severe environmental challenges to aquaculture systems. Extreme summer heat can disrupt the performance of pond-cultured ectotherms. The Chinese mitten crab (Eriocheir sinensis) is an economically important freshwater crustacean, but coordinated molecular differences following contrasting natural summers remain incompletely characterized. RESULTS: We performed a comprehensive multi-omics analysis integrating meteorological monitoring, mRNA/lncRNA transcriptomics, small-RNA profiling of miRNAs, DNA methylomics, and LC-MS metabolomics in E. sinensis populations collected from Yancheng, China, between 2020 and 2024. Across the ten farms, survival was significantly lower in 2024, whereas yield and the proportion of large individuals showed nonsignificant downward trends. Gene-set analyses showed negative enrichment of cellular heat-response, protein-folding, oxidative-phosphorylation, and mitochondrial ATP-production terms in the 2024 cohort at the time of sampling. The integrated transcript annotation contained 72,240 lncRNAs and 63,833 mRNAs, and CpG was the predominant methylation context. Differential methylation analysis identified 73 regions and 185 cytosines, with hypomethylated events predominating within the significant subset. Metabolomic profiles differed between annual cohorts and mapped to carbohydrate, lipid, and amino-acid pathways. Cross-omics integration prioritized eight candidate genes-ADCY9, UNC79, UBN1, IFT52, ACO2, LOC126986070, LOC127001126, and LOC126997895-and qPCR reproduced the reported directions of expression for selected RNAs. CONCLUSION: This study provides the first integrative multi-omics framework for understanding chronic heat adaptation in E. sinensis. By linking transcriptomic, epigenomic, and metabolic remodeling, we elucidate the molecular mechanisms underlying energy imbalance, epigenetic reprogramming, and immune dysregulation during prolonged thermal stress. These findings offer valuable insights and genomic resources for breeding heat-tolerant crab strains and improving aquaculture resilience under ongoing climate change.

DNA methylation

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

Role of lncRNA PVT1 in the progression of urological cancers: Novel insights into signaling pathways and clinical opportunities.

Urologic malignancies, encompassing cancers of the kidney, bladder, and prostate, represent approximately 25&#xa0;% of all cancer cases. Recent advances have enhanced our understanding of PVT1's crucial functions. Long noncoding RNAs influence both the onset and development of cancer, as well as epigenetic alterations. Recent findings have focused on PVT1's mechanism of action across several malignancies, particularly urologic cancers. Understanding the various functions of PVT1 linked to cancer is necessary for the development of cancer detection and treatment when PVT1 is dysregulated. Furthermore, recent advancements in genomic and epigenetic research have elucidated the complex regulatory networks that control PVT1 expression. Comprehending the intricate role of PVT1 Understanding the complex function of PVT1 in urologic cancers has substantial clinical implications. Here, we summarize some of the most recent findings about the carcinogenic effects of PVT1 signaling pathways and the possible treatment strategies for urological malignancies that target these pathways.

Humans

Comprehensive circRNA expression profile and hub genes screening during human liver development.

BACKGROUND: Understanding the expression of non-coding RNA in the liver during embryonic development provides important insights into liver diseases. Therefore, we investigated circular RNA (circRNA) roles in human liver development, an unexplored research domain. METHODS: Using high-throughput sequencing and bioinformatics, we analysed foetal liver samples across developmental stages (7-20&#x2009;weeks post-conception). Differentially expressed (DE) genes were identified and subjected to enrichment analysis using Gene Ontology (GO), Kyoto Encyclopaedia of Genes and Genomes (KEGG), and Disease Ontology (DO). Modular analysis was performed using the Search Tool for Retrieval of Interacting Genes (STRING), followed by construction of a protein-protein interaction (PPI) network using Cytoscape software. The key genes were screened using Molecular Complex Detection (MCODE). The mRNA levels of hub genes were validated using quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: There were 645 DE circRNAs and 5,145 DE mRNAs between human livers at the three growth stages (HB, EH, and LH). It was found that the activity of circRNAs was boosted remarkably in the hepatoblastic stage. Enrichment analysis found they mainly involved in nervous system regulation of liver function, embryonic organ development and digestive system development. In addition, DE circRNAs were primarily involved in the PI3K-AKT, MAPK and calcium pathways, potentially contributing to adult liver diseases. Notably, only hsa_circ_001471 and novel_circ_017382 were simultaneously identified at all stages and were persistently downregulated. A co-expression regulatory network involving these circRNAs was established. Three hub genes (LGR5, FOXL1 and RSPO3) were identified from the PPI network of 167 genes and may play key roles in human liver development. The RT-qPCR validation results were in agreement with the sequencing data. CONCLUSIONS: Our findings provide the first insights into the roles and regulatory networks of circRNAs in human liver development, laying the groundwork for further investigations of molecular and signalling networks.

Humans

The complete and symmetric transcription of the main non coding region of rat mitochondrial genome: in vivo mapping of heavy and light transcripts.

The experiments here reported demonstrate that the main non-coding region of rat mitochondrial DNA is symmetrically transcribed. We have identified stable heavy and light transcripts, whose pattern is rather complex, in the D-loop region of rat mitochondrial DNA. Their relative concentrations have been determined. We detected heavy transcripts which encompass the whole D-loop and more abundant heavy RNA species which we interpreted as transcripts terminating downstream of the 3' end of the last coded gene (Thr-tRNA). The processed heavy RNA species contain polyA, suggesting a strict association between cleavage and polyadenylation. The pattern of light transcripts shows a long RNA, which, starting from the light strand promoter, covers the whole segment, and shorter RNA species which seems to be actively processed at the level of the conserved sequence boxes, probably acting as primers. The symmetric transcription of the D-loop containing region of rat mitochondrial DNA, and in particular the presence of stable transcripts complementary to the putative RNA primers, suggest that mechanisms mediated by interaction between complementary transcripts (antisense RNAs) might play a role in the regulation of mitochondrial DNA replication and expression.

Amino Acid Sequence

Marburg virus, a filovirus: messenger RNAs, gene order, and regulatory elements of the replication cycle.

The genome of Marburg virus (MBG), a filovirus, is 19.1 kb in length and thus the largest one found with negative-strand RNA viruses. The gene order - 3' untranslated region-NP-VP35-VP40-GP-VP30-VP24-L-5' untranslated region-resembles that of other non-segmented negative-strand (NNS) RNA viruses. Six species of polyadenylated subgenomic RNAs, isolated from MBG-infected cells, are complementary to the negative-strand RNA genome. They can be translated in vitro into the known structural proteins NP, GP (non-glycosylated form), VP40, VP35, VP30 and VP24. At the gene boundaries conserved transcriptional start (3'-NNCUNCNUNUAAUU-5') and stop signals (3'-UAAUUCUUUUU-5') are located containing the highly conserved pentamer 3'-UAAUU-5'. Comparison with other NNS RNA viruses shows conservation primarily in the termination signals, whereas the start signals are more variable. The intergenic regions vary in length and nucleotide composition. All genes have relatively long 3' and 5' end non-coding regions. The putative 3' and 5' leader RNA sequences of the MBG genome resemble those of other NNS RNA viruses in length, conservation at the 3' and 5' ends, and in being complementary at their extremities. The data support the concept of a common taxonomic order Mononegavirales comprising the Filoviridae, Paramyxoviridae, and Rhabdoviridae families.

Base Sequence

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans