PubMed HealthSearch

SEARCH · PubMed Health

Results for “Single nucleotide variant”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans

Activity of natural single nucleotide variants of human alkyladenine DNA glycosylase AAG R145H, G163S and R197C involved in DNA binding.

Alkyladenine DNA glycosylase (AAG) is a critical enzyme in the base excision repair (BER) pathway that safeguards genome integrity by removing structurally diverse alkylated and deaminated purine lesions from DNA. It serves as a primary defense against alkylation-induced mutations, which are linked to cancer development, chronic inflammation, and neurodegenerative diseases. Single nucleotide variants (SNVs) in the gene coding region have the potential to alter the enzyme's functionality, potentially modulating the repair capacity and affecting response and prognosis following chemoradiotherapy. In our study, we investigated three SNVs that lead to amino acid class changes in regions involved in DNA substrate coordination: R145H, G163S, and R197C using an in vitro approach. Using biochemical assays and molecular dynamics simulations, we evaluated the thermal stability, DNA binding affinity, and glycosylase activity of AAG variants toward hypoxanthine (Hx) and 1,N6-ethenoadenosine (εA) containing substrates. The G163S variant showed reduced thermal stability due to the conformational strain in the β-hairpin loop that intercalates in DNA, but retained εA excision activity comparable to that of wild-type AAG, while losing activity against Hx-containing DNA. The R197C variant had a four-fold reduction in DNA binding affinity for both substrates, and was catalytically inactive, unable to excise either damaged bases. This loss of function correlated with the rearrangement of the 201-210 loop and the reorientation of Arg-201 and Arg-207, which disrupts critical DNA contacts. However, the R145H variant retained near-wild-type thermal stability and activity on both substrates, despite bioinformatic predictions of deleterious effect. Molecular dynamics simulations revealed variant-specific structural disruptions. The data obtained underscore the importance of experimental validation in assessing the functional impact SNVs.

DNA Glycosylases

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Polymorphism, Single Nucleotide

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values.

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R2 of 0.604 [0.593-0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156-5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

SARS-CoV-2

Substitutions of nucleotides at the 3' ends of COL6A1/2/3 exons induce exon skipping associated with collagen VI-related muscular dystrophies and therapeutic strategies.

PURPOSE: Collagen VI-related muscular dystrophies, characterized by proximal muscle weakness and joint contractures, are caused by pathogenic variants in the genes, COL6A1 to COL6A3. A monoallelic variant at the last nucleotide of a COL6A1 exon was initially classified as a missense variant but acted as a splicing variant, resulting in exon skipping. Here, we evaluated whether single-nucleotide variants at the 3'-ends of COL6A1 to COL6A3 exons cause aberrant splicing. METHODS: Ten relevant variants were identified in patients from our repository or public databases, and their muscle COL6A1 to COL6A3 transcripts were analyzed. The effects of the variants on splicing were also analyzed by minigene assay and SpliceAI in silico prediction. RESULTS: Transcripts from muscles of individuals with suspected collagen VI-related phenotypes showed exon skipping (skipping rate >12%). Findings of minigene assay and in silico prediction experiments supported these findings. Two therapeutic approaches, splicing correction of pre-messenger RNA or gene silencing of mature messenger RNA were assessed. Among them, gene silencing using short interfering RNAs targeting the skipped transcripts proved to be effective in restoring collagen VI in cells containing the pathogenic variant. CONCLUSION: Single-nucleotide variants at the 3'-ends of exons can lead to aberrant splicing, and allele-specific gene silencing targeting such variants is a promising therapeutic strategy.

Humans

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals

Survey of diagnostic laboratories highlights need for improved standards in somatic genomic testing and reporting.

There is a growing international need to support somatic genomic testing, standardised variant curation and improved patient access to molecular profiling for somatic conditions, including cancer. We conducted a survey of scope, curation, reporting and sharing practices of diagnostic laboratories performing somatic testing in Australia and New Zealand. Laboratories with accreditation (n = 41) were invited in 2023 to complete a semi-structured, 25-question interview. Responses were received for 27 laboratories (66% response rate) offering solid tumour, haematological malignancy and non-cancer services. Only 36% of laboratories offered tests capturing the full breadth of variants, from single-nucleotide variants to gene fusions. Knowledge sharing was rare, with only one laboratory submitting variant classifications to a public knowledge base. Most laboratories (96%) conducted somatic testing in oncology. Of cancer laboratories, 35% offered testing considered capable of comprehensive genomic profiling (CGP). Almost half of cancer laboratories had already adopted the 2022 ClinGen/CGC/VICC oncogenicity guidelines, and 84% were using AMP/ASCO/CAP 2017 clinical significance guidelines. Only 47% of mixed discipline cancer laboratories reported biomarkers such as tumour mutational burden, with wide variation in reporting of matched therapy options. Our study has generated a unique overview of somatic laboratory practices in the region, and areas for global standardisation in somatic molecular testing and reporting. We also provide a model for practice and guideline uptake assessment, for application by other country-wide networks. This is particularly relevant in anticipation of CGP mainstreaming, with the increasing complexity of sequencing interpretation for laboratories and clinicians.

Humans

Backtracking Cell Phylogenies in the Human Brain with Somatic Mosaic Variants.

Somatic mosaic variants, and especially somatic single nucleotide variants (sSNVs), occur in progenitor cells in the developing human brain frequently enough to provide permanent, unique, and cumulative markers of cell divisions and clones. Here, we describe an experimental workflow to perform lineage studies in the human brain using somatic variants. The workflow consists in two major steps: (1) sSNV calling through whole-genome sequencing (WGS) of bulk (non-single-cell) DNA extracted from human fresh-frozen tissue biopsies, and (2) sSNV validation and cell phylogeny deciphering through single nuclei whole-genome amplification (WGA) followed by targeted sequencing of sSNV loci.

Humans

Case Report: Early infantile drug-resistant epilepsy and gastrointestinal dysmotility associated with a de novo GNAO1 variant and a 16p13.11 microdeletion.

BACKGROUND: The GNAO1 gene is located on the long arm of chromosome 16 (16q13) and is associated with Developmental and Epileptic Encephalopathy 17. It encodes a protein involved in regulating various neurotransmitters and neuronal growth and development. The 16p13.11 microdeletion is a genomic deletion in the p13.11 region of the short arm of chromosome 16, previously reported to predispose individuals to neurodevelopmental disorders. These two regions do not overlap on the chromosome, and the 16p13.11 microdeletion does not involve the GNAO1 gene locus. CASE PRESENTATION: We reports a 1-month-old infant presenting with concurrent GNAO1 (p. G203R) gene variation and a 16p13.11 microdeletion. The patient initially presented with recurrent abdominal distension in the neonatal period, which gradually progressed to include focal seizures and intractable epileptic spasms. Interictal EEG showed multifocal discharges and burst-suppression patterns. Brain MRI revealed no abnormalities. A full gastrointestinal series suggested intestinal obstruction and gastro-esophageal reflux. Improved whole-exome sequencing identified: 1) A single nucleotide variantion: GNAO1, NM_020988.3:exon6: c.607G > A (p.Gly203Arg). 2) A copy number variation: A pathogenic 1.453 Mb deletion in the 16p13.11 region. CONCLUSIONS: We describe the developmental trajectory of an infant with overlapping features of gastrointestinal symptoms and drug-resistant epilepsy, carrying a GNAO1 variant and a 16p13.11 microdeletion. This case provides insights into the contribution of both single nucleotide variant and copy number variant to the phenotype.

16p13.11 microdeletion

Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants.

Variant interpretation in the era of massively parallel sequencing is challenging. Although many resources and guidelines are available to assist with this task, few integrated end-to-end tools exist. Here, we present the Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE), a web- and cloud-based platform for annotation, identification, and classification of variations in known or putative disease genes. Starting from a set of variants in variant call format (VCF), variants are annotated, ranked by putative pathogenicity, and presented for formal classification using a decision-support interface based on published guidelines from the American College of Medical Genetics and Genomics (ACMG). The system can accept files containing millions of variants and handle single-nucleotide variants (SNVs), simple insertions/deletions (indels), multiple-nucleotide variants (MNVs), and complex substitutions. PeCanPIE has been applied to classify variant pathogenicity in cancer predisposition genes in two large-scale investigations involving >4000 pediatric cancer patients and serves as a repository for the expert-reviewed results. PeCanPIE was originally developed for pediatric cancer but can be easily extended for use for nonpediatric cancers and noncancer genetic diseases. Although PeCanPIE's web-based interface was designed to be accessible to non-bioinformaticians, its back-end pipelines may also be run independently on the cloud, facilitating direct integration and broader adoption. PeCanPIE is publicly available and free for research use.

Child

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software

Variants in the interferon regulatory factor 5 gene confer genetic risk for systemic lupus erythematosus in a Han Chinese population.

BACKGROUND: Interferon regulatory factor 5 (IRF5), integral to interferon signaling pathways, has been identified as a susceptibility locus for systemic lupus erythematosus (SLE). Nevertheless, the relationship between IRF5 variants and SLE risk within the Han Chinese demographic remains inadequately characterized. MATERIALS AND METHODS: Genotyping of two functional single nucleotide variants (SNVs) in IRF5 was conducted in 167 individuals with SLE and 246 healthy controls utilizing sequence-specific primer polymerase chain reaction (PCR-SSP). Chi-square and Fisher's exact tests were employed to assess associations. RESULTS: The rs10954213 variant demonstrated a significant association with SLE susceptibility under the recessive model (GG vs. AG+AA, OR = 2.20, 95% CI: 1.30-3.75, p&#x2009;=&#x2009;0.003, adjusted p [pc]&#x2009;=&#x2009;0.030) and homozygous model (GG vs. AA, OR = 2.43, 95% CI: 1.36-4.42, p&#x2009;=&#x2009;0.003, pc = 0.032). Similarly, the rs2004640 variant was associated with an increased risk of SLE across allelic (T vs. G, OR = 1.66, 95% CI: 1.22-2.26, p&#x2009;=&#x2009;0.001, pc = 0.011), dominant (TG+TT vs. GG, OR = 1.77, 95% CI: 1.19-2.63, p&#x2009;=&#x2009;0.005, pc = 0.047), and homozygous models (TT vs. GG, OR = 3.72, 95% CI: 1.58-8.78, p&#x2009;=&#x2009;0.002, pc = 0.016). Haplotype analysis identified protective haplotype HT1 (A/G, OR = 0.54, 95% CI: 0.41-0.73, p&#x2009;<&#x2009;0.001) and risk haplotype HT4 (G/T, OR = 2.51, 95% CI: 1.42-4.42, p&#x2009;=&#x2009;0.001). CONCLUSIONS: These findings indicate that IRF5 gene variants substantially modulate susceptibility to SLE in the Han Chinese population. They hold potential as biomarkers for evaluating SLE risk and offer valuable perspectives into disease pathogenesis.

Adult

No receptor-binding domain adaptation detected in within-host H5N1 surveillance of 4,559 US dairy outbreak sequences.

BACKGROUND: The 2024-2026 US H5N1 clade 2.3.4.4b dairy cattle outbreak has been characterised primarily through consensus-level phylogenetics. Whether mammalian-adaptation variants are emerging at sub-consensus frequencies within infected hosts, particularly at the haemagglutinin receptor-binding domain (RBD), remains unknown because no systematic within-host variant analysis of the public sequencing corpus has been performed. METHODS: We conducted a pre-registered, corpus-wide intrahost single-nucleotide variant (iSNV) analysis of all publicly available H5N1 cattle, feline-spillover, and retail-milk sequences on the NCBI Sequence Read Archive (4559 samples across 7 BioProjects). A dual-caller concordance pipeline (iVar&#xa0;+&#xa0;LoFreq) with empirically determined allele frequency (AF) threshold (3%, set via four-criterion validation including synthetic spike-in controls) was applied to an 11-site Tier 1 mammalian-adaptation panel spanning the polymerase complex, haemagglutinin RBD, and accessory proteins. Within-host nucleotide diversity was compared across host categories. RESULTS: The HA RBD sites Q226L and G228S (H3 numbering) showed zero detections across >4300 adequately sequenced samples at all AF thresholds tested (1-5%), despite the pipeline detecting other non-synonymous variants at these exact codon positions (upper 95% CI for prevalence: 0.08%). Seven of eleven adaptation sites carried statistically significant iSNV signals after Bonferroni correction (corrected &#x3b1;&#x202f;=&#x202f;0.00417), though all at low prevalence (&#x2264;2.95%). Genotype stratification showed that most polymerase-site detections reflected genotype structure rather than within-host emergence: the apparent PB2 631&#x202f;L&#x2192;M "reversion" was largely the ancestral avian state of the D1.1 genotype (20 of 23 detections), which never acquired the 631L mammalian adaptation, with only two genuine sub-consensus events in the B3.13 background, while consensus-level PB2 701N was a fixed feature of the D1.1 genotype (10 of 14 detections) rather than independent sub-consensus emergence. Cattle exhibited significantly higher within-host nucleotide diversity than feline-spillover samples (&#x3c0;&#x202f;=&#x202f;1.59&#x202f;&#xd7;&#x202f;10-4 vs 6.11&#x202f;&#xd7;&#x202f;10-5; Kruskal-Wallis p&#x202f;=&#x202f;6.6&#x202f;&#xd7;&#x202f;10-15), a finding that persisted after depth-matching (p&#x202f;=&#x202f;4.6&#x202f;&#xd7;&#x202f;10-5); this may reflect prolonged mammary-gland infection, though sampling differences and host biology cannot be excluded. CONCLUSIONS: We did not detect HA receptor-switching adaptation (the acquisition of human-type &#x3b1;2,6 receptor binding via Q226L/G228S) at any tested allele frequency in the US dairy H5N1 outbreak. Sub-consensus mammalian-adaptation signals exist at polymerase-complex sites but at low prevalence, are genotype-structured rather than independently recurrent, and require functional characterisation before informing risk assessment.

Dairy cattle

Cell-type-specific enrichment of somatic aneuploidy in the mammalian brain.

Somatic mutations alter the genomes of a subset of an individual's brain cells, impacting gene regulation and contributing to disease processes. Mosaic single-nucleotide variants have been characterized with single-cell resolution in the brain, but we have limited information about large-scale structural variation such as whole-chromosome duplication or loss. We used a dataset of over 415,000 single-cell DNA methylation and chromatin conformation profiles from the adult mouse brain to comprehensively identify and characterize aneuploid cells. Somatic trisomy events were strongly enriched on chromosome 16, which is syntenic with human chromosome 21. We also observed a specific enrichment of chromosome gain and loss events in specific cell types, including Pons neurons and oligodendrocyte precursor cells. Chromosome 16 trisomy occurred in multiple cell types and across brain regions, suggesting that nondisjunction is a recurrent feature of somatic structural variation in the brain.

Animals

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans