PubMed HealthSearch

Biomedical subjects

Mina Ryten

Publications and source records attributed to Mina Ryten.

3 recordsLinked to original sources

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

Antisense oligonucleotide-mediated MSH3 suppression reduces somatic CAG repeat expansion in Huntington's disease iPSC-derived striatal neurons.

Expanded CAG alleles in the huntingtin (HTT) gene that cause the neurodegenerative disorder Huntington's disease (HD) are genetically unstable and continue to expand somatically throughout life, driving HD onset and progression. MSH3, a DNA mismatch repair protein, modifies HD onset and progression by driving this somatic CAG repeat expansion process. MSH3 is relatively tolerant of loss-of-function variation in humans, making it a potential therapeutic target. Here, we show that an MSH3-targeting antisense oligonucleotide (ASO) effectively engaged with its RNA target in induced pluripotent stem cell (iPSC)-derived striatal neurons obtained from a patient with HD carrying 125 HTT CAG repeats (the 125 CAG iPSC line). ASO treatment led to a dose-dependent reduction of MSH3 and subsequent stalling of CAG repeat expansion in these striatal neurons. Bulk RNA sequencing revealed a safe profile for MSH3 reduction, even when reduced by >95%. Maximal knockdown of MSH3 also effectively slowed CAG repeat expansion in striatal neurons with an otherwise accelerated expansion rate, derived from the 125 CAG iPSC line where FAN1 was knocked out by CRISPR-Cas9 editing. Last, we created a knock-in mouse model expressing the human MSH3 gene and demonstrated effective in vivo reduction in human MSH3 after ASO treatment. Our study shows that ASO-mediated MSH3 reduction can prevent HTT CAG repeat expansion in HD 125 CAG iPSC-derived striatal neurons, highlighting the therapeutic potential of this approach.

Huntington Disease

CLN3 transcript complexity revealed by long-read RNA sequencing analysis.

BACKGROUND: Batten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common pathogenic variant shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the 1-kb deletion: the 'major' and 'minor' transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. METHODS: We leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. RESULTS: We found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated 'major' transcripts are detected. Together, they have median usage of 1.5% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. CONCLUSION: Overall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the 1-kb deletion and rare pathogenic variants on CLN3 transcription and disease pathogenesis.

Humans