PubMed HealthSearch

SEARCH · PubMed Health

Results for “Long-read transcriptomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

41 records · Page 3Linked to original sources

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

Elimination of myotonia improves myopathy in a muscleblind knockout model of myotonic dystrophy.

A cardinal sign of myotonic dystrophy type 1 (DM1) is slow of muscle relaxation after voluntary contraction known as myotonia. Myotonia results from mis-regulated splicing of chloride channel 1 (ClC-1), leading to loss of channel function and runs of involuntary action potentials in muscle fibers. Heralding the onset of weakness, myotonia is often the first symptom of DM1, and raising the possibility that muscle hyperexcitability promotes the subsequent development of myopathy. We used genome editing to test this possibility by deleting the alternatively spliced and frameshift inducing ClC-1 exon 7a (E7a) in the Mbnl1 knockout model of DM1. Although several ClC-1 exons exhibit mis-regulated splicing in DM1, deletion of this single cryptic exon was sufficient to restore ClC-1 function and eliminate myotonia systemically and permanently. As determined by long-read sequencing, deletion of E7a reduced the frequency of other splicing defects in ClC-1 transcripts, likely as a passive consequence of restoring reading frame and nonsense surveillance. Furthermore, we observed significantly improved muscle force generation, fiber-type distribution, and histology, and partial restoration of the muscle transcriptome, including differential gene expression and alternative splicing, in non-myotonic Mbnl1 knockout mice. These results suggest that E7a inclusion is a lynchpin splice event that contributes to skeletal myopathy, highlighting myotonia as a therapeutic target and an outcome of interest in DM1.

Journal Article

Characterization and analysis of the full-length transcriptome of Frankliniella occidentalis (Thysanoptera: Thripidae).

BACKGROUND: Frankliniella occidentalis, an insect belonging to the order Thysanoptera, causes severe damage to agricultural and horticultural crops, resulting in significant economic losses worldwide. The development of molecular and sequencing technologies has helped elucidate the molecular mechanisms regulating its growth and development as well as its damaging activity. However, much remains to be explored. To further investigate the molecular complexity of this species, we sequenced the full-length transcriptome of mixed samples obtained from specimens at all developmental stages. RESULTS: Of all transcripts, 89.04% matched with the reference genome; additionally, 29,750 alternative splicing events, 2,342 genes with poly(A) sites, and 153 candidate fusion transcript events were identified, and 4,235 long noncoding RNAs were discovered. CONCLUSIONS: This is the first full-length transcriptome of F. occidentalis reported to date. This study greatly contributes to the understanding of the molecular complexity and diversity of this insect, providing a basis to develop specific molecular targets as well as resources for gene function studies in other insects.

Animals