PubMed HealthSearch

SEARCH · PubMed Health

Results for “Direct RNA sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Dogme: a nextflow pipeline for reprocessing nanopore RNA and DNA modifications.

MOTIVATION: Oxford Nanopore (ONT) sequencing allows for the direct detection of RNA and DNA modifications from unamplified nucleic acids, which is a significant advantage over other platforms. However, the rapid updates to ONT basecalling models and the evolving landscape of computational tools for modification detection bring about challenges for reproducible and standardized analyses. To address these challenges, we developed Dogme to automate basecalling, alignment, modification detection, and transcript quantification. Dogme automates the reprocessing of ONT POD5 files by integrating basecalling using Dorado, read mapping using minimap2 and subsequent analysis steps such as running modkit. The pipeline supports three major types of sequencing data-direct RNA (dRNA), complementary DNA (cDNA), and genomic DNA (gDNA). Dogme facilitates detection of diverse RNA modifications supported by Dorado such as N6-methyladenosine (m6A), 5-methylcytosine (m5C), inosine, pseudouridine, 2'-O-methylation (Nm) and DNA methylation, while concurrently quantifying full-length transcript isoforms LR-Kallisto for transcript quantification for dRNA and cDNA. RESULTS: We applied Dogme to three separate mouse C2C12 myoblast replicates using direct RNA sequencing on MinION flow cells. We detected 96 603 m6A, 43 476 m5C, 8829 inosine, 10 055 pseudouridine, and 30 320 Nm sites in three biological replicates. The pipeline produced reproducible modification profiles and transcript expression levels across replicates, demonstrating its utility for integrative long-read transcriptomic and epigenomic analyses. AVAILABILITY AND IMPLEMENTATION: Dogme is implemented in Nextflow and is freely available under the MIT license at https://github.com/mortazavilab/dogme, with documentation provided for installation and usage.

RNA

Analysis of Leishbuviridae from Trypanosomatids.

Over the last decade, considerable progress has been made in unraveling RNA virus diversity. This has contributed to our understanding of the evolution of these viruses, which include emerging zoonotic human pathogens. Current success has been greatly facilitated by the development of next-generation sequencing platforms instrumental for meta-transcriptomic studies. However, due to the rapid evolution of RNA viruses, there are numerous "blind spots" waiting to be explored; one of those is the RNA virome of unicellular eukaryotes. Here, we present the pipeline, which has been successfully used to characterize various types of RNA viruses, including Leishbuviridae (Bunyaviricetes, Hareavirales) in the parasitic flagellates of the family Trypanosomatidae. The pipeline relies on axenic in vitro cell culture and double-stranded RNA enrichment, followed by direct RNA-sequencing. A detailed procedure description starting from the initial total RNA preparation to the final assembly of the viral segments is provided.

High-Throughput Nucleotide Sequencing

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

Enzymes in high-throughput RNA sequencing: Applications and challenges.

High-throughput RNA sequencing provides genome-wide information on the dynamics of RNA in each cell and how the dynamics responds to environmental changes. Next-generation sequencing by the Illumina platform currently provides the highest information output as compared to other platforms. A key component of next generation sequencing of each RNA is the successful end-to-end reverse-transcription into a cDNA strand. This can be highly challenging given the propensity of each RNA to adopt ordered structures and to contain post-transcriptional modifications. While many reverse transcriptase (RT) enzymes have been developed over the years to maximize read-through of an RNA, their processivity and efficiency varies, raising the question of how to select the RT for the experiment at hand. Here, we use tRNA as a model for genome-wide sequencing, as tRNA has a stable secondary and tertiary structure and has a high density and wide variety of post-transcriptional modifications, presenting one of the most challenging problems of sequencing RNA. We compare the efficiency of end-to-end cDNA synthesis of tRNA among several recent RT enzymes and provide a general sequencing workflow that is applicable to most of these enzymes.

High-Throughput Nucleotide Sequencing

The use of nuclease P1 in sequence analysis of end group labeled RNA.

A method is described for the direct sequence analysis of 20-25 nucleotides from the termini of 5'- or 3'-end-group [32P] labeled RNA. The method involves partial endonucleolytic digestion of the labeled RNA with nuclease P1 (from Penicillium citrinum) followed by separation of the partial digestion products by two-dimensional homochromatography, the nucleotide sequence being determined by mobility shift analysis. This procedure has been applied to the sequence analysis of the terminal regions of tRNAs and of high molecular weight RNA, such as messenger RNA or viral RNA. A further application involves its use in conjunction with snake venom phosphodiesterase to determine the sequence of 5'-end group labeled oligonucleotides, containing modified bases, derived from T1 or pancreatic RNase digestion of tRNA.

Base Sequence

The structure of a transcriptional unit on colicin E1 plasmid.

In an RNA-synthesizing system in vitro, a low-molecular-weight RNA consisting of about 110 residues (RNA-I) was efficiently synthesized on DNA of colicin E 1 plasmid (ColE1) and its deletion derivatives. The promoter site for RNA-I was analysed by testing the RNA polymerase-binding ability and template activity of restriction fragments; it was mapped in the region between the replication initiation site and the colicin immunity gene of ColE1. The direction of transcription was determined by hybridization tests to the separated strands of the template. The DNA region directing RNA-I was sequenced, and RNA-I was assigned on the sequence based on the nearest-neighbour data of RNA. The sequences of its promoter and terminator regions were also deduced. Although the function of this small RNA species is unknown, a unique secondary structure could be constructed from its sequence and sensitivity to RNase.

Bacteriocin Plasmids

Genome sequencing reveals the impact of pseudoexons in rare genetic disease.

PURPOSE: Advancements in sequencing technologies have significantly improved clinical genetic testing; yet, the diagnostic yield remains around 30% to 40%. Emerging technologies are now being deployed to address the remaining diagnostic gap. METHODS: We tested whether short-read genome sequencing could increase the diagnostic yield in individuals enrolled into the UCI-GREGoR research study, who had suspected Mendelian conditions and prior inconclusive testing. Two other collaborative research cohorts, focused on aortopathy and dilated cardiomyopathy, consisted of individuals who were undiagnosed but had not undergone harmonized prior testing. RESULTS: We sequenced 353 families (754 participants) and found a molecular diagnosis in 54 (15.3%) of them. Of these diagnoses, 55.5% were previously missed because the causative variants were in regions not originally interrogated. In 5 cases, they were deep intronic variants, all of which led to abnormal splicing and pseudoexons, as directly shown by RNA sequencing. All 5 of these variants had inconclusive spliceAI scores. In 26% of newly diagnosed cases, the causal variant could have been detected by exome sequencing reanalysis. CONCLUSION: Genome sequencing can overcome limitations of clinical genetic testing, such as the inability to call intronic variants. Our findings highlight pseudoexons as a common mechanism via which deep intronic variants cause Mendelian disease.

Humans

Model-directed generation of artificial CRISPR-Cas13a guide RNA sequences improves nucleic acid detection.

CRISPR guide RNA sequences deriving exactly from natural sequences may not perform optimally in every application. Here we implement and evaluate algorithms for designing maximally fit, artificial CRISPR-Cas13a guides with multiple mismatches to natural sequences that are tailored for diagnostic applications. These guides offer more sensitive detection of diverse pathogens and discrimination of pathogen variants compared with guides derived directly from natural sequences and illuminate design principles that broaden Cas13a targeting.

CRISPR-Cas Systems

Characterization of double-stranded ribonucleic acid sequences present in the initial transcription products of rat liver chromatin.

At low ionic strength and with a low exogenous RNA polymerase/DNA ratio, rat liver chromatin directs the synthesis in vitro of RNA sequences rich in double-stranded segments. All the transcripts contain at least one double-stranded sequence. Most of the double-stranded segments are formed by intramolecular base-pairing of inverted complementary sequences separated by a single-stranded loop. They are heterogeneous in size, 35-45% of them being more than 80 nucleotides long. They contain 61-64% G+C, whether synthesized by rat liver RNA polymerase (form B) or Escherichia coli RNA polymerase. The largest double-stranded sequences are found in the largest transcripts, and are the most thermostable. The fidelity of base-matching is better in double-stranded transcripts synthesized on rat liver chromatin by homologous polymerase than in those synthesized on it by a bacterial polymerase, or in those synthesized by either of the two polymerases on pure DNA.

Animals

A redefinition of the Asp-Asp domain of reverse transcriptases.

The rules defining the Asp-Asp domain of RNA-dependent polymerases deduced by Argos (1988) were tested in a set of 53 putative reverse transcriptases (RTs) sequences. Since it was found that some of these rules are not followed by RTs coded by bacteria, group II introns, and non-LTR retrotransposons, we present here a more strict definition of the Asp-Asp domain.

Amino Acid Sequence

The lymphoproliferative disease virus of turkeys represents a distinct class of avian type-C retrovirus.

The lymphoproliferative disease virus of turkeys (LPDV) is the etiological agent of a rapidly developing lymphoproliferative process in turkeys. To better understand the genetic relationships of LPDV to other retroviruses we determined the nucleotide sequence of its pol gene. Comparative computer analyses of the deduced amino acid sequences of the reverse transcriptase and integrase domains within pol established that LPDV represents a distinct class of avian retroviruses that is most closely related to the avian leukemia-sarcoma viruses.

Amino Acid Sequence

Site specific enzymatic cleavage of RNA.

The hybridization of a DNA oligonucleotide a specific tetramer or longer) will direct a cleavage by RNase H (EC 3.1.4.34) to a specific site in RNA. The resulting fragments can then be labeled at their 5' or 3' ends, purified, and sequenced directly. This procedure is demonstrated with two RNA molecules of known sequence: 5.8S rRNA from yeast (158 nucleotides) and satellite tobacco necrosis virus (STNV) RNA (1240 nucleotides).

Animals

Temperature-dependent template switching during in vitro cDNA synthesis by the AMV-reverse transcriptase.

Reverse transcriptase template switching has been invoked to explain several aspects of retroviral replication and recombination, and has been reported in vitro for the Moloney murine leukemia virus (M-MuLV) reverse transcriptase. During in vitro cDNA synthesis, the avian myeloblastosis virus (AMV) reverse transcriptase can switch from one template to another in a homology-dependent and temperature-dependent manner. Chimeric cDNA molecules are generated within 30 min at high incubation temperatures, with an increasing efficiency from 42 degrees C to 50 degrees C. Such products are detectable only after much longer incubation times when primer extension reactions are carried out at lower temperatures (90 min at 37 degrees C).

Avian Myeloblastosis Virus

Contacts between Escherichia coli RNA polymerase and thymines in the lac UV5 promoter.

I have identified those 5 positions of thymines in the lac UV5 promoter that lie close to bound Escherichia coli RNA polymerase (nucleosidetriphosphate:RNA nucleotidyltransferase, EC 2.7.7.6). Although ultraviolet irradiation of DNA with 5-bromouracil substituted in place of thymine normally cleaves the DNA at the bromouracils, a protein bound to the DNA can perturb these cleavages at those locations at which the protein lies close to the bromine. In the lac promoter most of these contacts lie in three regions. Four contacts lie in the region where transcription initiates; four lie in the "Pribnow box," which is located about 10 base pairs upstream from the initiation site; and three more lie in the "-35 region," located about 35 base pairs upstream from the initiation site. The "Pribnow box" and the "-35 region" are regions whose sequences are partially conserved between promoters and in which most promoter mutations are located; thus, contacts in these two regions probably represent sites of sequence-specific recognition by RNA polymerase.

Base Sequence

Structure and biosynthesis of unbranched multicopy single-stranded DNA by reverse transcriptase in a clinical Escherichia coli isolate.

It has been shown that retrons, retro-elements in bacteria, produce a reverse transcriptase (RT) and multicopy single-stranded DNA (msDNA) whose 5' end is covalently linked to RNA (msdRNA) by a 2'-5' phosphodiester bond. Here, I show that a retron in clinical Escherichia coli strain 161 produces an msDNA unlinked to RNA. The msDNA produced by this retron is a 79-nucleotide-long single-stranded DNA with monophosphate on its 5' terminus. When the retron in strain 161 is cloned into E. coli K-12, the majority of msDNA produced in the clone is the same as the msDNA in the clinical strain. However, in the K-12 clone, about 10% of the msDNA produced is present as a DNA covalently linked to RNA. The DNA part of this RNA-DNA compound is an 83 nucleotides long with the same sequence as the unbranched msDNA, except for the presence of four additional nucleotides at the 5' side. From the analysis of the RNA-DNA compound and the results of in vitro synthesis, I show that the primary product of reverse transcription in this retron is an 83-nucleotide-long DNA covalently linked to RNA. This RNA-DNA compound is further processed to the final product, the 79-nucleotide-long msDNA with a terminal 5' monophosphate, by an endonucleolytic cleavage between the fourth and fifth positions of the DNA component of the RNA-DNA compound. The minimum region required for the production of such msDNA free of RNA contains only genes known to be required for the synthesis of branched msDNA-RNA compound in other retrons (msd, msr and ret). This suggests that either the RT has an endonuclease activity or that the msDNA-RNA compound is autocatalytically processed.

Amino Acid Sequence

Nucleotide sequence at the 5' terminus of the avian sarcoma virus genome.

Transcription of DNA from the RNA genome of avian sarcoma virus by RNA-directed DNA polymerase in vitro initiates on a primer (tRNATrp) located near the 5'-terminus of the viral genome. One of the major products of transcription is a single-stranded DNA chain complementary to a sequence of 101 nucleotides immediately distal to the site of initiation of DNA synthesis. We have determined the complete nucleotide sequence of this transcribed chain for the Prague strain of avian sarcoma virus, a partial sequence of the transcribed chain for the Bratislava 77 strain of avian sarcoma virus, and the sequence of a DNA transcript that is shorter than the transcribed single-stranded chain. Our data define the location of tRNATrp on the genome of avian sarcoma virus and provide the sequence of 119 nucleotides at the 5'-terminus of the genome. Portions of this sequence may be involved in the binding of RNA-directed DNA polymerase, the initiation of translation from viral messenger RNA, the extension of RNA-directed DNA synthesis from the 5'- to the 3'-terminus of viral RNA, and the integration of viral DNA into the host genome.

Avian Sarcoma Viruses