PubMed HealthSearch

SEARCH · PubMed Health

Results for “Direct RNA sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Translating single-cell RNA sequencing into monocyte direct leukocyte subpopulation-transcript abundance assay ratio-based biomarkers (IFI27/PSAP or IFI27/CTSS) for clinical detection of viral infection.

A rapid method for triaging febrile patients by aetiology (e.g., viral or bacterial infection) using gene expression in peripheral blood (PB) is an intensively researched area. However, gene expression in blood represents a composite sum of gene expression of all the component cell types present in the sample. As a result, numerous genes are measured in most proposed signatures. Herein, we propose a simple ratio-based biomarker (RBB) called direct leukocyte subpopulation-transcript abundance assay (DIRECT LS-TA) that recapitulates gene expressions of a single cell type in PB (i.e., monocytes). Based on single-cell RNA sequencing (scRNAseq) data and bulk expression data, IFI27 and SIGLEC1 are found as interferon-stimulated genes (ISGs) predominantly expressed by monocytes. The DIRECT LS-TA method can use a simple ratio of two genes measured in PB as an RBB to represent the target gene expression in monocytes without the need for monocyte purification. Both scRNAseq and bulk RNA sequencing datasets were used to evaluate the correlation between ISG expression in monocytes and PB, with a particular focus on monocyte expression of IFI27. An iceberg plot of bulk transcriptome data was used to identify genes that were predominantly expressed by monocytes in PB. DIRECT LS-TA RBBs of the three genes (IFI27, IFI44L and SIGLEC1) were evaluated by group-wise comparison, receiver operating characteristic and meta-analysis. In addition, the conventional interferon (IFN) score was evaluated for comparison of diagnostic performance. In viral infection datasets, DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) was most intensely activated (p value by t test <1e-9) and had the best area under the curve (0.94) among the three potential monocyte ISGs analysed. DIRECT LS-TA SIGLEC1 was also another monocyte biomarker but showed a lower activation (p<9e-5). IFI27/PSAP showed better diagnostic performance than the conventional IFN score. On the other hand, IFI44L was not a predominant monocyte expression gene. DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) measured in PB was the best biomarker of viral infection and IFN activation among ISGs predominantly expressed by monocytes. It performed even better than the conventional IFN score which required quantification of eight genes. The results suggest that DIRECT LS-TA of IFI27 is a monocyte-informative biomarker which is easy to determine in PB without the need for cell sorting.

Humans

Analysis of formaldehyde-induced Adh mutations in Drosophila by RNA structure mapping and direct sequencing of PCR-amplified genomic DNA.

Two formaldehyde-induced mutations at the Drosophila Adh locus (Adhfn45 and Adhfn46) were analyzed by determining RNA structures at different developmental stages, polymerase chain reaction (PCR) amplification of the affected genomic regions, and direct sequencing of the resulting double-stranded DNA fragments. Adhfn46 adults and larvae accumulate abundant ADH-like distal (adult) and proximal (larval) transcripts that are shorter than transcripts in wild-type flies by a lesion located in the second ADH protein-coding exon. Direct sequencing of the amplified DNA region showed that Adhfn46 contains a 69-bp in-frame deletion that removes 23 amino acids near one border of the second exon. Consistent with these findings, we observed a shorter ADHfn46 protein present at only 3% of wild-type levels. In contrast, Adhfn45 adults and larvae accumulate much smaller amounts of ADH-like distal and proximal transcripts. Both RNAs have an identical aberration in RNA splicing of the 65-base intron sequence. Direct sequencing of the amplified mutated DNA region showed that Adhfn45 contains a 21-bp deletion that removed and rearranged DNA at the 5' splice junction of the 65-bp intron. No ADH cross-reacting material is detected in Adhfn45 flies. Direct-repeat sequences (3-11 bp) are present flanking and within the mutated DNA regions. The patterns of DNA deletion and deletion accompanied by sequence addition at the mutant sites suggest a slipped mispairing mechanism during DNA replication or repair that involves local DNA homology.

Alcohol Dehydrogenase

Analysis of Pteridium ribosomal RNA sequences by rapid direct sequencing.

A total of 864 bases from 5 regions interspersed in the 18S and 26S rRNA molecules from various clones of Pteridium covering the general geographical distribution of the genus was analysed using a rapid rRNA sequencing technique. No base difference has been detected amongst the three major lineages, two of which apparently separated before the breakup of the ancient supercontinent, Pangaea. These regions of the rRNA sequences have thus been conserved for at least 160 million years and are here compared with other eukaryotic, especially plant rRNAs.

Biological Evolution

Direct nucleotide sequencing using total cellular RNA from dengue-virus-infected mosquito cells.

In recent years, a large amount of nucleotide sequence data for dengue viruses has been published. Most of it was derived by sequencing cDNA synthesized from highly purified genomic viral RNA. This paper presents a simple and rapid method for the isolation of total RNA from mosquito cells infected with dengue viruses. This RNA can be used for direct nucleotide sequencing with specific primers without the need for further purification.

Animals

RNA sequence containing hexanucleotide AAUAAA directs efficient mRNA polyadenylation in vitro.

To determine whether a specific nucleotide sequence is required to direct polyadenylation of a simian virus 40 early pre-mRNA in a soluble HeLa whole-cell lysate, we constructed a series of rearranged and deleted DNA templates, transcribed them in vitro, and determined whether the resultant RNAs could be polyadenylated when incubated in whole-cell lysate. When a 237-base-pair DNA fragment encoding the 3' end of the simian virus 40 early pre-mRNA was transferred to recombinant plasmids encoding RNAs that were not substrates for polyadenylation, the resultant RNAs could now be polyadenylated efficiently. In one case, the chimeric RNA was polyadenylated even more efficiently than was the original simian virus 40 early transcript. Analysis of the RNAs produced from the deletion mutant templates revealed that only RNAs containing at least one copy of the AAUAAA sequence situated near the 3' end and implicated in 3'-end formation and polyadenylation in vivo could be polyadenylated in vitro. Surprisingly, this sequence directed polyadenylation of pre-mRNAs not only when near the RNA 3' end, i.e., 50 nucleotides or less away, but also when the 3' end was situated over 400 nucleotides downstream. Thus, our results show that a polyadenylic acid polymerase activity in HeLa lysates can recognize a specific nucleotide sequence in pre-mRNA and then, in the absence of the nucleolytic cleavage that presumably occurs in vivo, locate the RNA 3' end and use it as a primer for polyadenylic acid synthesis.

Base Sequence

Hybridization properties of DNA sequences directing the synthesis of messenger RNA and heterogeneous nuclear RNA.

The relationship of the DNA sequences from which polyribosomal messenger RNA (mRNA) and heterogeneous nuclear RNA (NRNA) of mouse L cells are transcribed was investigated by means of hybridization kinetics and thermal denaturation of the hybrids. Hybridization was performed in formamide solutions at DNA excess. Under these conditions most of the hybridizing mRNA and NRNA react at values of D(o)t (DNA concentration multiplied by time) expected for RNA transcribed from the nonrepeated or rarely repeated fraction of the genome. However, a fraction of both mRNA and NRNA hybridize at values of D(o)t about 10,000 times lower, and therefore must be transcribed from highly redundant DNA sequences. The fraction of NRNA hybridizing to highly repeated sequences is about 1.7 times greater than the corresponding fraction of mRNA. The hybrids formed by the rapidly reacting fractions of both NRNA and mRNA melt over a narrow temperature range with a midpoint about 11 degrees C below that of native L cell DNA. This indicates that these hybrids consist of partially complementary sequences with approximately 11% mismatching of bases. Hybrids formed by the slowly reacting fraction of NRNA melt within 4 degrees -6 degrees C of native DNA, indicating very little, if any, mismatching of bases. Hybrids of the slowly reacting components of mRNA, formed under conditions of sufficiently low RNA input, have a high thermal stability, similar to that observed for hybrids of the slowly reacting NRNA component. However, when higher inputs of mRNA are used, hybrids are formed which have a strikingly lower thermal stability. This observation can be explained by assuming that there is sufficient similarity among the relatively rare DNA sequences coding for mRNA so that under hybridization conditions, in which these DNA sequences are not truly in excess, reversible hybrids exhibiting a considerable amount of mispairing are formed. The fact that a comparable phenomenon has not been observed for NRNA may mean that there is less similarity among the relatively rare DNA sequences coding for NRNA than there is among the rare sequences coding for mRNA.

Animals

The Euglena gracilis chloroplast rpoB gene. Novel gene organization and transcription of the RNA polymerase subunit operon.

The rpoB gene coding for a beta-like subunit of the chloroplast DNA-dependent RNA polymerase has been located on the chloroplast genome of Euglena gracilis distal to the rrnC ribosomal RNA operon. We have determined 5760 base-pairs of DNA sequence, including 97 bp of the 5S rRNA gene, an intergenic spacer of 1264 bp, the rpoB gene of 4249 bp, 84 bp spacer and 67 bp of the rpoC1 gene. The rpoB gene is of the same polarity as the rRNA operons. The organization of the rpoB and rpoC genes resembles the E. coli rpoB-rpoC and higher plant chloroplast rpoB-rpoC1-rpoC2 operons. The Euglena rpoB gene (1082 codons) encodes a polypeptide with a predicted molecular weight of 124,288. The rpoB gene is interrupted by seven Group III introns of 93, 95, 94, 99, 101, 110 and 99 bp respectively and a Group II intron of 309 bp. All other known rpoB genes lack introns. All the exon-exon junctions were experimentally determined by cDNA cloning and sequencing or direct primer extension RNA sequencing. Transcripts from the rpoB locus were characterized by Northern hybridization. Fully-spliced, monocistronic rpoB mRNA, as well as rpoB-rpoC1 and rpoB1-rpoC1-rpoC2 mRNAs were identified.

Amino Acid Sequence

Dogme: a nextflow pipeline for reprocessing nanopore RNA and DNA modifications.

MOTIVATION: Oxford Nanopore (ONT) sequencing allows for the direct detection of RNA and DNA modifications from unamplified nucleic acids, which is a significant advantage over other platforms. However, the rapid updates to ONT basecalling models and the evolving landscape of computational tools for modification detection bring about challenges for reproducible and standardized analyses. To address these challenges, we developed Dogme to automate basecalling, alignment, modification detection, and transcript quantification. Dogme automates the reprocessing of ONT POD5 files by integrating basecalling using Dorado, read mapping using minimap2 and subsequent analysis steps such as running modkit. The pipeline supports three major types of sequencing data-direct RNA (dRNA), complementary DNA (cDNA), and genomic DNA (gDNA). Dogme facilitates detection of diverse RNA modifications supported by Dorado such as N6-methyladenosine (m6A), 5-methylcytosine (m5C), inosine, pseudouridine, 2'-O-methylation (Nm) and DNA methylation, while concurrently quantifying full-length transcript isoforms LR-Kallisto for transcript quantification for dRNA and cDNA. RESULTS: We applied Dogme to three separate mouse C2C12 myoblast replicates using direct RNA sequencing on MinION flow cells. We detected 96&#xa0;603 m6A, 43&#xa0;476 m5C, 8829 inosine, 10&#xa0;055 pseudouridine, and 30&#xa0;320 Nm sites in three biological replicates. The pipeline produced reproducible modification profiles and transcript expression levels across replicates, demonstrating its utility for integrative long-read transcriptomic and epigenomic analyses. AVAILABILITY AND IMPLEMENTATION: Dogme is implemented in Nextflow and is freely available under the MIT license at https://github.com/mortazavilab/dogme, with documentation provided for installation and usage.

RNA

Nucleotide sequence at the 3' end of Japanese encephalitis virus genomic RNA.

Japanese encephalitis (JE) virus genomic RNAs were purified from virions. Two hundred nucleotides at the 3' end of JE virus genomic RNA were directly sequenced by using reverse transcriptase. The nucleotide sequence at the 3' end of the viral RNA was conserved among four kinds of JE virus strains. The sequence has no AU-rich region that is present at the 3' termini of alphavirus RNAs. We also compared the nucleotide sequences at the 3' ends of RNAs from three different flaviviruses and found several common sequence elements. A secondary structure at the 3' end of JE virus genomic RNA was proposed that may be common among flavivirus genomic RNAs. Such structures and other common stretches of nucleotide sequences may be related to the biological properties of flaviviruses.

Base Sequence

EIAV genomic organization: further characterization by sequencing of purified glycoproteins and cDNA.

Nucleotide sequence analyses of two different proviral clones of equine infectious anemia virus (EIAV), designated lambda 12 (K. Rushlow et al., 1986, Virology 155, 309-321) and 1369 (T. Kawakami et al., 1987, Virology 158, 300-312), indicate significant differences in the organization of two critical regions of the viral genome, i.e., in the short open reading frames in the pol-env intergenic region and in the 5'-end of the env gene. To determine the correct structure of the EIAV genome, we have performed nucleotide sequence analyses of cDNA clones produced from viral RNA and direct sequencing of purified EIAV envelope glycoproteins (gp90 and gp45). The results of the cDNA sequencing confirm the presence of two short open reading frames in the pol-env intergenic region, as reported previously for the lambda 12 clone. The protein sequencing data correlated exactly with the amino-terminal sequences of gp90 and gp45 deduced from lambda 12 nucleotide sequences. However, the protein sequencing also revealed that the putative signal sequence of EIAV gp90 is not removed during processing. Thus, EIAV apparently contains short open reading frames analogous to human immunodeficiency virus, but differs in its mode of env polyprotein processing.

Amino Acid Sequence

Phased adenine tracts in double-stranded RNA do not induce sequence-directed bending.

Tracts of four to six adenines phased with the DNA helix produce a sequence-directed bending of the helix axis. Here, using gel electrophoresis and electron microscopy (EM), we have asked whether a similar motif will induce bending in a duplex RNA helix. Single-stranded RNAs were transcribed either from short synthetic DNA templates or from Crithidia fasciculata kinetoplast bent DNA, and the complementary single-stranded RNAs were annealed to produce duplex RNA molecules containing blocks of four to six adenines. Electrophoresis on polyacrylamide gels revealed no retardation of the RNAs containing phased blocks of adenines relative to duplex RNAs lacking such blocks. Examination by EM showed most of the molecules to be straight or only slightly bent. Thus, in contrast to DNA duplexes, phased adenine tracts do not induce sequence-directed bending in double-stranded RNA. Analysis of the distribution of molecule shapes for the highly bent C. fasciculata DNA showed that the adenine blocks do not act cooperatively to induce DNA bending and that the molecules must equilibrate between a spectrum of bent shapes.

Animals

Apolipoprotein B mRNA editing is an intranuclear event that occurs posttranscriptionally coincident with splicing and polyadenylation.

The subcellular compartment in which apolipoprotein (apo) B mRNA is edited is unknown. We studied the site of endogenous apoB mRNA editing and correlated the extent of editing with mRNA maturation in the rat liver. RNA editing activity was demonstrated in both nuclear and cytoplasmic extracts. The specific activity of the editing activity was 5.5-fold higher in the nuclear extract, which was not accounted for by activators, inhibitors, or modulators. However, the total editing activity was 3.1 times higher in the cytoplasmic extract. Highly purified rat liver nuclear apoB mRNA contained 17.3 +/- 1.45% edited sequences compared with 56 +/- 2.5% and 62.15 +/- 6.2% edited sequences in hepatic total and polysomal RNAs, respectively. Because of the significant extent of editing of total nuclear RNA, we fractionated it into a poly(A-) and poly(A+) fraction. While the poly(A-) nuclear fraction contained only 10.4 +/- 1.1% edited sequences, which represents a maximum estimate, the poly(A+) nuclear apoB mRNA contained 50 +/- 1.8% edited sequences, a value very similar to that for polysomal RNA. By direct sequencing of cDNA and genomic clones, we found that as in the case of the human apoB gene, the rat apoB gene contains an intron 25 immediately upstream of the edited exon 26. Using this information, we developed a method to examine in a highly selective manner apoB mRNA that is present in the nucleus before splicing of intron 25 and after splicing of this intron. The unspliced nuclear pre-mRNA contained 7.4 +/- 0.2% edited sequences compared with 51.0 +/- 0.9% edited sequences in the spliced nuclear apoB mRNA. Furthermore, in the poly(A-) pool of apoB pre-mRNA, unspliced nuclear pre-mRNA contained hardly any (1.56%) edited sequences, and the spliced nuclear pre-mRNA contained 7.8 +/- 0.6% edited mRNA. In the poly(A+) fraction, unspliced nuclear pre-mRNA had 25.4 +/- 0.05% and spliced nuclear mRNA 53 +/- 0.6% of its apoB mRNA in an edited form. We conclude that in the rat liver apoB mRNA editing is not a cotransciptional event. It occurs posttranscriptionally, but the process is essentially complete in the spliced polyadenylated apoB mRNA before it leaves the nucleus. Little, if any, additional editing occurs in the cytoplasmic compartment.

Animals

Sequence organization and RNA structural motifs directing the mouse primary rRNA-processing event.

The first processing step in the maturation of mouse precursor rRNA involves cleavage at nucleotide ca. +650, at the 5' border of a 200-nucleotide region that is conserved across mammals and contains the sequences that direct the processing. To identify the relevant sequence elements, we used rRNAs with small internal mutations and short pre-rRNA substrates. Much of the region can be mutated without appreciable effect, but nucleotides +655 to +666 appear to be absolutely required and short segments surrounding +750 and +810 markedly stimulate processing. The minimal processing signal corresponds to rRNA nucleotides +645 to +672. Formation of a ribonucleoprotein complex of retarded electrophoretic mobility is evidently necessary but not sufficient for processing. Computer-assisted analysis suggested a phylogenetic- and mutant-supported secondary structure in which the minimal processing signal forms a stem with the +655 region in the loop, and there is a separate branched duplex containing the downstream stimulatory sequences. Use of antisense RNA, in trans and in cis, to sequester the +655 region in a duplex supported the hypothesis that this critical region was needed in a single-stranded conformation for processing and for specific complex formation.

Animals

Vegetal messenger RNA localization directed by a 340-nt RNA sequence element in Xenopus oocytes.

Contained within a single cell, the fertilized egg, is information that will ultimately specify the entire organism. During early embryonic cleavages, cells acquire distinct fates and their differences in developmental potential might be explained by localization of informational molecules in the egg. The mechanisms by which Vg1 RNA, a maternal mRNA, is translocated to the vegetal pole of Xenopus oocytes may indicate how developmental signals are localized. Data presented here show that a 340-nucleotide localization signal present in the 3' untranslated region of Vg1 RNA is sufficient to direct RNA localization to the vegetal pole.

Animals

The genome structure of turnip crinkle virus.

The nucleotide sequence of turnip crinkle virus (TCV) genomic RNA has been determined from cDNA clones representing most of the genome. Segments were confirmed using dideoxynucleotide sequencing directly from viral RNA, and the 3' terminal sequence was confirmed by chemical sequencing of end-labeled genomic RNA. Three open reading frames (ORFs) have been identified by examination of the deduced amino acid sequences and by comparison with the ORFs found in the genome of carnation mottle virus. ORF 1 initiates near the 5' terminus of the genome and is punctuated by an amber termination codon. Translation of ORF 1 would yield a 28-kDa protein and an 88-kDa read-through product. The read-through domain possesses amino acid sequence similarities with putative viral RNA polymerases. ORFs 2 and 3 encode products of 38 (coat protein) and 8 kDa, respectively, which are expressed from subgenomic mRNAs. The organization of the TCV genome suggests that TCV is closely related to carnation mottle virus and distinct from members classified in other small RNA virus groups, such as the tombus- and sobemoviruses.

Amino Acid Sequence

Glucagon gene 3'-flanking sequences direct formation of proglucagon messenger RNA 3'-ends in islet and nonislet cells lines.

Glucagon and the glucagon-like peptides are encoded within a larger precursor, proglucagon. Transcription of the glucagon gene in pancreas, intestine, and brain gives rise to identical proglucagon mRNA transcripts, after which tissue-specific post-translational processing produces different profiles of proglucagon-derived peptides in each tissue. The importance of glucagon gene 3'-untranslated and 3'-flanking sequences in the control of glucagon mRNA production was studied by transfecting a series of 3'-deleted glucagon genes into fibroblast and islet cell lines. Glucagon genes containing 2 kilobases of 3'-flanking sequences gave rise to accurately processed mRNA transcripts in both baby hamster kidney fibroblasts and InR1-G9 islet cell lines. Deletion of all but 50 basepairs of 3'-flanking sequence had no effect on glucagon mRNA 3'-end formation. In contrast, additional deletion of 3'-flanking and 3'-untranslated sequences resulted in the production of read-through mRNA transcripts with aberrant 3'-ends. The results of these studies define a 50-basepair region in the 3'-flanking sequence of the glucagon gene important for the accurate processing of proglucagon mRNA transcripts.

Animals

Analysis of Leishbuviridae from Trypanosomatids.

Over the last decade, considerable progress has been made in unraveling RNA virus diversity. This has contributed to our understanding of the evolution of these viruses, which include emerging zoonotic human pathogens. Current success has been greatly facilitated by the development of next-generation sequencing platforms instrumental for meta-transcriptomic studies. However, due to the rapid evolution of RNA viruses, there are numerous "blind spots" waiting to be explored; one of those is the RNA virome of unicellular eukaryotes. Here, we present the pipeline, which has been successfully used to characterize various types of RNA viruses, including Leishbuviridae (Bunyaviricetes,&#xa0;Hareavirales) in the parasitic flagellates of the family Trypanosomatidae. The pipeline relies on axenic in vitro cell culture and double-stranded RNA enrichment, followed by direct RNA-sequencing. A detailed procedure description starting from the initial total RNA preparation to the final assembly of the viral segments is provided.

High-Throughput Nucleotide Sequencing

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals