PubMed HealthSearch

SEARCH · PubMed Health

Results for “Library preparation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 × 150 bp and 2 × 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 × 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 × 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 × 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 × 150 bp and 2 × 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 × 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

High-throughput DNA extraction and cost-effective miniaturized metagenome and amplicon library preparation of soil samples for DNA sequencing.

Reductions in sequencing costs have enabled widespread use of shotgun metagenomics and amplicon sequencing, which have drastically improved our understanding of the microbial world. However, large sequencing projects are now hampered by the cost of library preparation and low sample throughput, comparatively to the actual sequencing costs. Here, we benchmarked three high-throughput DNA extraction methods: ZymoBIOMICS™ 96 MagBead DNA Kit, MP BiomedicalsTM FastDNATM-96 Soil Microbe DNA Kit, and DNeasy® 96 PowerSoil® Pro QIAcube® HT Kit. The DNA extractions were evaluated based on length, quality, quantity, and the observed microbial community across five diverse soil types. DNA extraction of all soil types was successful for all kits, however DNeasy® 96 PowerSoil® Pro QIAcube® HT Kit excelled across all performance parameters. We further used the nanoliter dispensing system I.DOT One to miniaturize Illumina amplicon and metagenomic library preparation volumes by a factor of 5 and 10, respectively, with no significant impact on the observed microbial communities. With these protocols, DNA extraction, metagenomic, or amplicon library preparation for one 96-well plate are approx. 3, 5, and 6 hours, respectively. Furthermore, the miniaturization of amplicon and metagenome library preparation reduces the chemical and plastic costs from 5.0 to 3.6 and 59 to 7.3 USD pr. sample. This enhanced efficiency and cost-effectiveness will enable researchers to undertake studies with greater sample sizes and diversity, thereby providing a richer, more detailed view of microbial communities and their dynamics.

Metagenome

Library strategies differentially shape microbial, functional, and host signals in clinical metagenomic sequencing.

Metagenomic next-generation sequencing (mNGS) is increasingly used in infectious disease diagnostics, yet how library preparation shapes the microbial, functional, and host signals recovered from clinical samples remains poorly defined. Here, we performed a within-sample parallel comparison of three mNGS library preparation strategies-DNA-based libraries (DNAlib), RNA-based libraries (RNAlib), and total nucleic acid-based libraries (TNAlib)-across a diverse range of clinical specimens spanning five sample types. Using a curated clinical infectome as a benchmark, we show that library strategies are not interchangeable but capture distinct biological dimensions of the same specimen. RNAlib provided the most comprehensive standalone recovery of the clinical infectome, with improved detection of RNA viruses and cellular pathogens, enhanced resolution of resistance and virulence signals, and preservation of infection-associated host immune signatures. DNAlib showed stronger baseline recovery of DNA viruses and broader host genome coverage, whereas the TNAlib workflow evaluated here largely behaved as an intermediate strategy rather than a consistent improvement over dedicated DNA- or RNA-based workflows. Together, these results establish that the library preparation protocol is a major determinant of how clinical mNGS data should be interpreted and provide a framework for selecting sequencing strategies according to specific diagnostic and biological questions.IMPORTANCEMetagenomic sequencing is increasingly used in infectious disease research and clinical diagnostics, but different library preparation strategies may recover fundamentally different biological signals from the same sample. These signals include not only pathogens but also background microbes, microbial functional activity, and host immune-response patterns. Here, we systematically compared DNA-, RNA-, and total nucleic acid-based metagenomic sequencing libraries using the same clinical samples processed in parallel. We found that the three strategies did not provide equivalent information. RNA-based sequencing generated the most informative single-library view of infection, particularly for RNA viruses, cellular pathogens, functional microbial signals, and host immune-response patterns. DNA-based sequencing was more effective for DNA virus and host genome recovery, whereas the total nucleic acid sequencing workflow evaluated here generally behaved as an intermediate strategy. These findings show that library preparation can substantially influence the interpretation of metagenomic data.

functional characterization

RNA Sequencing Protocols for Short-Read Sequencing.

RNA sequencing (RNA-seq) methodologies allow the discovery of novel variants and transcripts. These comprise three general steps: (1) capture of RNA species of interest, (2) conversion of RNA to complementary DNA (cDNA), and (3) modification of cDNA to fit the sequencing platform. Here we describe four different library preparation protocols for short-read sequencing: cDNA synthesis with poly(A) selection, library preparation with ribosomal depletion, and cDNA synthesis with SMART® (Switching Mechanism at 5' end of RNA Template) technology for low and Pico inputs.

Gene Library

SMART-RNA-Metavirome: a practical RNA metavirome platform compatible with high-throughput sequencing of both short and long reads.

BACKGROUND: The RNA virosphere's extensive diversity and its role in emerging infectious diseases underscore the importance of non-targeted sequencing for identifying unknown or rare pathogens, including co-infections. However, enriching low-abundance viral sequences in RNA metaviromics, particularly in the preparation of cDNA libraries and their compatibility with next-generation sequencing (NGS) and third-generation sequencing (TGS), remains challenging. Therefore, our objective is to develop and systematically assess a practical RNA metavirome methodology specifically tailored for the enrichment of low-abundance viral sequences within samples. METHODS: We developed the SMART-RNA-Metavirome platform, integrating SMART-9n library preparation with NGS and TGS technologies. Total RNA was extracted from two field-collected wild Aedes albopictus pools, along with one laboratory-infected Ae. albopictus pool harboring dengue virus (DENV). This RNA was subjected to reverse transcription using both this optimized protocol and random primer-based methods, followed by high-throughput sequencing on Illumina, Oxford Nanopore, and QitanTech Nanopore technologies. Welch's t-test was employed for comparative analysis of the subsequent RNA metavirome data, specifically to evaluate differences in viral species composition and abundance of viral reads between experimental groups. Furthermore, the effectiveness of this platform was systematically validated via RT-qPCR and SMART-RNA-Metavirome-based Oxford Nanopore sequencing across multiple sample types, including mosquito specimens from DENV-infected Ae. albopictus, serum samples from dengue patients and viral isolates of Japanese encephalitis virus (JEV) and Zika virus (ZIKV). RESULTS: The SMART-RNA-Metavirome platform has been systematically validated to excel in enriching the composition and diversity of the RNA virome (P = 0.04), providing sufficient coverage for the complete reconstruction of viral genomes. When employed in the detection of DENV-infected Ae. albopictus, clinical serum samples, and viral isolates of JEV and ZIKV, this technique exhibits a robust correlation with RT-qPCR (r2 > 0.95). Notably, it demonstrates exceptional sensitivity, ensuring sufficient coverage even in samples of DENV-infected Ae. albopictus with a Ct-value of 35.3, attaining an impressive 99.88% genome coverage. Furthermore, this platform possesses the capability to identify virus species and determine their serotypes. CONCLUSIONS: In our study, the SMART-RNA-Metavirome platform outperforms traditional methods, enriching RNA virome composition and diversity, enabling practical compatibility with both NGS and TGS technologies. It demonstrates significant proficiency in detecting both known and unknown arboviruses, even in low-titer samples such as those from wild mosquitoes and clinical sera. This platform facilitates comprehensive monitoring, risk assessment, and early warning of RNA virus transmissions, enhancing our understanding of RNA virome diversity and ecological patterns.

High-Throughput Nucleotide Sequencing

Protocol for Duplex Sequencing of Mitochondrial DNA in Single Human Oocytes.

Oocytes are densely packed with mitochondria, the energy-producing organelles that contain their own genome, mitochondrial DNA (mtDNA). Each cell contains multiple copies of mtDNA, with copy number varying among tissue types. Oocytes possess the highest mtDNA copy number, containing hundreds of thousands of mtDNA molecules per cell. Because mitochondria are inherited exclusively through the maternal lineage, accurate detection of mtDNA variants is essential for studies of inheritance, aging, and disease. The presence of multiple mtDNA copies allows wild-type and mutant molecules to coexist within the same cell, a condition known as heteroplasmy, in which low-frequency and de novo variants may occur at frequencies below 1%. Conventional next-generation sequencing (NGS) lacks sufficient accuracy to reliably distinguish these rare variants from errors introduced during library preparation and sequencing. Here, we present a protocol for enriching mtDNA from single human oocytes using Exonuclease V to remove linear DNA, followed by duplex sequencing library preparation for highly accurate mtDNA analysis. This workflow enables error-corrected sequencing of individual oocytes, facilitating reliable detection of low-frequency mtDNA variants and analysis of heteroplasmy and de novo mutagenesis. The protocol provides a reproducible approach for investigating mitochondrial genome variation in single oocytes using Illumina-compatible sequencing platforms.

Humans

URMD-Seq: A high-throughput method for scalable detection of ultra-rare mutations in the human mitochondrial genome.

The study of mitochondrial genetics has long been limited to polymorphisms and high frequency mutations owing in part to technical and technological limitations in reliably detecting and quantifying rare somatic mutations. Over the past decade or so, the study of rare somatic mitochondrial DNA (mtDNA) variants has expanded and continues to garner increasing interest in a wide range of research fields. Here, we describe Ultra-Rare Mutation Detection-Sequencing (URMD-Seq), a high-throughput method that combines unique molecular identifier (UMI)-based library preparation and Next Generation Sequencing (NGS) for the accurate and scalable detection of ultra-rare mutations in the mtDNA control region. Our method exploits degenerate primers to label individual mtDNA molecules. This is followed by several purification, quantification and amplification steps, to obtain high quality amplicons for sequencing on the Illumina MiSeq platform. Our approach enables the use of total genomic DNA extract as starting point for the assay, overcoming the need for organelle isolation and/or mtDNA enrichment, hence broadening the type of specimen that can be studied, while offering cost and time benefits. The assay described herein has been demonstrated to reliably measure variants present at on average 0.09%, but as low as 0.03%, variant allele frequency in a variety of tissues, including fresh and frozen biobanked specimens. Using this protocol, library preparation of 300 specimens can be completed by a single individual with general nucleic acid handling experience in approximately 20 days. Given its flexibility and scalability, URMD-Seq is particularly well suited for epidemiological studies using a large number of specimens.

Humans

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4&#x2013;6&#xa0;h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30&#x2009;>&#x2009;90%, mapping rates&#x2009;>&#x2009;95%, on-target rates 85&#x2013;90%). The duplication rate was reduced to 5&#x2013;8%, compared with 10&#x2013;15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

Targeted Next-Generation Sequencing in Rare Diseases.

Targeted next-generation sequencing (NGS) in rare disease focuses on genetic analysis of specific regions in genome that are linked to a rare disease. In addition to library preparation, sequencing, and data analysis, targeted NGS includes an additional step of target enrichment of selected genes and regions. It allows for more sensitive and profound sequencing, as it is a fast and cost-effective approach with less data burden and is therefore often a method of choice for identifying rare variants in known genes, especially in diagnostics of rare diseases. Several in silico tools address the pathogenicity predictions of rare variants of unknown significance (VUS) and can therefore facilitate clinical interpretation.

Rare Diseases

A contextual activity score (CAS) for inferring ADAR-associated transcriptional activity across RNA-seq, single-cell, and spatial transcriptomics.

BACKGROUND AND OBJECTIVE: Adenosine-to-inosine RNA editing, catalyzed by Adenosine Deaminases Acting on RNA (ADARs), is a widespread modification involved in neural function, immune regulation, and cancer. The Alu Editing Index (AEI) is the standard metric to estimate ADAR activity but requires raw sequencing reads and is poorly suited for single-cell and spatial transcriptomic data. This study aimed to develop an alternative framework for inferring ADAR-associated transcriptional activity from gene expression data across diverse transcriptomic technologies. METHODS: We developed the Contextual Activity Score (CAS), a framework based on transcriptional signatures from ADAR perturbation experiments. Context-specific signatures were generated for human neurons, mouse neurons, and cancer models to infer ADAR1 and ADAR2 activity. CAS was computed from normalized gene expression matrices using regulon-based enrichment analysis. Performance was evaluated by comparing with the Alu Editing Index across bulk RNA sequencing datasets, simulated sequencing depths, and library preparation protocols. RESULTS: CAS showed strong concordance with the Alu Editing Index across multiple datasets, while remaining robust to reduced sequencing depth and different library protocols. Unlike the Alu Editing Index, CAS can be applied to single-cell and spatial transcriptomic data and enables the independent assessment of ADAR2 activity. In cancer and neuronal contexts, CAS captured biologically meaningful variations in ADAR-associated transcriptional activity at sample, cell-type, and spatial levels. CONCLUSION: CAS provides a scalable approach applicable across multiple RNA-seq protocols for estimating ADAR-associated transcriptional activity using gene expression data. This method, implemented in an open-source R package for broad adoption, expands the ability to study ADAR-associated transcriptional activity across transcriptomic modalities where direct editing quantification is challenging, such as single-cell and spatial transcriptomics.

Adenosine Deaminase

Identification of RBP binding sites using RNA deaminases.

RNA-binding proteins (RBPs) are critical regulators of gene expression and RNA processing. Identification of their binding sites has important implications for their physiological and disease-related functions. Crosslinking and immunoprecipitation, followed by sequencing (CLIP-seq) and its derivatives, are the most commonly used methods to identify RBP binding sites, but are laborious and require a large amount of starting material. Recent advancements harnessing RNA deaminases in fusion to any RBP of interest, allow for the profiling of RBP binding sites from low-input samples in simpler procedures. Among these efforts, we developed STAMP (Surveying Targets by APOBEC-Mediated Profiling), which efficiently detects RBP-RNA interactions. This chapter describes the detailed protocol for the STAMP method, including plasmid construction, delivery and sorting, library preparation and bioinformatic data analysis.

RNA-Binding Proteins

Whole-Genome Bisulfite Sequencing with a Small Amount of DNA.

Whole-genome bisulfite sequencing (WGBS) is the most widely used method to study DNA methylation profiles across the genome. Since the bisulfite reaction causes DNA degradation, a new approach called post-bisulfite adapter tagging (PBAT) was developed to overcome this problem by adding adapters after bisulfite treatment. In mammals, the PBAT method is used for single-cell bisulfite sequencing (scBS-seq), which enables DNA methylation analysis using a very small amount of DNA from only a few cells, including single-cell input. This protocol involves bisulfite conversion, followed by preamplification and tagging with random hexamer primers prior to Illumina library preparation. Since many procedures are completed in one single test tube, the loss of DNA can be minimized, enabling highly sensitive experiments to study DNA methylation profiles from a very small amount of input material.

Sulfites

Evaluation of bone preparation approaches using length-based analysis and targeted sequencing for forensic human identification of historic skeletal remains.

Advances in DNA technology have significantly enhanced the forensic community's ability to develop genetic profiles from unidentified human skeletal remains. However, sampling requires mechanical grinding of hard tissues before DNA isolation. This processing can compromise genetic profiles, particularly in aged bones. We compared the industry-standard pulverization method with an alternative powder-free preparation involving prolonged demineralization and subsequent slicing of 19th-century cortical bone. Data from DNA quantification, STR genotyping, and targeted SNP sequencing were used to evaluate powdered samples versus demineralized slices from paired human bones. Average human DNA yields for pulverized samples and demineralized slices were 0.032&#x2009;ng and 0.692&#x2009;ng, respectively. Demineralized slices recovered more amplifiable DNA than traditional homogenization methods (p&#x2009;<&#x2009;0.05). No pulverized samples produced STR profiles, whereas demineralized slices from the same bone samples yielded partial profiles. Samples underwent DNA repair, library preparation, and hybridization capture using the FORensic Capture Enrichment (FORCE) panel. Applying low-coverage (1X) analysis of high-throughput sequencing (HTS) data, demineralized slices outperformed those prepared by traditional pulverization methods (p&#x2009;<&#x2009;0.05) and substantially increased the information recovered compared with conventional STR analysis methods. Based on HTS data from pulverized samples, DNA fragment length ranged from 27 to 95&#x2009;bp, and FORCE SNP recovery was 33.23%. In contrast, for demineralized slices, DNA fragment length ranged from 85 to 114&#x2009;bp, and FORCE SNP recovery was 83.24%. The required reagents and equipment are typically available in forensic labs, and the workflow outlined herein significantly increases the success of DNA recovery from challenging skeletal samples.

Humans

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n&#x202f;=&#x202f;53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals

Workflow for Long-Read Amplicon Sequencing of Chikungunya Virus Using Oxford Nanopore Technology.

This protocol provides a comprehensive, step-by-step workflow for whole-genome sequencing of Chikungunya virus (CHIKV) using an amplicon-based strategy optimized for Oxford Nanopore Technologies (ONT) platforms. The procedure includes detailed instructions for sample handling, viral RNA extraction, quality control, cDNA synthesis, multiplex PCR amplification, library preparation, sequencing, and primary bioinformatic processing. The protocol is designed to maximize reproducibility across laboratories and is suitable for genomic surveillance applications, including outbreak investigation and molecular epidemiology, even when working with low-to-moderate viral loads.

Chikungunya virus

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum