PubMed HealthSearch

SEARCH · PubMed Health

Results for “Workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals

Tractor Workflow Pipeline: A Scalable Nextflow Framework for Local Ancestry-Aware Genome-Wide Association Studies.

The routine exclusion of admixed individuals from traditional Genome-Wide Association Studies (GWAS) due to concerns about spurious associations has hindered genetic analyses involving multiple ancestries. Tractor GWAS addresses this issue by incorporating local ancestry into its analysis, empowering identification of ancestry-enriched hits and generating ancestry-specific summary statistics. However, Tractor requires accurate genomic phasing and local ancestry inference as prerequisite steps, which requires additional bioinformatics expertise and decision points regarding reference panel setup. To streamline, harmonize, and automate this process, we present a scalable Nextflow workflow that integrates all necessary steps, minimizing the need for manual intervention while remaining modular and customizable. The workflow supports multiple commonly used tools and offers flexibility in how Tractor is implemented. To demonstrate its utility, we applied this pipeline to analyze 32 blood biomarkers in 6,245 two-way AFR-EUR admixed individuals from the UK Biobank. This pipeline ran efficiently at scale, replicated known associations, and identified novel ancestry-specific loci. These novel associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously missed genetic signals. By enabling the efficient analysis of admixed individuals, our workflow facilitates Tractor use, paving the way for more broader genetic discovery.

Journal Article

Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: a multicentre clinical study of two end-to-end commercial workflows.

BACKGROUND: Drug-resistant tuberculosis remains a major obstacle in ending the global tuberculosis epidemic. Deployment of molecular tools for comprehensive drug resistance profiling is imperative for successful detection and characterisation of tuberculosis drug resistance. We aimed to assess the diagnostic accuracy of a new class of molecular diagnostics for drug-resistant tuberculosis. METHODS: We conducted a prospective, cross-sectional, multicentre clinical evaluation of the performance of two targeted next-generation sequencing (tNGS) assays for drug-resistant tuberculosis at reference laboratories in three countries (Georgia, India, and South Africa) to assess diagnostic accuracy and index test failure rates. Eligible participants were aged 18 years or older, with molecularly confirmed pulmonary tuberculosis, and at risk for rifampicin-resistant tuberculosis. Sensitivity and specificity for both tNGS index tests (GenoScreen Deeplex Myc-TB and Oxford Nanopore Technologies [ONT] Tuberculosis Drug Resistance Test) were calculated for rifampicin, isoniazid, fluoroquinolones (moxifloxacin, levofloxacin), second line-injectables (amikacin, kanamycin, capreomycin), pyrazinamide, bedaquiline, linezolid, clofazimine, ethambutol, and streptomycin against a composite reference standard of phenotypic drug susceptibility testing and whole-genome sequencing. FINDINGS: Between April 1, 2021, and June 30, 2022, 832 individuals were invited to participate in the study, of whom 720 were included in the final analysis (212, 376, and 132 participants in Georgia, India, and South Africa, respectively). Of 720 clinical sediment samples evaluated, 658 (91%) and 684 (95%) produced complete or partial results on the GenoScreen and ONT tNGS workflows, respectively, with 593 (96%) and 603 (98%) of 616 smear-positive samples producing tNGS sequence data. Both workflows had sensitivities and specificities of more than 95% for rifampicin and isoniazid, and high accuracy for fluoroquinolones (sensitivity approximately ≥94%) and second line-injectables (sensitivity 80%) compared with the composite reference standard. Importantly, these assays also detected mutations associated with resistance to critical new and repurposed drugs (bedaquiline, linezolid) not currently detectable by any other WHO-recommended rapid diagnostics on the market. We note that the current format of assays have low sensitivity (≤50%) for linezolid and more work on mutations associated with drug resistance is needed. INTERPRETATION: This multicentre evaluation demonstrates that culture-free tNGS can provide accurate sequencing results for detection and characterisation of drug resistance from Mycobacterium tuberculosis clinical sediment samples for timely, comprehensive profiling of drug-resistant tuberculosis. FUNDING: Unitaid.

Humans

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics

sedimix: a workflow for the analysis of hominin nuclear DNA sequences from sediments.

SUMMARY: Sediment DNA-the recovery of genetic material from archaeological sediments-is an exciting new frontier in ancient DNA research, offering the potential to study individuals at a given archaeological site without destructive sampling. In recent years, several studies have demonstrated the promise of this approach by extracting hominin DNA from prehistoric sediments, including those dating back to the Middle or Late Pleistocene. However, a lack of open-source workflows for analysis of hominin sediment DNA samples poses a challenge for data processing and reproducibility of findings across studies. Here, we introduce a snakemake workflow, sedimix, for processing genomic sequences from archaeological sediment DNA samples to identify hominin sequences and generate relevant summary statistics to assess the reliability of the pipeline. By performing simulations and comparing our results to two published studies with human DNA from ∼25,000 years ago (including shotgun data from a sediment sample and capture data from touch DNA recovered from a deer tooth pendant) we demonstrate that sedimix yields accurate and reliable inferences. sedimix offers a reliable and adaptable framework to aid in the analysis of sediment DNA datasets and improve reproducibility across studies. AVAILABILITY AND IMPLEMENTATION: sedimix is available as an open-source software with the associated code, example data, and user manual with installation instructions available at https://github.com/jierui-cell/sedimix. A permanent archived version of this release is available via Zenodo: https://doi.org/10.5281/zenodo.17244854.

Animals

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics

LncRAnalyzer: a robust workflow for long non-coding RNA discovery using RNA-Seq.

Long non-coding RNA (lncRNA) is a major transcript category that lacks protein-coding capabilities, with relatively low abundance and complex expression patterns. Distinguishing lncRNAs from protein-coding genes is a complex process involving multiple filtering steps. We developed an automated pipeline named LncRAnalyzer featuring retrained models for 60 species. This workflow aims to reduce the likelihood of obtaining protein-coding or partial protein-coding transcripts during lncRNA identification by utilizing eight distinct approaches. We conducted a 10-fold cross-validation of the sorghum models and training sets with their standard ones and other approaches using real-life RNA-Seq datasets and known lncRNA and CDS sequences of sorghum. The results showed that the sorghum models and training sets were outperformed. The pipeline output comprises upset plots illustrating the number of lncRNA/NPCTs identified by the approaches, commonly identified lncRNA and their classes, NPCTs, and expression count tables. A feature-level comparison and benchmarking analysis of LncRAnalyzer with four existing pipelines, namely, LncPipe, LncEvo, lncRNA-Annotation, and Plant-LncPipe, demonstrated that LncRAnalyzer is more comprehensive, easier to implement, and accurate in lncRNA predictions. This workflow also ascertains lncRNA origins from various Transposable Elements (TEs) in plants using TE annotations from APTEdb [http://apte.cp.utfpr.edu.br/]. LncRAnalyzer is publicly available on GitLab [https://gitlab.com/nikhilshinde0909/LncRAnalyzer.git] for academic users.

RNA, Long Noncoding

Development of a clinical metagenomics workflow for the diagnosis of wound infections.

BACKGROUND: Wound infections are a common complication of injuries negatively impacting the patient's recovery, causing tissue damage, delaying wound healing, and possibly leading to the spread of the infection beyond the wound site. The current gold-standard diagnostic methods based on microbiological testing are not optimal for use in austere medical treatment facilities due to the need for large equipment and the turnaround time. Clinical metagenomics (CMg) has the potential to provide an alternative to current diagnostic tests enabling rapid, untargeted identification of the causative pathogen and the provision of additional clinically relevant information using equipment with a reduced logistical and operative burden. METHODS: This study presents the development and demonstration of a CMg workflow for wound swab samples. This workflow was applied to samples prospectively collected from patients with a suspected wound infection and the results were compared to routine microbiology and real-time quantitative polymerase chain reaction (qPCR). RESULTS: Wound swab samples were prepared for nanopore-based DNA sequencing in approximately 4 h and achieved sensitivity and specificity values of 83.82% and 66.64% respectively, when compared to routine microbiology testing and species-specific qPCR. CMg also enabled the provision of additional information including the identification of fungal species, anaerobic bacteria, antimicrobial resistance (AMR) genes and microbial species diversity. CONCLUSIONS: This study demonstrates that CMg has the potential to provide an alternative diagnostic method for wound infections suitable for use in austere medical treatment facilities. Future optimisation should focus on increased method automation and an improved understanding of the interpretation of CMg outputs, including robust reporting thresholds to confirm the presence of pathogen species and AMR gene identifications.

Humans

REAPER: a project-centric workflow layer for comparative repeatome analysis.

INTRODUCTION: Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments - a gap that motivates a project-centric workflow layer for repeatome analysis. METHODS: We present REAPER (Repeatome Extended Analysis Pipeline-Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. RESULTS: Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers - including pTa794- and pSc119-class entries in curated databases. DISCUSSION: These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit - and by supporting caching and structured logging - REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.

TAREAN

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

A Computational Workflow for Prioritizing Microbial Metabolite-Associated Host Genes in Constipation-Predominant Irritable Bowel Syndrome.

No standardized computational pipeline exists for systematically prioritizing microbial metabolite-associated host genes and protein-ligand complexes from publicly available chemical, genomic, and structural databases. This article describes an eight-stage workflow that accepts a user-defined set of gut microbiota-derived metabolites and produces a ranked shortlist of candidate metabolite-associated host genes, enriched biological pathways, and structurally prioritized protein-ligand complexes for experimental follow-up. The pipeline integrates (i) chemoinformatic metabolite profiling; (ii) multi-database candidate target prediction using protein-chemical interaction and ligand-based target-prediction tool and a molecular docking program; (iii) differential gene expression analysis of publicly available transcriptomic data; (iv) target-differentially expressed gene overlap; (v) protein-protein interaction network construction and pathway enrichment; (vi) molecular docking with a molecular docking program; (vii) 200 ns molecular dynamics simulation using a molecular dynamics engine with a protein force field used for molecular dynamics simulations; and (viii) MM-PBSA binding free-energy estimation. As a worked example, nine gut microbiota-derived or microbiota-modified metabolites representing short-chain fatty acids, bile acids, tryptophan-derived metabolites, and urolithin A were processed using the public IBS-C rectal mucosal transcriptomic dataset GSE36701. The workflow ranked 17 unique predicted metabolite-associated genes that were differentially expressed in this dataset. Docking, molecular dynamics simulation, and MM-PBSA analyses structurally prioritized five metabolite-protein complexes: lithocholic acid-VDR, lithocholic acid-NR1H4/FXR, ursodeoxycholic acid-NR1H4/FXR, tryptamine-HTR2A (simulated in an explicit 1-Palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC) lipid bilayer), and urolithin A-CASP3. The protocol is designed to be adaptable to other metabolite sets, disease transcriptomic datasets, and target classes; all outputs are hypothesis-generating computational predictions that require independent transcriptomic replication, protein-level validation, and functional ligand-response assays before causal or therapeutic conclusions can be drawn.

Irritable Bowel Syndrome

Surgical management of jugular foramen meningiomas: a function-prioritized perioperative workflow.

OBJECTIVE: Jugular foramen meningiomas are challenging because of their deep, neurovascularly crowded location and multicompartment extension; hyperostosis and rigid dural attachment further narrow the corridor and increase the risk of lower cranial nerve morbidity, causing dysphagia and airway complications that may rarely require tracheostomy. This study aimed to describe a contemporary function-first workflow integrating compartment-based anatomy, venous sinus status, preoperative embolization, and continuous vagus nerve monitoring and its relation to clinically actionable recovery endpoints. METHODS: The authors retrospectively reviewed 26 consecutive patients who underwent primary surgery for jugular foramen meningiomas (2014-2025). Tumors were classified as intradural + intrajugular (IJ) or intradural + intrajugular + extracranial extension (IJE). Retrosigmoid, suprajugular, or transjugular approaches were selected by tumor extension and sigmoid-jugular venous status. Selective embolization and continuous vagus nerve monitoring were used when feasible. Outcomes included extubation timing, time to oral intake, 1-year swallowing/voice severity, extent of resection, and salvage stereotactic radiosurgery (SRS) for progression/regrowth. RESULTS: Twenty tumors were IJ and 6 were IJE. Selective embolization was performed in 16 patients (62%) without complications. Continuous vagus nerve monitoring was implemented in 16 patients (62%); lower preservation rates showed an exploratory association with worse 1-year swallowing. All patients were extubated immediately after surgery. Oral intake began by postoperative day ≤ 7 in 20 patients (77%); only 1 required > 14 days before resuming oral intake. At 1 year, swallowing and hoarseness remained worse in 54% and 46% of patients, respectively, but almost all cases were mild; the same patient had moderate dysphagia/hoarseness, and none required tracheostomy, gastrostomy, long-term tube feeding, or phonosurgery. Simpson grade IV comprised 69% of cases but predominantly reflected intrajugular/extracranial residual rather than persistent intradural disease. No patient without preoperative facial nerve palsy developed new palsy; serviceable hearing was preserved in 70%, and 38% with preoperative nonserviceable hearing improved to serviceable hearing. During a median 55.6-month follow-up, 3 patients (12%) underwent salvage SRS for regrowth; none required reoperation. CONCLUSIONS: A function-first workflow guided by anatomical compartment extension and intraoperative monitoring can support rapid recovery and durable functional independence in jugular foramen meningiomas. The IJE phenotype identifies a higher-risk subgroup for delayed oral intake and postoperative subjective dysphagia/hoarseness, while continuous vagus nerve monitoring may provide actionable insights to calibrate surgical aggressiveness and support function-prioritized acceptance of intrajugular/extracranial residual with close surveillance and salvage SRS when needed.

Humans

Hard to Halt: Automation Bias in Agent-Driven Sequencing Prior Authorization Workflows.

PURPOSE: Prior authorization (PA) for exome or genome sequencing is a time-consuming process that impedes timely rare disease diagnosis. Large language model-based browser agents offer potential for automating these workflows, but their clinical reliability remain uncharacterized. METHODS: We developed a sandbox compromising a simulated ES/GS PA submission payer portal and a synthetic EHR containing 836 patient records spanning compliant profiles and deficient profiles with different types of issues. Gemini 3 Pro, Gemini 3 Flash, and Claude Opus 4.5 were evaluated on task completion rate, form completion accuracy, and appropriate withholding for deficient profiles. RESULTS: Larger models achieved much higher task completion rates (Gemini 3 Pro 95.45%, Claude Opus 4.5 93.67%) compared to Gemini 3 Flash (56.05%), but nearly universally failed to withhold submission for deficient profiles whereas Gemini 3 Flash ironically demonstrated superior withholding performance (17.33%). In a non-agentic setting, Gemini 3 Pro correctly identified 91% of the issues in deficient profiles, indicating that withholding failure is attributable to the browser interaction rather than the model's reasoning limitations. CONCLUSION: Current LLM-based browser agents exhibit a systematic bias towards form submission that poses risks in PA workflows. A modular, multi-agent architecture with human supervision is necessary for a safe clinical deployment.

Journal Article

Workflow for Long-Read Amplicon Sequencing of Chikungunya Virus Using Oxford Nanopore Technology.

This protocol provides a comprehensive, step-by-step workflow for whole-genome sequencing of Chikungunya virus (CHIKV) using an amplicon-based strategy optimized for Oxford Nanopore Technologies (ONT) platforms. The procedure includes detailed instructions for sample handling, viral RNA extraction, quality control, cDNA synthesis, multiplex PCR amplification, library preparation, sequencing, and primary bioinformatic processing. The protocol is designed to maximize reproducibility across laboratories and is suitable for genomic surveillance applications, including outbreak investigation and molecular epidemiology, even when working with low-to-moderate viral loads.

Chikungunya virus

ModiCal: A Targeted Calibration Workflow for Site-Specific m5C Validation by Nanopore Direct RNA Sequencing.

Accurate identification of RNA 5-methylcytidine (m5C) at the single-nucleotide resolution remains a central challenge in nanopore direct RNA sequencing (DRS). Current global scanning and modification-aware basecalling methods enable transcriptome-wide profiling but often yield high false-positive rates and lack site-specific accuracy. To address this, we repurposed ModiDeC, originally a de novo multimodification classifier, into a targeted, high-precision validation tool for RNA modification sites with prior biochemical knowledge. This was implemented through a three-step calibration workflow that alternates between biochemical and computational modules using the well-characterized m5C2278 site in 25S rRNA as a starting point. Baseline training uses short synthetic RNAs carrying either a methylated or unmodified C2278 as ground truth, followed by IVT-derived calibration and validation in methyltransferase knockout yeast. The baseline model accurately detected the bona fide m5C2278 site but initially produced off-target predictions. Iterative retraining with unmodified IVT signals progressively reduced and ultimately eliminated false positives while maintaining a strong signal at the bona fide site. The final model retained enzyme-dependent detection in wild-type versus knockout yeast and, when explicitly targeted, was also able to detect the second rRNA site, C2870, which remained invisible in the initial analysis. Application to native human prerRNA processing intermediates further resolved two distinct m5C deposition regimes on 28S rRNA, while generalization to dengue virus genomic RNA confirmed that the same calibration logic transfers across diverse RNA contexts. Together, this study establishes a reproducible and transferable framework that integrates biochemical validation with iterative neural network refinement, providing a route toward reliable site-specific m5C confirmation by nanopore direct RNA sequencing.

RNA Methylation

A scalable, low-cost, sample hashing workflow for multiomic single-cell analysis using the Seq-Well S3 platform.

In-depth analyses of clinical samples have the potential to provide unparalleled insights into the cellular mechanisms that underlie both health and disease, as well as therapeutic and prophylactic responses. However, these specimens are often paucicellular, necessitating the use of workflows that maximize the amount of information that can be learned. Here we provide a detailed protocol for generating and analyzing single-cell multiomic data from low-input samples with the Seq-Well S3 platform. We further describe a matched pipeline for sample hashing that reduces costs and sources of technical variation in the resulting data while also enhancing throughput. In brief, our streamlined and efficient methodology involves: (1) optionally staining single-cell suspensions with antibody-oligonucleotide conjugates for cell surface protein quantification and/or sample multiplexing; (2) generating Seq-Well S3 sequencing libraries; (3) optionally producing bulk-RNA sequencing libraries via SMART-seq2 to support genetic demultiplexing; and (4) computationally analyzing the resulting data. Each step herein has been designed to leverage readily available reagents and standard laboratory equipment, substantially lowering barriers to entry for researchers. The overall Protocol can yield high-quality multiomic insights from samples in under a week.

Single-Cell Analysis

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software