PubMed HealthSearch

SEARCH · PubMed Health

Results for “High throughput workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A streamlined workflow for high throughput metaproteomic analysis of the rumen microbiome.

Metaproteomics can provide direct functional insights into complex microbial communities, yet its application in rumen research remains limited due to labor-intensive and low-throughput sample preparation workflows before the MS analysis. This work aimed to develop and characterize a streamlined, high throughput metaproteomic workflow optimized for rumen samples. Key steps, including microbial cell extraction, cell lysis, protein digestion, and LC-MS/MS acquisition, were systematically assessed and optimized to reduce hands-on time while maintaining deep proteome coverage. The optimized workflow integrates a minimized cell extraction protocol using 0.5 g starting material and in-solution tryptic digestion. Application of the final workflow to 72 samples from in vitro fermentation revealed that biological variability between inocula dominated technical variability, which remained moderate (median CV of 21-24% across batches). Overall, the optimized workflow supports robust taxonomic and functional characterization of the rumen microbiome with improved scalability. These advances provide a foundation for applying metaproteomics to larger experimental designs, including nutritional trials and cohort studies, thereby enabling broader functional interrogation of rumen microbial ecosystems. SIGNIFICANCE: This study addresses current limitations in the application of metaproteomics to rumen microbiome research by developing a streamlined and scalable sample preparation workflow. By optimizing key steps and reducing sample input while maintaining reproducibility and proteome coverage, this work enables more efficient processing of larger sample sets. These advances support the broader use of metaproteomics in rumen studies and facilitate functional investigations relevant to animal nutrition and sustainable livestock production.

Animals

Herd-level heterogeneity of antimicrobial resistance in commensal Escherichia coli: A nationwide high-throughput survey of Australian pig herds.

Antimicrobial resistance in commensal Escherichia coli provides a useful indicator for overall antimicrobial resistance burden. We applied this approach to assess antimicrobial resistance within and between commercial pig herds across Australia. A high-throughput robotic workflow was used to isolate 2730 E. coli colonies from rectal contents collected in 2022 from healthy slaughter pigs (n = 300) representing 30 herds (∼70% of national production). Up to 94 isolates per herd underwent antimicrobial susceptibility testing using the Robotic Antimicrobial Susceptibility Platform. Isolate- and herd-level antimicrobial resistance indices were calculated, weighting antimicrobials by their human health importance. Resistance to first-line agents was widespread: ampicillin 77% and tetracycline 79%. By contrast, resistance to critically important antimicrobials was rare (ciprofloxacin 0.11%; extended-spectrum cephalosporins 0.04%), and no clinical resistance to carbapenems or colistin was detected. Overall, 56.9% of isolates were multi-class resistant. Herd-level antimicrobial resistance within indices ranged from 1.51 to 5.76, revealing substantial between-herd heterogeneity. Three herds carried critically important antimicrobials-resistant isolates that would likely have been missed using conventional, lower-density sampling approaches. Whole-genome sequencing identified fluoroquinolone-resistant isolates belonging to ST10 and ST69 (both qnrS1), and ST744 (Quinolone Resistance Determining Region mutations plus blaCTX-M-27). By testing approximately tenfold more isolates than conventional surveys, we uncovered considerable antimicrobial resistance with heterogeneity within and between animals and herds, including farm-specific variability. This expanded sampling also enabled detection of critically important antimicrobial resistance at very low prevalence. In conclusion, high-throughput, high-density testing offers a practical early-warning system and herd-level benchmark to inform surveillance and targeted interventions.

Animals

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5 ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91 µg of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8 million filtered reads per sample, with 63±7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9 million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans

MetaChrome: An Open-Source, User-Friendly Tool for Automated Metaphase Chromosome Analysis.

DNA Fluorescence In Situ Hybridization (FISH) is an essential technique to study chromosome biology and genetics, enabling precise visualization of specific genomic loci to study structural abnormalities, gene mapping, and chromosomal rearrangements. High-Throughput Imaging (HTI) can automate the analysis of DNA-FISH chromosome images, but the accurate and automated segmentation of mitotic chromosomes and simultaneous colocalization of FISH signals remains a challenge. While several commercial automated karyotyping tools partially solve these issues, open-source software that effectively combines robust chromosome segmentation with comprehensive colocalization analysis capabilities remains necessary. To address this unmet need, we developed MetaChrome, an open-source software platform built around a graphical user interface and explicitly designed for automated metaphase chromosome analysis. MetaChrome leverages fine-tuned deep learning models to automate metaphase chromosome segmentation, together with colocalization analysis of chromosome-specific FISH probes and immunofluorescent-labeled proteins. Importantly, MetaChrome achieves enhanced segmentation accuracy compared to traditional image processing methods by adopting a Cellpose segmentation model fine-tuned with manually annotated metaphase chromosome datasets. The fine-tuned model ensures precise assignment of DNA-FISH spots to individual chromosomes in an automated manner. This facilitates rapid identification of chromosomal abnormalities, reduces human error, and advances high-throughput chromosome analysis workflows, addressing a key bottleneck in chromosome biology research.

Chromosome segmentation

Analysis of Confounding Factors in Reactive Cysteine Profiling Reveals Enhanced Chromatin-Protein Association via CDK7 Inhibition by THZ1.

Recent advances in activity-based proteome profiling (ABPP) have enabled the global mapping of cysteine ligandability, uncovering novel biological insights and opportunities for identifying disease vulnerabilities. While both live-cell-based and native-lysate-based ABPP have been applied, how cysteine ligandability differs between these systems and what factors influence these measurements remain unclear. Building on our previous development of a high-throughput TMT-ABPP workflow for native lysates, here we adapt the protocol for live cells and systematically compare cysteine ligandability across both platforms. Our analysis reveals three major contributors to the discrepancies: in-cellular cysteine accessibility, protein abundance changes, and protein relocalization. Notably, we highlight that the CDK7 inhibitor THZ1 induces substantial protein relocalization and promotes chromatin binding. Together, these results provide a practical framework for ABPP experimental design and data interpretation, supporting the more accurate application of ABPP in functional proteomics and drug discovery.

Cysteine

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Turbo-charging crop improvement: harnessing multiplex editing for polygenic trait engineering and beyond.

Multiplex CRISPR editing has emerged as a transformative platform for plant genome engineering, enabling the simultaneous targeting of multiple genes, regulatory elements, or chromosomal regions. This approach is effective for dissecting gene family functions, addressing genetic redundancy, engineering polygenic traits, and accelerating trait stacking and de novo domestication. Its applications now extend beyond standard gene knockouts to include epigenetic and transcriptional regulation, chromosomal engineering, and transgene-free editing. These capabilities are advancing crop improvement not only in annual species but also in more complex systems such as polyploids, undomesticated wild relatives, and species with long generation times. At the same time, multiplex editing presents technical challenges, including complex construct design and the need for robust, scalable mutation detection. We discuss current toolkits and recent innovations in vector architecture, such as promoter and scaffold engineering, that streamline workflows and enhance editing efficiency. High-throughput sequencing technologies, including long-read platforms, are improving the resolution of complex editing outcomes such as structural rearrangements-often missed by standard genotyping-when targeting repetitive or tandemly spaced loci. To fully realize the potential of multiplex genome engineering, there is growing demand for user-friendly, synthetic biology-compatible, and scalable computational workflows for gRNA design, construct assembly, and mutation analysis. Experimentally validated inducible or tissue-specific promoters are also highly desirable for achieving spatiotemporal control. As these tools continue to evolve, multiplex CRISPR editing is poised to become a foundational technology of next-generation crop improvement to address challenges in agriculture, sustainability, and climate resilience.

Gene Editing

A scalable, low-cost, sample hashing workflow for multiomic single-cell analysis using the Seq-Well S3 platform.

In-depth analyses of clinical samples have the potential to provide unparalleled insights into the cellular mechanisms that underlie both health and disease, as well as therapeutic and prophylactic responses. However, these specimens are often paucicellular, necessitating the use of workflows that maximize the amount of information that can be learned. Here we provide a detailed protocol for generating and analyzing single-cell multiomic data from low-input samples with the Seq-Well S3 platform. We further describe a matched pipeline for sample hashing that reduces costs and sources of technical variation in the resulting data while also enhancing throughput. In brief, our streamlined and efficient methodology involves: (1) optionally staining single-cell suspensions with antibody-oligonucleotide conjugates for cell surface protein quantification and/or sample multiplexing; (2) generating Seq-Well S3 sequencing libraries; (3) optionally producing bulk-RNA sequencing libraries via SMART-seq2 to support genetic demultiplexing; and (4) computationally analyzing the resulting data. Each step herein has been designed to leverage readily available reagents and standard laboratory equipment, substantially lowering barriers to entry for researchers. The overall Protocol can yield high-quality multiomic insights from samples in under a week.

Single-Cell Analysis

Integrative Multi-PTM Proteomics Reveals Dynamic Global, Redox, Phosphorylation, and Acetylation Regulation in Cytokine-Treated Pancreatic Beta Cells.

Studying regulation of protein function at a systems level necessitates an understanding of the interplay among diverse posttranslational modifications (PTMs). A variety of proteomics sample processing workflows are currently used to study specific PTMs but rarely characterize multiple types of PTMs from the same sample inputs. Method incompatibilities and laborious sample preparation steps complicate large-scale physiological investigations and can lead to variations in results. The single-pot, solid-phase-enhanced sample preparation (SP3) method for sample cleanup is compatible with different lysis buffers and amenable to automation, making it attractive for high-throughput multi-PTM profiling. Herein, we describe an integrative SP3 workflow for multiplexed quantification of protein abundance, cysteine thiol oxidation, phosphorylation, and acetylation. The broad applicability of this approach is demonstrated using cell and tissue samples, and its utility for studying interacting regulatory networks is highlighted in a time-course experiment of cytokine-treated β-cells. We observed a swift response in the global regulation of protein abundances consistent with rapid activation of JAK-STAT and NF-κB signaling pathways. Regulators of these pathways as well as proteins involved in their target processes displayed multi-PTM dynamics indicative of complex cellular response stages: acute, adaptation, and chronic (prolonged stress). PARP14, a negative regulator of JAK-STAT, had multiple colocalized PTMs that may be involved in intraprotein regulatory crosstalk. Our workflow provides a high-throughput platform that can profile multi-PTMomes from the same sample set, which is valuable in unraveling the functional roles of PTMs and their co-regulation.

Proteomics

DURABLE: A Workflow for Determining Corrosion-Driving and Protective Microbial Mechanisms.

Microbiologically influenced corrosion (MIC) threatens global infrastructure, causing billions of dollars in annual losses. Its persistence stems from unresolved mechanisms─particularly the metabolites produced by microorganisms that drive or inhibit corrosion─and the microbial community structures. Progress has been hindered by the absence of systematic workflows to rapidly and accurately identify MIC-relevant microorganisms and their functions. Here, we present DURABLE (Detection of Unique Corrosion Resistant or Accelerating Biologics in a Laboratory Environment), a pipeline that couples high-throughput microbial screening with genomic and metabolic workflows. We applied the DURABLE workflow to six diesel tank samples and revealed fuel-dependent microbial community structures, which showed greater diversity and evenness in bacterial communities than their fungal counterparts. The workflow used carbon steel beads to rapidly screen over 80 bacterial isolates for corrosive activity, reducing assay time to approximately 2 days compared with the conventional 30-day metal coupon test. More than 40 isolates were identified as corrosive. Further testing using mass spectrometry analysis revealed corrosion-associated metabolites, which were further validated using electrochemical assays. Thus, DURABLE achieved a ∼15-fold increase in screening speed and provided a scalable and mechanistic framework for dissecting MIC dynamics. We expect this advance will enable the development of precision mitigation strategies in hydrocarbon fuel infrastructure.

Bacteria

Functional screening and single-cell cultivation of marine CO2-fixing bacteria via flow-mode Raman-activated cell sorting.

Most marine CO2-fixing microorganisms remain uncultivated due to strong culture bias and low throughput of conventional approaches, which fail to link in situ function with isolated strains and render slow-growing or low-abundance taxa virtually inaccessible. This study presents an integrated single-cell workflow that incorporates 13C-NaHCO3 labeling, high-throughput flow-mode Raman-activated cell sorting (RACS) and microwell cultivation for the isolation of active CO2-fixing bacteria from the Yellow Sea. Function-guided sorting was achieved by monitoring the 13C-induced Raman shifts of carotenoids (ν1 band: ∼1507 to ∼ 1503.78 cm-1 at 24 h). Genomic and physiological analyses identified Paraburkholderia aromaticivorans FR-4 as a novel facultative chemoautotrophic nitrite-oxidizing bacterium (NOB). Its genome encodes complete nitrite oxidation and Calvin cycle pathways, together with key carbon acquisition genes (carbonic anhydrase, bicarbonate transporter). FR-4 grows autotrophically using NO2- as the electron donor and CO2/HCO3- as the carbon source, confirming its ability to couple nitrite oxidation with carbon fixation, while retaining metabolic flexibility for heterotrophic growth. By directly linking in situ carbon-fixing activity, genotype, and phenotype, this workflow provides a targeted strategy for exploring elusive marine CO2-fixing bacteria and overcomes critical limitations of conventional cultivation.

Carbon-fixing

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals

Workflow for Long-Read Amplicon Sequencing of Chikungunya Virus Using Oxford Nanopore Technology.

This protocol provides a comprehensive, step-by-step workflow for whole-genome sequencing of Chikungunya virus (CHIKV) using an amplicon-based strategy optimized for Oxford Nanopore Technologies (ONT) platforms. The procedure includes detailed instructions for sample handling, viral RNA extraction, quality control, cDNA synthesis, multiplex PCR amplification, library preparation, sequencing, and primary bioinformatic processing. The protocol is designed to maximize reproducibility across laboratories and is suitable for genomic surveillance applications, including outbreak investigation and molecular epidemiology, even when working with low-to-moderate viral loads.

Chikungunya virus

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

A python based automated computational framework to classify and comparative genomics analysis of the global diversity of chili leaf curl virus (ChiLCV) strains to understand virus host interactions.

Chili leaf curl virus (ChiLCV) is a Begomovirus chillicapsici that is one of the most devastating viruses impacted on the production of chili in the world, especially in South Asia. In the present study, we combined high-throughput computational genomics with experimental analysis of global diversity. A workflow was created using automated Python scripts to download, curate and process ChiLCV genomes from public database. About 410 complete ChiLCV genomes download from public databases. Using a phylogenetic approach, these isolates were subdivided into 34 strains, belonging to 10 major clades, showing significant genetic diversity. Geographic analysis revealed that Pakistan (207 isolates) and India (148 isolates) were the main sources of ChiLCV diversity and the remainder of the isolates were from Oman, Bangladesh, Iran, Saudi Arabia and Sri Lanka. Recombination was observed as a major evolutionary force as more than twenty recombination events were detected. Analysis of cis-regulatory elements showed a complex structure of the viral promoter, including multiple binding sites for transcription factors, hormone-response elements, light-responsive elements, and stress-responsive elements, indicating a high number of interactions between viral regulatory elements and host signaling pathways. Pangenome analysis showed the presence of a highly dynamic open pangenome made up of strain-specific orthologous groups (species-specific orthogroups). Experimental inoculation of chili plants was also carried out to assess the biological effects of infection, along with phytochemical, FTIR, HPLC, and qPCR analyses.

Begomovirus

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics