PubMed HealthSearch

SEARCH · PubMed Health

Results for “Workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software

LCR-modules: a collection of workflows for cancer genome analysis.

MOTIVATION: The surge of genomic data from advanced sequencing technologies is outpacing current analytical pipelines. We introduce LCR-modules, an open-source suite of bioinformatics tools designed for flexible and automated cancer genome data analysis. LCR-modules enables reproducible analysis of diverse cancer genomics data at scale. The suite comprises 49 Snakemake-based workflows organized into three levels, facilitating tasks from low-level quality control to complex cohort-level analyses. LCR-modules supports various sequencing types and integrates pipelines such as mutation calling, expression quantification, and cohort-level aggregation, ensuring flexibility and reproducibility. LCR-modules represents a significant advancement in genomic data analysis, reducing barriers in reproducibility and scalability and has already been applied to a combination of exomes and genomes from over 10 800 samples. AVAILABILITY: No new data were generated in support of this research. The source code for the LCR-modules is openly available at https://github.com/LCR-BCCRC/lcr-modules.

Software

DiaReport: reproducible workflow for differential expression analysis and interactive reporting in DIA-based proteomics.

MOTIVATION: Data-independent acquisition (DIA) has become the preferred data acquisition method for mass spectrometry-based proteomics, yet, reproducible workflows for differential expression (DE) analysis and results reporting remain limited. We present DiaReport, an R package that performs precursor- and protein-level DE analysis from DIA-NN output using MSqRob and QFeatures, while generating high-quality, interactive HTML reports through Quarto. DiaReport integrates precursor data, filtering of missing values, normalization, protein summarization and statistical modeling within a single function, supporting both simple pairwise as well as complex experimental designs. The package provides structured outputs and configuration files to ensure computational reproducibility across different studies. To accommodate diverse research needs, DiaReport includes multiple reporting templates tailored to different proteomic applications. Applying DiaReport to an extracellular vesicle (EV) proteomics dataset demonstrates its ability to efficiently analyze DIA data and provide rapid insights into sample quality and protein level differences. AVAILABILITY: DiaReport is an open-source R package available at https://github.com/Gevaert-Lab/diareport (DOI: 10.5281/zenodo.20120604). The package is platform-independent and distributed under the MIT license. Reports are generated using Quarto and require only standard R dependencies. Detailed documentation, installation guides and usage vignettes are provided within the repository. The interactive HTML reports discussed in this study, including the UPS2 benchmark and EV case study, are archived on Zenodo (10.5281/zenodo.20122506 and 10.5281/zenodo.20123378).

Proteomics

UnionLoops: a workflow for calling chromatin loops across related Hi-C datasets with improved specificity, precision, and sensitivity.

Chromatin loop calling from chromatin interaction data often exhibits substantial variability across related samples. We present UnionLoops, a computational workflow for chromatin loop calling across multiple related samples. UnionLoops integrates information across datasets to determine positions and dataset-specificity of looping interactions. It constructs a unified candidate loop set, applies consistent filtering and aggregation, and evaluates loop support across samples. We demonstrate that UnionLoops increases sensitivity for detecting shared chromatin loops, reduces spurious sample-specific calls, and improves concordance with independent genomic features, including CTCF and cohesin occupancy. UnionLoops enables improved biological interpretation of chromatin loop organization and dynamics across related conditions.

Chromatin

Application of PathoChip to urine-derived nucleic acids for broad microbial profiling in men with suspected prostate cancer: setup of a methodological workflow and pilot feasibility study.

BACKGROUND: Urine-based liquid biopsy is an attractive non-invasive source of prostate cancer (PCa) biomarkers, but urinary microbiome studies have mainly relied on 16S rRNA sequencing or shotgun metagenomics. This pilot study optimized and evaluated a practical workflow using PathoChip - a broad-spectrum microarray designed to detect bacterial, viral, fungal, and parasitic signatures - for microbial profiling of urine sediments from men with suspected PCa, an application not previously established. METHODS: First-morning urine was collected without prostatic massage from 35 men scheduled for biopsy; 19 were diagnosed with PCa and 16 were biopsy-negative. Different urine volumes and extraction strategies were evaluated to optimize DNA/RNA recovery. A setup phase compared 25 ng versus 50 ng of urine DNA and RNA input. DNA/RNA isolated from human B cells was used as reference control. An analysis pipeline was developed to detect outlier probes and create a presence/absence matrix. Reproducibility was assessed via library yield, Pearson correlation, blank-control subtraction, outlier probe detection. Prevalence comparisons were performed between clinical groups. RESULTS: An 8 mL starting volume was chosen as consistently available from self-collected urine. Sequential DNA/RNA extraction using the AllPrep DNA/RNA Micro Kit from sediment provided the best balance between nucleic-acid recovery, purity, and clinical compatibility. Reducing the input from 50 ng to 25 ng preserved highly concordant hybridization profiles, with matched samples clustering together with strong correlations. Exploratory analysis revealed PCa- and grade-associated patterns involving Actinomycetaceae, Aerococcaceae, and Streptococcaceae, with Streptococcaceae enriched in PCa of higher grades (ISUP GG ≥ 2). Other signatures, including Mobiluncus, Prevotella, Rhodotorula, Hymenolepis, and JC polyomavirus, were broadly detected but not PCa-discriminating. CONCLUSIONS: PathoChip can be adapted to urine sediments, generating reproducible microbial profiles from limited DNA/RNA input without prostatic massage. This platform provides a quick and accessible approach to broad screening, extending beyond 16S rRNA sequencing by enabling simultaneous multi-kingdom detection. The observed PCa- and grade-associated patterns are hypothesis-generating and require validation in larger independent cohorts.

Pathochip

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Proteome-Scale Tissue Mapping Using Mass Spectrometry Based on Label-Free and Multiplexed Workflows.

Multiplexed bimolecular profiling of tissue microenvironment, or spatial omics, can provide deep insight into cellular compositions and interactions in healthy and diseased tissues. Proteome-scale tissue mapping, which aims to unbiasedly visualize all the proteins in a whole tissue section or region of interest, has attracted significant interest because it holds great potential to directly reveal diagnostic biomarkers and therapeutic targets. While many approaches are available, however, proteome mapping still exhibits significant technical challenges in both protein coverage and analytical throughput. Since many of these existing challenges are associated with mass spectrometry-based protein identification and quantification, we performed a detailed benchmarking study of three protein quantification methods for spatial proteome mapping, including label-free, TMT-MS2, and TMT-MS3. Our study indicates label-free method provided the deepest coverages of ∼3500 proteins at a spatial resolution of 50 μm and the highest quantification dynamic range, while TMT-MS2 method holds great benefit in mapping throughput at >125 pixels per day. The evaluation also indicates both label-free and TMT-MS2 provides robust protein quantifications in identifying differentially abundant proteins and spatially covariable clusters. In the study of pancreatic islet microenvironment, we demonstrated deep proteome mapping not only enables the identification of protein markers specific to different cell types, but more importantly, it also reveals unknown or hidden protein patterns by spatial coexpression analysis.

Proteome

Targeted Modulation of Abundant Proteins Enhances Proteomic Profiling of Ovarian Cancer Ascites: A Pilot Technical Workflow Comparison.

Ascites from ovarian cancer patients are increasingly recognized as a valuable biofluid for cancer research, as its protein composition reflects the disease state and may reveal biomarkers of treatment sensitivity and response. However, the detection of low-abundance proteins is hindered by the presence of highly abundant proteins such as albumin. In this study, we evaluated five protein preparation methods for their effectiveness in depleting high-abundance or enriching low-abundance proteins in ovarian cancer ascites. The Norgen (Nor), Minutes (Min), and Perchloric acid (PerCA) methods were based on abundant protein depletion, while the Urine (Uri) and Nanomics (Nano) kits focused on low-abundance protein enrichment. Processed samples were analyzed using label-free quantitative bottom-up proteomics by LC-MS/MS, followed by a bioinformatics assessment. Compared with undepleted ascites (UnD), Min, Nor, Nano, and PerCA increased protein identifications, whereas Uri produced profiles similar to those of UnD. Notably, PerCA and Nano enabled the identification of distinct protein subsets associated with cancer-related pathways, including immune responses and autophagy. PerCA enriched transmembrane and secreted immunomodulatory glycoproteins, whereas Nano enrichment primarily captured secreted, nuclear, and cytoplasmic soluble proteins. Overall, our results show that both high-abundance protein depletion and low-abundance enrichment improve ascites proteome coverage, each offering distinct advantages in identifying biologically relevant low-abundance proteins.

Female

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

Proteome-scale tissue mapping using mass spectrometry based on label-free and multiplexed workflows.

Multiplexed bimolecular profiling of tissue microenvironment, or spatial omics, can provide deep insight into cellular compositions and interactions in healthy and diseased tissues. Proteome-scale tissue mapping, which aims to unbiasedly visualize all the proteins in a whole tissue section or region of interest, has attracted significant interest because it holds great potential to directly reveal diagnostic biomarkers and therapeutic targets. While many approaches are available, however, proteome mapping still exhibits significant technical challenges in both protein coverage and analytical throughput. Since many of these existing challenges are associated with mass spectrometry-based protein identification and quantification, we performed a detailed benchmarking study of three protein quantification methods for spatial proteome mapping, including label-free, TMT-MS2, and TMT-MS3. Our study indicates label-free method provided the deepest coverages of ~3500 proteins at a spatial resolution of 50 µm and the highest quantification dynamic range, while TMT-MS2 method holds great benefit in mapping throughput at >125 pixels per day. The evaluation also indicates both label-free and TMT-MS2 provide robust protein quantifications in identifying differentially abundant proteins and spatially co-variable clusters. In the study of pancreatic islet microenvironment, we demonstrated deep proteome mapping not only enables the identification of protein markers specific to different cell types, but more importantly, it also reveals unknown or hidden protein patterns by spatial co-expression analysis.

Journal Article

Novel insights into tomato leaf curl New Delhi virus introduction and evolution in Southeastern France using an advanced long-read sequencing workflow.

The Mediterranean population of tomato leaf curl New Delhi virus (ToLCNDV-ES) is characterized by a high genetic uniformity, distinguishing it from its Asian counterparts. ToLCNDV-ES is thought to have a monophyletic origin, likely resulting from a single recombination event, prior to its spread throughout the Mediterranean region. Following its first detection in southeastern France in 2020, ToLCNDV-ES re-emerged in France in 2022. Our analysis based on advanced long-read sequencing, circular DNA profiling, and phylogeny indicates both local persistence of French ToLCNDV-ES and multiple independent introduction events. Signatures of positive selection were identified in French ToLCNDV-ES populations, whereas no clear evidence of recombination was found. Bayesian time-structured phylogenetic analyses suggest that introductions in France occurred between 2018 and 2021 from the major ToLCNDV-ES clade, while several Italian ToLCNDV-ES isolates diverged prior to the virus introduction in the Mediterranean basin. Overall, this study demonstrates the value of an optimized long-read sequencing approach for resolving circular DNA virus diversity, and sheds light on the complex evolutionary history of ToLCNDV-ES in the Mediterranean Basin, particularly in southeastern France.

France

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

CBIcall: a configuration-driven framework for variant calling in large sequencing cohorts.

MOTIVATION: Variant calling for next-generation sequencing (NGS) data relies on a diverse ecosystem of tools and workflows. Large-scale collaborative studies increasingly adopt federated analysis, where each institution processes sensitive data locally using standardized pipelines. Deploying identical pipelines across multiple centers remains challenging because heterogeneous software environments and computing policies can cause workflow divergence and inconsistent results. RESULTS: We developed CBIcall, a workflow backend-flexible, configuration-driven framework that runs standardized variant-calling pipelines from raw FASTQ files to analysis-ready VCFs. Users define each analysis in a single YAML parameters file, which CBIcall resolves against a controlled workflow registry and resource catalog. The execution driver validates parameters and checks compatibility among pipelines, analysis modes, workflow backends, genome builds, tool versions, and resource bundles. CBIcall supports reproducibility auditing by comparing executions using recorded provenance and output fingerprints. CBIcall dispatches validated workflows natively through Bash, Cromwell, Nextflow and Snakemake backends and provides production-ready pipelines for germline WES, WGS (single-sample or cohort joint genotyping following GATK Best Practices), and mitochondrial DNA analysis. We evaluated analytical performance using public benchmark datasets and validated reproducibility across four computing environments. We further deployed CBIcall in the EU HEREDITARY project, where it processed 1102 samples with both WES and mtDNA pipelines on an institutional HPC system, supporting its suitability for reproducible cohort-scale genomic analyses. AVAILABILITY AND IMPLEMENTATION: CBIcall is open source (GPLv3) and distributed with ready-to-run pipelines; full dependency and installation documentation is available at https://github.com/CNAG-Biomedical-Informatics/cbicall.

Journal Article

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals

Human-AI Interaction With AI-Assisted Tumor Overlays in Pediatric Whole-Body Magnetic Resonance Imaging: Exploratory Reader Study.

BACKGROUND: AI tools have the potential to enhance personalized clinical care, particularly in radiology. However, their integration into clinical workflows remains complex, especially in pediatric oncology, where early cancer detection is critical. Children with Li-Fraumeni syndrome (LFS), a rare cancer predisposition disorder, undergo regular surveillance whole-body magnetic resonance imaging (wbMRI), which presents an opportunity for AI-assisted tumor detection. OBJECTIVE: We evaluated the feasibility of an AI-assisted overlay for highlighting tumor-like regions in pediatric surveillance wbMRI and explored how access to the overlay influenced radiologist workflow, candidate-lesion marking behavior, follow-up recommendations, and perceived workload. METHODS: We developed a patch-based AI segmentation model trained on augmented 2D slices from 675 surveillance wbMRI volumes of pediatric patients with LFS. The model was designed to highlight regions with high tumor probability. A reader study was conducted with 2 radiologists who independently reviewed wbMRI cases both with and without AI assistance. We measured evaluation time, number and location of reader-marked candidate lesions, type of follow-up recommendation, and subjective feedback using structured questionnaires. RESULTS: AI assistance altered interpretation workflows for both radiologists, with mixed effects. On average, the time required to evaluate each case increased when using the AI tool for both radiologists. However, one radiologist had an increase in the number of candidate lesion locations selected with the tool, and one had a decrease in the number of candidate lesion locations selected with the tool. Subjective feedback indicated that one of the radiologists reported lower mental demand with the AI tool, while both radiologists reported lower stress with the AI tool. Interrater variability was evident, underscoring the need for personalized calibration of AI tools. CONCLUSIONS: AI-assisted wbMRI interpretation can improve tumor detection in pediatric cancer surveillance by reducing false negatives. However, its influence on workflow efficiency and interradiologist variability highlights the importance of careful implementation. Successful integration requires addressing challenges such as improving the predictive precision of AI models, offering intuitive end-user designs and instructions, and building trust in AI outputs. AI outputs can influence workflow and behavior in reader-specific ways. Clinical translation will require larger, randomized, multireader studies and model refinement to reduce false positives and quantify lesion-level reader performance. This can help ensure better patient outcomes in addition to reduced clinician burnout.

Humans

Applications of artificial intelligence in robot-assisted surgery: a systematic review.

To characterize applications of artificial intelligence (AI) in robot-assisted surgery, summarize technical and clinical performance, and assess the quality of the available evidence. PubMed, Web of Science Core Collection, and Scopus were searched for English-language journal articles published from 1 January 2020 through 31 October 2025. Randomized, observational, model-development, validation, and feasibility studies evaluating AI in robot-assisted surgery or closely related image-guided minimally invasive workflows were eligible. Two reviewers independently performed study selection, data extraction, and risk-of-bias assessment. Owing to heterogeneity in surgical procedures, AI tasks, analytical units, validation strategies, and outcomes, findings were synthesized descriptively without statistical pooling. The review was registered in the International Prospective Register of Systematic Reviews (CRD420251175699). Seventeen studies were included: seven clinical prediction or decision-support studies, eight intraoperative recognition, segmentation, or image-guided studies, and two training or workflow studies. Five prediction studies reported area-under-the-curve values of 0.74-0.95. Technical studies reported F1 or Dice scores of 0.525-0.995 and task-specific accuracies of 0.840-0.998. Two randomized studies suggested benefits for personalized suturing feedback and automated camera control, but neither established improved patient outcomes. Only one study had low overall risk of bias; the remaining studies were at high or unclear risk or raised some concerns. AI applications in robot-assisted surgery show promise for prediction, intraoperative perception, training, and workflow support. Evidence primarily demonstrates technical feasibility rather than established clinical effectiveness. Independent multicenter validation and prospective evaluation of patient, educational, and workflow outcomes are required before widespread implementation.

Robotic Surgical Procedures