PubMed HealthSearch

SEARCH · PubMed Health

Results for “reproducible workflows”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

DiaReport: reproducible workflow for differential expression analysis and interactive reporting in DIA-based proteomics.

MOTIVATION: Data-independent acquisition (DIA) has become the preferred data acquisition method for mass spectrometry-based proteomics, yet, reproducible workflows for differential expression (DE) analysis and results reporting remain limited. We present DiaReport, an R package that performs precursor- and protein-level DE analysis from DIA-NN output using MSqRob and QFeatures, while generating high-quality, interactive HTML reports through Quarto. DiaReport integrates precursor data, filtering of missing values, normalization, protein summarization and statistical modeling within a single function, supporting both simple pairwise as well as complex experimental designs. The package provides structured outputs and configuration files to ensure computational reproducibility across different studies. To accommodate diverse research needs, DiaReport includes multiple reporting templates tailored to different proteomic applications. Applying DiaReport to an extracellular vesicle (EV) proteomics dataset demonstrates its ability to efficiently analyze DIA data and provide rapid insights into sample quality and protein level differences. AVAILABILITY: DiaReport is an open-source R package available at https://github.com/Gevaert-Lab/diareport (DOI: 10.5281/zenodo.20120604). The package is platform-independent and distributed under the MIT license. Reports are generated using Quarto and require only standard R dependencies. Detailed documentation, installation guides and usage vignettes are provided within the repository. The interactive HTML reports discussed in this study, including the UPS2 benchmark and EV case study, are archived on Zenodo (10.5281/zenodo.20122506 and 10.5281/zenodo.20123378).

Proteomics

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Towards a Robust cell-free DNA Isolation Protocol for NGS Applications in a Clinical Molecular Diagnostics Setting.

Cell-free DNA (cfDNA), released from apoptotic and necrotic cells into body fluids, is a non-invasive source of genetic information for disease prediction, diagnosis, and monitoring. However, its low abundance makes cfDNA highly susceptible to various pre-analytical influences, potentially increasing high molecular weight (HMW) or genomic DNA (gDNA) compromising downstream cfDNA analyses. This study evaluated the impact of different cfDNA-stabilizing blood collection tubes (BCT; Cell-Free DNA BCT, Streck; S-Monovette cfDNA Exact, Sarstedt) stored at room temperature for 1, 5, or 10 days, prior to plasma isolation using different isolation methods (magnetic bead-based or silica column-based) on cfDNA stability and yield. DNA quantity and quality were assessed by fluorometric quantification, automated fragment analysis, and gene-specific quantitative PCR. Streck-based workflows maintained stable cfDNA yields and characteristic mononucleosomal fragmentation profiles across all storage times. In contrast, Sarstedt tubes showed reduced cfDNA concentrations after 5 days and a pronounced increase at 10 Days, accompanied by high-molecular weight DNA patterns consistent with white-blood cells (WBC) lysis. These trends were largely independent of the extraction method. Overall, the results demonstrate that blood collection tube chemistry critically influences cfDNA integrity during delayed processing. Streck tubes, particularly when combined with silica column-based isolation method, provided the most robust and reproducible workflow for routine molecular diagnostics, whereas Sarstedt tubes produced physiologically implausible results after extended storage.

blood collection tubes

Liquid biopsy-based detection of circulating and exfoliated cholangiocarcinoma tumor cells from blood and bile using heparan sulfate octasaccharides on integrated microfluidic systems.

Early diagnosis of cholangiocarcinoma (CCA) remains challenging because existing diagnostic approaches often lack sufficient sensitivity for reliable detection of early-stage disease. Circulating tumor cells (CTCs) in blood and exfoliated tumor cells (ETCs) in bile represent valuable targets for liquid biopsy-based detection; however, their low abundance and the complexity of clinical sample analysis pose substantial technical challenges for reliable enrichment and identification. Herein, we present a reproducible workflow for isolating and identifying CCA tumor cells from blood for CTCs and bile for ETCs using synthetic cell-surface heparan sulfate (HS) octasaccharide-functionalized magnetic beads (MBs) on integrated microfluidic systems. The method combined sample pre-processing, magnetic bead-based enrichment, controlled low-shear mixing and immunofluorescence-based identification into a unified workflow compatible with distinct clinical sample types. Key operational parameters, including MB concentration, mixing frequency, and pressure settings, were detailed to facilitate consistent performance. Using this workflow, tumor cell capture rates of approximately 70% in bile (for ETCs) and blood (for CTCs) were achieved, with a total processing time of 60-90 min per sample under clinically relevant low-abundance conditions. The platform enables reliable detection of as few as 1 tumor cell per mL of blood or bile. This method provides a practical and adaptable strategy for glycosaminoglycan-mediated liquid biopsy applications and may be extended to other tumor-cell enrichment workflows involving heterogeneous cell-surface interactions.

Humans

VINE-seq and MultiVINE-seq for single-nucleus and multiome profiling of the brain vasculature.

The human cerebrovasculature is a critical yet historically understudied component of neurological health. Dysfunction of the diverse endothelial, mural, and perivascular cells that comprise cerebral vessels is central to diseases ranging from stroke to Alzheimer's disease. However, characterizing these cell populations at a molecular level has proven exceptionally challenging. Encased within a robust basement membrane, vascular cells resist standard dissociation methods, leading to their systematic depletion and underrepresentation in existing single-nucleus genomic atlases. This has created a major blind spot in neuroscience. To overcome this barrier, we developed vessel isolation and nucleus extraction for sequencing (VINE-seq) and its advanced iteration, MultiVINE-seq. The protocol provides a robust, reproducible workflow for the enrichment and high-resolution profiling of vascular, perivascular, and immune cells from fresh or frozen human and mouse brain tissue. First, intact vessels (predominantly capillaries and small arterioles/venules, 100 µm in diameter) are isolated from homogenized brain tissue via dextran-based density-gradient centrifugation, separating the vascular pellet from myelin and the parenchymal fraction. Second, the collected vessels are rigorously washed over a cell strainer to remove trapped contaminants. A critical innovation lies in the third stage: the optimized extraction of nuclei from purified vessels using enzymatic digestion. After extraction, the protocol uses fluorescence-activated cell sorting (FACS) to ensure collection of high-purity nuclei suitable for widely used droplet-based sequencing platforms (e.g., 10x Genomics single cell 3' or multiome). This protocol requires 4-5 h to complete and can be carried out by researchers with single-cell and flow cytometry training.

Journal Article

The clinical promise of mass spectrometry-based single-cell proteomics: from bedside to bench.

INTRODUCTION: Single-cell proteomics (SCP) is entering into a transformative phase, moving beyond technically demanding benchmarking studies toward robust and reproducible workflows capable of quantifying thousands of proteins per cell. These advances highlight SCP's potential to address clinically relevant questions by resolving cellular and pathological heterogeneity that remains obscured in bulk proteomics. AREAS COVERED: This review discusses current advances, challenges, and clinical applications of SCP based on literature identified through searches in major scientific databases. Many clinically relevant samples remain underexplored in SCP studies, in part because their application requires careful evaluation of pre-analytical variables that can strongly influence proteomic readouts. Current SCP methodologies vary according to sample type, experimental conditions, and available resources. Compared with single-cell RNA sequencing, SCP remains limited in cellular throughput, making it challenging to define optimal sample sizes and to reliably detect both abundant and rare cell populations. These limitations also make dataset integration difficult, as reduced cellular coverage and sampling depth increase data sparsity. Moreover, implementing quality control strategies across sequential SCP experiments is essential to ensure data robustness, comparability, and accurate biological interpretation. EXPERT OPINION: Applying SCP to clinical samples advances our understanding of biological complexity and holds potential to drive progress in translational and precision medicine.

Humans

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Cytokines and Inflammatory Gene Polymorphisms Associated With Nosocomial Pulmonary Infection After Spontaneous Intracerebral Hemorrhage.

Nosocomial pulmonary infection is a frequent complication after spontaneous intracerebral hemorrhage and may worsen neurological recovery, prolong hospitalization, and increase clinical burden. This retrospective clinical-laboratory study presents a reproducible workflow for evaluating inflammatory biomarker and host immune-genetic profiles associated with nosocomial pulmonary infection after primary spontaneous intracerebral hemorrhage. Patients are classified according to whether nosocomial pulmonary infection occurs after admission. Peripheral venous blood is collected in the early post-admission period under standardized pre-analytical conditions. Serum is separated, aliquoted, and stored for enzyme-linked immunosorbent assay measurement of IL-1β, IL-6, IL-10, IL-17, IFN-γ, TNF-α, TLR2, TLR4, and TLR9. In parallel, genomic DNA is extracted from anticoagulated whole blood and used for polymerase chain reaction-restriction fragment length polymorphism genotyping of selected cytokine- and Toll-like receptor-related loci. The workflow also includes quality-control procedures for sample handling, duplicate ELISA measurements, DNA purity assessment, genotype calling, and repeat genotyping. Statistical analysis includes between-group comparison of clinical characteristics and biomarker levels, Hardy-Weinberg equilibrium testing, logistic regression analysis for genotype and allele associations, adjustment for relevant clinical covariates, and false-discovery-rate correction for multiple genetic comparisons. This combined clinical, inflammatory, and immune-genetic workflow may help characterize infection-risk profiles after spontaneous intracerebral hemorrhage, although prospective multicenter validation is still required before routine clinical application.

Humans

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals

pmultiqc: An Open-Source, Lightweight, and Metadata-Oriented QC Reporting Library for MS Proteomics.

The increasing scale and complexity of proteomics data demand robust, scalable, and interpretable quality control (QC) frameworks to ensure data reliability and reproducibility. Here, we present pmultiqc, an open-source Python package that standardizes and generates web-based QC reports across multiple proteomics data analysis platforms. Built on top of the widely adopted MultiQC framework, pmultiqc offers specialized modules tailored to mass spectrometry workflows, with full initial support for quantms, DIA-NN, MaxQuant/MaxDIA, FragPipe, and mzIdentML/mzML-based pipelines. The package computes a wide range of QC metrics, including raw intensity distributions, identification rates, retention time consistency, and missing value patterns, and presents them in interactive, publication-ready reports. By leveraging sample metadata in the Sample and Data Relationship Format format, pmultiqc enables metadata-aware QC and introduces, for the first time in proteomics, QC reports and metrics guided by standardized sample metadata. Its modular architecture allows easy extension to new workflows and formats. Alongside comprehensive documentation and examples for running pmultiqc locally or integrated into existing workflows, we offer a cloud-based service that enables users to generate QC reports from their own data or public PRIDE datasets.

Proteomics

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps

A reproducible computational transcriptomic framework for cell-type-resolved fibroinflammatory-AKT remodeling in human heart failure.

BACKGROUND: Human heart failure involves multicellular transcriptional remodeling, but public transcriptomic studies often remain disconnected from cell-type localization and perturbational interpretation. METHODS: We developed a reproducible computational workflow integrating human left-ventricular bulk transcriptomes, donor-level cell-type pseudobulk results from a human heart-failure single-cell/single-nucleus atlas, external snRNA-seq support, curated module scoring, focused ligand-receptor prioritization and LINCS/L1000 perturbational matching. RESULTS: Cross-cohort analysis identified 14,358 same-direction HF-associated genes, including 1633 replicated HF-up and 785 replicated HF-down genes. Donor-level pseudobulk analysis localized disease remodeling to cardiomyocyte, fibroblast and myeloid compartments. Activated fibroblast and inflammatory myeloid programs defined a fibroinflammatory remodeling axis connected to context-dependent AKT-associated transcriptional shifts. External snRNA-seq support was strongest for fibroblast activation and AKT-associated remodeling, with etiology-dependent heterogeneity across validation resources. L1000FWD screening prioritized safety-aware perturbational hypotheses, including glimepiride and simvastatin as interpretable candidates requiring experimental validation. CONCLUSIONS: This study provides a computational transcriptomic framework linking reproducible human HF signatures, cell-type-resolved fibroinflammatory remodeling and perturbational genomic prioritization without claiming drug efficacy or AKT causality.

Humans