PubMed HealthSearch

SEARCH · PubMed Health

Results for “snakemake”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

PopGLen-a Snakemake pipeline for performing population genomic analyses using genotype likelihood-based methods.

SUMMARY: PopGLen is a Snakemake workflow for performing population genomic analyses within a genotype-likelihood framework, integrating steps for raw sequence processing of both historical and modern DNA, quality control, multiple filtering schemes, and population genomic analysis. Currently, the population genomic analyses included allow for estimating linkage disequilibrium, kinship, genetic diversity, genetic differentiation, population structure, inbreeding, and allele frequencies. Through Snakemake, it is highly scalable, and all steps of the workflow are automated, with results compiled into an HTML report. PopGLen provides an efficient, customizable, and reproducible option for analyzing population genomic datasets across a wide variety of organisms. AVAILABILITY AND IMPLEMENTATION: PopGLen is available under GPLv3 with code, documentation, and a tutorial at https://github.com/zjnolen/PopGLen. An example HTML report using the tutorial dataset is included in the Supplementary Material.

Software

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics

REAPER: a project-centric workflow layer for comparative repeatome analysis.

INTRODUCTION: Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments - a gap that motivates a project-centric workflow layer for repeatome analysis. METHODS: We present REAPER (Repeatome Extended Analysis Pipeline-Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. RESULTS: Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers - including pTa794- and pSc119-class entries in curated databases. DISCUSSION: These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit - and by supporting caching and structured logging - REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.

TAREAN

Paralog-aware assembly and filtering strategies reveal minimal nucleotide variation on the macro germline-restricted chromosome of the zebra finch.

The germline-restricted chromosome (GRC) of passerines is a remarkable tissue-specific chromosome that accumulated paralogs of genes from the regular "A chromosomes" over millions of years, often amplified into dozens of gene copies. In addition to its repetitive content, typically uniparental inheritance, and lack of recombination, the GRC resembles non-recombining sex chromosomes and some B chromosomes, for all of which assembly and single-nucleotide polymorphisms (SNPs) calling are difficult. Here, we first show that much of the Australian zebra finch macro-GRC can be assembled using accurate long reads. We then describe a paralog-aware Snakemake pipeline, ParaVar, to map short reads from the GRC to retrieve GRC regions suitable for haplotype-based analysis. ParaVar reliably calls hundreds of SNPs across the GRC, thereby providing an estimate of nucleotide diversity on the highly repetitive zebra finch macro-GRC. Our results show significantly lower nucleotide diversity (20- to 50-fold lower) on the GRC compared to the mitogenome and autosomes, and a strong phylogenetic discordance between the GRC and the mitochondrial genome. Beyond the contribution of background selection, our results suggest that a single GRC haplotype recently spread through the populations while jumping across matrilines via occasional paternal inheritance. We anticipate that our paralog-aware pipeline will be useful for SNP calling and population genetics analyses of repetitive GRCs, sex chromosomes, and B chromosomes.

Animals

CBIcall: a configuration-driven framework for variant calling in large sequencing cohorts.

MOTIVATION: Variant calling for next-generation sequencing (NGS) data relies on a diverse ecosystem of tools and workflows. Large-scale collaborative studies increasingly adopt federated analysis, where each institution processes sensitive data locally using standardized pipelines. Deploying identical pipelines across multiple centers remains challenging because heterogeneous software environments and computing policies can cause workflow divergence and inconsistent results. RESULTS: We developed CBIcall, a workflow backend-flexible, configuration-driven framework that runs standardized variant-calling pipelines from raw FASTQ files to analysis-ready VCFs. Users define each analysis in a single YAML parameters file, which CBIcall resolves against a controlled workflow registry and resource catalog. The execution driver validates parameters and checks compatibility among pipelines, analysis modes, workflow backends, genome builds, tool versions, and resource bundles. CBIcall supports reproducibility auditing by comparing executions using recorded provenance and output fingerprints. CBIcall dispatches validated workflows natively through Bash, Cromwell, Nextflow and Snakemake backends and provides production-ready pipelines for germline WES, WGS (single-sample or cohort joint genotyping following GATK Best Practices), and mitochondrial DNA analysis. We evaluated analytical performance using public benchmark datasets and validated reproducibility across four computing environments. We further deployed CBIcall in the EU HEREDITARY project, where it processed 1102 samples with both WES and mtDNA pipelines on an institutional HPC system, supporting its suitability for reproducible cohort-scale genomic analyses. AVAILABILITY AND IMPLEMENTATION: CBIcall is open source (GPLv3) and distributed with ready-to-run pipelines; full dependency and installation documentation is available at https://github.com/CNAG-Biomedical-Informatics/cbicall.

Journal Article

Fedflow: cloud orchestration for federated learning with the FeatureCloud platform.

MOTIVATION: Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. RESULTS: We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. AVAILABILITY: Fedflow is open-source and available at https://github.com/W-L/fedflow.

Journal Article

CREMSA: compressed indexing of (ultra) large multiple sequence alignments.

MOTIVATION: Recent viral outbreaks motivate the systematic collection of pathogenic genomes in order to accelerate their study and monitor the apparition/spread of variants. Due to their limited length and temporal proximity of their sequencing, viral genomes are usually organized, and analyzed as oversized Multiple Sequence Alignments (MSAs). Such MSAs are largely ungapped, and mostly homogeneous on a column-wise level but not at a sequential level due to local variations, hindering the performances of sequential compression algorithms. RESULTS: In order to enable an efficient handling of MSAs, including subsequent statistical analyses, we introduce CREMSA (Column-wise Run-length Encoding for MSAs), a new index that builds on sparse bitvector representations to compress an existing or streamed MSA, all the while allowing for an expressive set of accelerated requests to query the alignment without prior decompression. Using CREMSA, a 65 GB MSA consisting of 1.9M SARS-CoV 2 genomes could be compressed into 22 MB using less than half a gigabyte of main memory, while executing access requests in the order of 100 ns. Such a speed up enables a comprehensive analysis of covariation over this very large MSA. We further assess the impact of the sequence ordering on the compressibility of MSAs and propose a resorting strategy that, despite the proven NP-hardness of an optimal sort, induces greatly increased compression ratios at a marginal computational cost. AVAILABILITY AND IMPLEMENTATION: CREMSA is freely accessible at https://gitlab.univ-lille.fr/cremsa/cremsa. The Snakemake workflow for the benchmarks is available at: https://gitlab.univ-lille.fr/cremsa/bench. The data used in the paper is on Zenodo at https://zenodo.org/records/14698859 and https://zenodo.org/records/15100011.

SARS-CoV-2

Mapler: a pipeline for assessing assembly quality in taxonomically rich metagenomes sequenced with HiFi reads.

SUMMARY: Metagenome assembly seeks to reconstruct the most high-quality genomes from sequencing data of microbial ecosystems. Despite technological advancements that facilitate assembly, such as Hi-Fi long reads, the process remains challenging in complex environmental samples consisting of hundreds to thousands of populations. Mapler is a metagenome assembly and evaluation pipeline with a focus on evaluating the quality of Hi-Fi long read metagenome assemblies. It incorporates several state-of-the-art metrics, as well as novel metrics assessing the diversity that remains uncaptured by the assembly process. Mapler facilitates the comparison of assembly strategies and helps identify methodological bottlenecks that hinder genome reconstruction. AVAILABILITY AND IMPLEMENTATION: Mapler is open source and publicly available under the AGPL-3.0 licence at https://github.com/Nimauric/Mapler. Source code is implemented in Python and Bash as a Snakemake pipeline. A snapshot of the code is available on Software Heritage at swh:1:snp:df4f5f02e22ebbab285ec14af58d4d88436ee5d6. Raw data and results are available at https://entrepot.recherche.data.gouv.fr/dataset.xhtml?persistentId=doi:10.57745/2SA8AB.

Metagenome

sedimix: a workflow for the analysis of hominin nuclear DNA sequences from sediments.

SUMMARY: Sediment DNA-the recovery of genetic material from archaeological sediments-is an exciting new frontier in ancient DNA research, offering the potential to study individuals at a given archaeological site without destructive sampling. In recent years, several studies have demonstrated the promise of this approach by extracting hominin DNA from prehistoric sediments, including those dating back to the Middle or Late Pleistocene. However, a lack of open-source workflows for analysis of hominin sediment DNA samples poses a challenge for data processing and reproducibility of findings across studies. Here, we introduce a snakemake workflow, sedimix, for processing genomic sequences from archaeological sediment DNA samples to identify hominin sequences and generate relevant summary statistics to assess the reliability of the pipeline. By performing simulations and comparing our results to two published studies with human DNA from ∼25,000 years ago (including shotgun data from a sediment sample and capture data from touch DNA recovered from a deer tooth pendant) we demonstrate that sedimix yields accurate and reliable inferences. sedimix offers a reliable and adaptable framework to aid in the analysis of sediment DNA datasets and improve reproducibility across studies. AVAILABILITY AND IMPLEMENTATION: sedimix is available as an open-source software with the associated code, example data, and user manual with installation instructions available at https://github.com/jierui-cell/sedimix. A permanent archived version of this release is available via Zenodo: https://doi.org/10.5281/zenodo.17244854.

Animals

Managing workflow executions with WESkit.

SUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section.

Workflow

CIRCE: a scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data.

MOTIVATION: Chromatin 3D folding creates numerous DNA interactions, participating in gene expression regulation. Single-cell chromatin-accessibility assays now profile hundreds of thousands of cells, challenging existing methods for mapping cis-regulatory interactions. RESULTS: We present CIRCE, a fast and scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. CIRCE re-implements the Cicero workflow to analyse single-cell atlases, cutting runtime and memory use by several orders of magnitude. We also provide new options to compute metacells, grouping similar cells to reduce data sparsity. We benchmarked CIRCE against Cicero on two datasets of different sizes and demonstrated the improvement from CIRCE's metacells' strategy with promoter capture Hi-C data. We also evaluated how DNA interaction predictions are impacted by different pre-processing. We observed a negative impact of Cicero's count normalization, and the best performance was obtained with the single-cell count matrix directly. Finally, we demonstrated the scalability of CIRCE by processing a dataset of more than 700 000 cells and 1 million DNA regions in less than an hour. CIRCE should greatly facilitate the prediction of DNA region interactions for scverse and Python users, while providing new and up-to-date pre-processing insights. AVAILABILITY AND IMPLEMENTATION: CIRCE is released as an open-source software under the AGPL-3.0 licence. The package source code is available on GitHub at https://github.com/cantinilab/CIRCE, and its documentation is accessible at https://circe.readthedocs.io. The code to reproduce the presented results is available as a Snakemake pipeline at https://github.com/cantinilab/circe_reproducibility.s.

Software

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence