PubMed HealthSearch

SEARCH · PubMed Health

Results for “TAREAN”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1 recordsLinked to original sources

REAPER: a project-centric workflow layer for comparative repeatome analysis.

INTRODUCTION: Repeatome characterization from short-read sequencing data is widely performed using RepeatExplorer2/TAREAN. However, long-lived multisample projects and explicit comparative designs are often executed as ad hoc command sequences that are hard to version, rerun, and monitor on shared compute environments - a gap that motivates a project-centric workflow layer for repeatome analysis. METHODS: We present REAPER (Repeatome Extended Analysis Pipeline-Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses. REAPER does not implement a new repeat-discovery algorithm; it is an orchestration layer, and biological accuracy for clustering and satellite calling depends on the underlying RepeatExplorer2/TAREAN and satMiner methods it coordinates. REAPER standardizes: Read QC Deterministic subsampling and preparation RepeatExplorer2/TAREAN execution via seqclust, with satMiner-inspired iterative assembly Post-TAREAN BLAST-based annotation against curated repeat collections (optionally including taxon-scoped NCBI-derived resources with freshness checks) Optional graph-based comparative reports The pipeline makes comparative read allocation, prefix policy, and analysis-ready tables explicit; caching supports incremental reruns and structured logs support monitoring. Performance was assessed using a Triticeae short-read dataset (five samples), with rule-level logging of runtime and memory across pipeline stages. RESULTS: Rule-level performance logs show that graph-based clustering dominates runtime and memory, while QC and preparation steps are lightweight by comparison. Graph-report annotations for the Triticeae project additionally link high-ranking clusters to established repeat markers - including pTa794- and pSc119-class entries in curated databases. DISCUSSION: These findings illustrate biologically interpretable outputs (recovery of known Triticeae repeat markers) alongside quantitative performance metrics (identification of graph-based clustering as the dominant computational cost). By making comparative read allocation, prefix policy, and analysis-ready tables explicit - and by supporting caching and structured logging - REAPER supports reproducible comparative repeatome analysis in evolving multisample projects. As an orchestration layer rather than a discovery algorithm, REAPER's contribution lies in reproducibility, monitorability, and comparative-analysis infrastructure, with biological accuracy remaining contingent on the underlying RepeatExplorer2/TAREAN and satMiner methods.

TAREAN