PubMed HealthSearch

SEARCH · PubMed Health

Results for “High throughput workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Systematic Analysis of Tumor Microenvironment Using IOBR.

The Immuno-Oncology Biological Research (IOBR) package is an R-based analysis tool for exploring the tumor microenvironment (TME) and its influence on anti-tumor immunity. Built for high-throughput data-spanning both transcriptomic and genomic profiles-IOBR integrates six analytical modules, including transcriptomic data preprocessing, TME profiling, TME pattern identification, ligand-receptor interaction analysis, genome-TME interaction assessment, and visualization. In this chapter, we walk through a multi-omics workflow using example datasets, illustrating data preparation, distribution analyses, result interpretation, and graphical output. IOBR is open source and is available at https://github.com/IOBR/IOBR and a detailed GitBook ( https://iobr.github.io/book/ ) offers a complete manual and analysis guide for each function.

Tumor Microenvironment

Optimized Amplicon Strategy for Long-Read Sequencing of the Chikungunya Virus Genome.

Chikungunya virus (CHIKV) is a positive-sense RNA alphavirus transmitted to humans primarily by Aedes aegypti and Aedes albopictus mosquitoes. Its global circulation and significant public health impact underscore the need to better understand the molecular mechanisms driving CHIKV pathogenesis and transmission. Although robust molecular biology methods exist for CHIKV genome sequencing, a major limitation for surveillance and research is the inability to determine whether two nucleotide variations co-occur within the same viral genome when they are separated beyond the span of typical short-read designs. Here, we describe an optimized approach for processing CHIKV RNA samples that generates large amplicons suitable for long-read nanopore sequencing. This protocol enables amplification of the complete CHIKV genome in only two or three amplicons and facilitates detection of co-occurring nucleotide variations across 4-7.5 kb within the same molecule, thereby simplifying sequencing workflows and improving resolution in studies of viral evolution.

Chikungunya virus

Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: a multicentre clinical study of two end-to-end commercial workflows.

BACKGROUND: Drug-resistant tuberculosis remains a major obstacle in ending the global tuberculosis epidemic. Deployment of molecular tools for comprehensive drug resistance profiling is imperative for successful detection and characterisation of tuberculosis drug resistance. We aimed to assess the diagnostic accuracy of a new class of molecular diagnostics for drug-resistant tuberculosis. METHODS: We conducted a prospective, cross-sectional, multicentre clinical evaluation of the performance of two targeted next-generation sequencing (tNGS) assays for drug-resistant tuberculosis at reference laboratories in three countries (Georgia, India, and South Africa) to assess diagnostic accuracy and index test failure rates. Eligible participants were aged 18 years or older, with molecularly confirmed pulmonary tuberculosis, and at risk for rifampicin-resistant tuberculosis. Sensitivity and specificity for both tNGS index tests (GenoScreen Deeplex Myc-TB and Oxford Nanopore Technologies [ONT] Tuberculosis Drug Resistance Test) were calculated for rifampicin, isoniazid, fluoroquinolones (moxifloxacin, levofloxacin), second line-injectables (amikacin, kanamycin, capreomycin), pyrazinamide, bedaquiline, linezolid, clofazimine, ethambutol, and streptomycin against a composite reference standard of phenotypic drug susceptibility testing and whole-genome sequencing. FINDINGS: Between April 1, 2021, and June 30, 2022, 832 individuals were invited to participate in the study, of whom 720 were included in the final analysis (212, 376, and 132 participants in Georgia, India, and South Africa, respectively). Of 720 clinical sediment samples evaluated, 658 (91%) and 684 (95%) produced complete or partial results on the GenoScreen and ONT tNGS workflows, respectively, with 593 (96%) and 603 (98%) of 616 smear-positive samples producing tNGS sequence data. Both workflows had sensitivities and specificities of more than 95% for rifampicin and isoniazid, and high accuracy for fluoroquinolones (sensitivity approximately ≥94%) and second line-injectables (sensitivity 80%) compared with the composite reference standard. Importantly, these assays also detected mutations associated with resistance to critical new and repurposed drugs (bedaquiline, linezolid) not currently detectable by any other WHO-recommended rapid diagnostics on the market. We note that the current format of assays have low sensitivity (≤50%) for linezolid and more work on mutations associated with drug resistance is needed. INTERPRETATION: This multicentre evaluation demonstrates that culture-free tNGS can provide accurate sequencing results for detection and characterisation of drug resistance from Mycobacterium tuberculosis clinical sediment samples for timely, comprehensive profiling of drug-resistant tuberculosis. FUNDING: Unitaid.

Humans

Quantitative Fluorescence Imaging of Alphavirus Infection for Antiviral Screenings.

Fluorescence microscopy offers a highly sensitive and versatile approach for investigating alphavirus infection at the cellular level. By combining fluorescently labeled viruses with quantitative image analysis, this method enables detailed spatial and temporal characterization of infection dynamics, including the detection of subtle differences in replication kinetics and cell-to-cell spread. A central aim of this protocol is its application in antiviral screening assays. Image-based quantification of fluorescence intensity provides a robust and reproducible means to assess the efficacy of antiviral compounds, allowing early and sensitive detection of inhibitory effects in infected cells. This facilitates the identification of promising antiviral hits and supports the evaluation of dose-dependent responses. The approach is also well-suited for comparative studies of different alphavirus strains or mutants, as variations in replication behavior and dissemination patterns become readily apparent. Its flexibility, compatibility with multiple cell lines, and straightforward integration into automated imaging platforms makes the method scalable and suitable for high-throughput screening campaigns. Overall, this protocol advances the discovery and evaluation of antiviral strategies. Given that several alphaviruses cause significant human and veterinary diseases, lack approved antiviral therapies, and continue to expand geographically with emerging outbreaks, the identification of novel antivirals remains an urgent priority. Therefore, this fluorescence-based workflow represents a valuable and timely contribution to modern alphavirus research.

Antiviral Agents

RUMINA: high-throughput deduplication of unique molecular identifiers for amplicon and whole-genome sequencing with enhanced error correction.

MOTIVATION: Unique molecular identifiers (UMIs) are widely used in next-generation sequencing to enable accurate molecular counting and error correction. However, challenges remain in accurately collapsing UMI clusters, especially when read counts are low or sparse read clusters arise from barcode sequencing errors. RESULTS: We present RUMINA, a Rust-based pipeline for UMI-aware deduplication and error correction, optimized for both amplicon and shotgun sequencing. RUMINA supports multiple UMI cluster strategies, alongside majority-rule read selection independent of mapping quality, as well as discrete handling of 1-2 read clusters, paired-end merging, and read-length stratification. Benchmarking using simulated HIV population sequencing data and real-world iCLIP and TCR datasets showed that RUMINA improves ultra-low frequency SNV detection (0.01%-1%), reduces false positives, enhances reproducibility, and processes sequencing data up to 10-fold faster than existing tools. By integrating UMI- and sequence-level correction in a high-performance framework, RUMINA offers a fast, scalable, and robust solution for UMI-enabled sequencing workflows. AVAILABILITY AND IMPLEMENTATION: RUMINA is implemented in Rust and distributed as open-source code and precompiled binaries. Source code and installation instructions are available at https://github.com/greninger-lab/rumina. Documentation associated with this manuscript is available at https://github.com/greninger-lab/rumina_paper.

High-Throughput Nucleotide Sequencing

Clinically actionable stratification of uncommon MET fusions: a precision oncology framework.

BACKGROUND: MET fusions represent emerging therapeutic targets in solid tumors; however, functional interpretation of non-canonical variants remains poorly understood, posing a major challenge for precision oncology. METHODS: We conducted a multicenter, pan-cancer study analyzing 23,299 clinical samples using DNA-based next-generation sequencing (NGS) to profile MET fusions. Transcriptional validation was performed using RNA-based NGS on available samples. Preliminary clinical outcomes were assessed in four patients with advanced malignancies harboring uncommon MET fusions who received MET tyrosine kinase inhibitor therapy. RESULTS: We identified 116 MET fusions (incidence: 0.5%), with 55.2% (64/116) classified as uncommon fusions. These uncommon fusions were stratified into: Group A (5’-retained, n = 12), Group B (intergenic/exonic breakpoints, n = 19), Group C (rare partners, n = 23), and Group D (dual fusions, n = 10). RNA validation revealed an overall low transcriptional consistency of 43.8% (14/32) for uncommon fusions, versus 100% for canonical fusions (PTPRZ1::MET, CAPZA2::MET). Notably, most 5’-retained fusions were transcriptionally silent, while some intergenic fusions resolved into expressed canonical partners (e.g. PTPRZ1::MET). Therapeutically, all four MET inhibitor-treated patients achieved partial responses, including pediatric diffuse midline gliomas (DMG) (median OS: 11.2 months) and lung adenocarcinoma (median OS: 34 months), demonstrating preliminary clinical activity. CONCLUSIONS: uncommon MET fusions are heterogeneous at genomic and transcriptional levels. DNA-level findings often do not predict functional transcripts, underscoring the necessity of RNA-based confirmation for clinical interpretation. Despite low overall consistency, a subset retains therapeutic potential. We propose a refined diagnostic framework integrating DNA-based stratification and RNA validation to guide the management of MET-altered cancers in precision oncology workflows.

Humans

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry

Modtector: ultra-fast modification signal mining on mapped sequencing reads.

SUMMARY: Existing tools for RNA epitranscriptomic modification and structural signal analysis are often fragmented, inefficiency, and limited to single signal types. We developed Modtector, an unified tool for extracting mutation and reverse-transcription stop signals from aligned sequencing reads. By using a "count-then-correct" strategy, Modtector reduces computational complexity and enables efficient dual-signal analysis. It achieves multi-fold speedups on large-genome and high-coverage datasets, including completing HEK293 22G data analysis in 5 minutes, and show strong scalability on single-cell datasets with speedups exceeding 50-fold. AVAILABILITY: The source code is available at GitHub (https://github.com/TongZhou2017/modtector) and Crates.io (https://crates.io/crates/modtector). The archived source-code snapshot used in this study is available at Zenodo (DOI: 10.5281/zenodo.20967747), corresponding to GitHub commit 7c60e9d. Workflow examples, datasets, and analysis scripts are available at Zenodo (DOI: 10.5281/zenodo.17316476 and 10.5281/zenodo.18523297).

Humans

A streamlined protocol for small-scale protoplast generation and CRISPR/Cpf1-mediated genome editing in Fusarium oxysporum.

Fusarium oxysporum is a significant threat to agriculture and One Health, requiring advanced molecular tools for functional genomic analyses and biological control agent development. Existing gene-editing methods are hampered by costly protoplast preparation protocols and by CRISPR-Cas9 limitations, such as restricted protospacer adjacent motif (PAM) sequences and complex guide RNA requirements. We engineered an efficient CRISPR/Cpf1 system that overcomes these issues through three main innovations: small-scale protoplast generation using filter column-based methods that greatly reduce enzyme consumption while simplifying workflows, a CRISPR/Cpf1 system with shorter guide RNA design and staggered DNA cleavage to promote homologous recombination, and minimal homology arm strategies that significantly decrease cloning complexity. Extensive validation confirms successful gene targeting with molecular verification and functional analysis via standardized pathogenicity assays. This integrated platform offers affordable, accessible tools for systematic F. oxysporum research, enhancing fundamental understanding of plant-pathogen interactions and supporting high-throughput screening vital for agricultural biotechnology and biological agent development.

CRISPR/Cpf1

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

Toward a unified approach: Considerations for bioinformatic and sequencing activities & data in wastewater surveillance of biologic public health threats.

Genomic technologies such as PCR and next-generation sequencing (NGS) have greatly advanced public health surveillance, especially during COVID-19, by enabling detailed tracking of pathogen spread, origins, and variants. While PCR is vital for targeted detection, falling NGS costs have made large-scale, high-throughput sequencing more feasible, supporting broader pathogen monitoring-including the detection of vaccine escape variants and new strains. Applying NGS to wastewater offers valuable population-level insights but faces challenges such as variable sample complexity, the need for skilled staff, suitable platforms, and robust IT infrastructure. Although there are currently a lot of efforts towards defining guidelines for sampling, analysis, and integrating wastewater data into public health policy, such as the recently published International Cookbook for Wastewater Practitioners, they often lack universal applicability, emphasizing the analytical approaches in favour of the NGS-based approaches. However, standardising protocols for sampling, sequencing, and analysis is crucial to ensure reliable, comparable data across surveillance systems worldwide. Pilot studies and continuous refinement are recommended to overcome implementation hurdles and fully realise the benefits of NGS in wastewater surveillance. This work attempts to outline these challenges and opportunities across the entire wastewater surveillance workflow, from data generation to reporting, and provide some concrete suggestions and considerations across the spectrum of activities. We further highlight that the infrastructure, funding and government-policy context in which surveillance operates acts as an enabling condition for these activities, and that technical standardisation alone is unlikely to deliver durable, comparable surveillance in its absence.

considerations

Decoding sequence recognition code of nucleic acid-binding proteins of human-infecting DNA viruses.

Human-infecting DNA viruses remain major health threats, yet the DNA-recognition mechanisms of their nucleic acid-binding proteins (NBPs) are poorly understood. Here, we systematically profiled 103 viral NBPs from human-infecting DNA viruses, with three NBPs from non-human-infecting DNA viruses as controls, using high-throughput screening. This analysis identified diverse DNA-binding motifs and specificity modules, including convergent recognition of a conserved CCACC motif across phylogenetically distant viruses. Notably, viral NBP binding-site distributions varied with genome size, and several NBPs from small-genome viruses showed enrichment on mitochondrial DNA. Functional assays further supported their mitochondrial association and effects on mitochondrial membrane potential. By integrating an ivTRT-based ssDNA-SELEX workflow, we further found that ssDNA viral NBPs recognize dimer-like and inverted-repeat sequences with potential to form stem-loop structures. Collectively, this study constructs a comprehensive viral NBP DNA-recognition atlas, offering a fundamental resource for elucidating viral genome recognition mechanisms, virus-mitochondria interactions, and developing future antiviral strategies.

Letter

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software

Identification of Genome-Wide Chromatin Structural Aberration in Cancer by Hi-C Analysis.

Aberrant three-dimensional genome organization is a hallmark of cancer, often driving oncogene activation through mechanisms such as enhancer hijacking. High-throughput chromosome conformation capture (Hi-C) maps these interactions on a genome-wide scale. Unlike earlier dilution-based methods, in situ Hi-C performs proximity ligation within intact nuclei, minimizing random ligation noise and enabling fine-scale structure detection. This chapter describes an optimized in situ Hi-C protocol tailored for cancer cell lines using MboI digestion and biotin-mediated pull-down to generate high-complexity libraries. We further outline a computational workflow that extends beyond standard topological mapping of compartments and topologically associating domains to identify cancer-specific aberrations. Specifically, we focus on detecting chromosomal rearrangements (structural variants) and characterizing the distinct circular topology of extrachromosomal DNA. This integrated experimental and analytical framework provides the necessary tools to dissect the spatial dysregulation underlying tumor evolution.

Humans

Integrating Radiogenomics and CSF-Based Liquid Biopsy Sequencing for Precision Neuro-Oncology.

Glioblastoma and diffuse gliomas pose major therapeutic challenges due to marked intratumoral heterogeneity, limited tissue accessibility, and the blood-brain barrier. Tissue-based next-generation sequencing (NGS) remains essential for WHO CNS5 molecular classification, yet it is invasive and poorly suited to serial monitoring. Two complementary non- or minimally invasive approaches have advanced rapidly: radiogenomics, which correlates multiparametric MRI features with genomic alterations, and cerebrospinal fluid (CSF) liquid biopsy sequencing, which detects circulating tumor DNA with high tissue concordance. This review examines the independent progress and synergistic integration of radiogenomics and CSF-NGS. Imaging signatures can non-invasively predict key drivers (IDH1/2, EGFR, TERT, PTEN, TP53) and molecular subtypes, while CSF-ctDNA sequencing enables real-time assessment of clonal evolution, therapy resistance (including post-temozolomide hypermutation), and residual disease. We discuss technical considerations, performance metrics, multimodal artificial-intelligence fusion, and emerging clinical applications for diagnosis, prognosis, treatment selection, and longitudinal surveillance. Critical challenges, standardization, prospective validation, and workflow integration are highlighted. By combining the spatial phenotypic information of radiogenomics with the temporal genomic resolution of CSF sequencing, this multimodal strategy offers a promising path toward precision neuro-oncology and reduced reliance on repeated invasive sampling.

Humans

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning