PubMed HealthSearch

SEARCH · PubMed Health

Results for “bioinformatics workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Comparative performance of portable DNA extraction protocols and bioinformatics workflows for rapid detection of gram-negative bacteria and antimicrobial resistance using Oxford Nanopore sequencing.

Oxford Nanopore Technology (ONT) enables rapid, portable pathogen identification and antimicrobial resistance (AMR) detection, but the reliability of downstream genomic analyses is highly dependent on DNA extraction quality, particularly in resource-limited settings. This study comparatively evaluated four portable bacterial DNA extraction protocols derived from three commercial kits to determine their impact on nanopore sequencing performance, bioinformatics workflow completion, and field deployability. Six gram-negative bacterial isolates (Escherichia coli, n = 4; Pseudomonas sp., n = 1; and Salmonella sp., n = 1) were processed using four extraction protocols: SwiftX DNA, SwiftX DNA with proteinase K (ProtK), SwiftX ParaBact, and NucleoSpin Microbial. Twenty-four resulting DNA extracts were sequenced on a single multiplexed MinION R10.4.1 flow cell. Sequencing data were analyzed using validated Galaxy-based generic and species-specific pipelines. Workflow completion was defined as successful progression through quality control, assembly, virulence, plasmid, and AMR detection modules. DNA purity varied substantially by extraction protocol and was strongly associated with successful workflow completion (Kruskal-Wallis, P = 0.0006). Accordingly, NucleoSpin Microbial achieved 100% workflow completion, and SwiftX ParaBact achieved 83%, while both SwiftX DNA-based protocols failed to complete full workflows. Importantly, key AMR genes required to classify isolates as multidrug-resistant were consistently detected using both NucleoSpin Microbial and SwiftX ParaBact extractions. However, NucleoSpin Microbial assemblies showed significantly higher contiguity and enabled a broader, more complete detection of virulence factors, pathogenicity islands, plasmid replicons, and accessory AMR genes, reflecting enhanced genomic resolution.IMPORTANCERapid whole-genome sequencing is increasingly used to detect antimicrobial resistance and guide public health responses, but its reliability depends strongly on how bacterial DNA is extracted. In this study, we have shown that DNA extraction method choice has a major impact on Oxford Nanopore sequencing performance across clinically relevant gram-negative bacteria. While silica column-based extraction maximized genomic completeness and analytical depth, paramagnetic bead-based reverse purification offered superior portability with sufficient resolution for frontline AMR surveillance. These findings highlight a practical trade-off between field deployability and high-resolution genomic characterization in low-resource settings.

DNA extraction

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

A Simplified Workflow for the Prediction of Putative Viral Reads Using NIPT Data.

OBJECTIVE: Non-invasive prenatal testing (NIPT) identifies fetal chromosomal abnormalities by sequencing cell-free fetal DNA (cffDNA). Recent studies suggest the prediction of viral sequences from NIPT data, but current methods lack cost-effectiveness for routine use. This study develops a straightforward workflow to investigate potential viral signatures in pregnant women using NIPT data from 888 Iranian participants. METHOD: Two bioinformatic workflows were compared for predicting viral reads: the traditional method involved mapping reads to the human genome, followed by mapping unmapped reads to viral references, and a direct mapping approach to viral genomes, as proposed in this research. RESULTS: While maintaining reproducibility comparable to the conventional method, the proposed workflow minimizes computational complexity and time usage for data processing. Ultimately, this analysis suggested viral DNA in 24.2% of samples, encompassing 29 distinct species, implying the diversity of the maternal virome. CONCLUSION: This study presents a computationally efficient workflow for the in silico prediction of viral-like sequences from routine NIPT data. Further experimental validation is essential to verify the presence, viability, or clinical relevance of these sequences.

Humans

Teratoma Formation and Genomic Profiling Using Multi-Omics Approaches.

Teratoma formation is the gold standard assay for evaluating the developmental pluripotency of human and mouse embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs). Following subcutaneous injection into immunodeficient mice, pluripotent stem cells spontaneously differentiate into derivatives representing all three embryonic germ layers-ectoderm, mesoderm, and endoderm. Beyond serving as a functional assay for pluripotency, teratomas provide a unique three-dimensional model system for studying early human development and lineage specification in vivo. This chapter describes comprehensive protocols for teratoma formation in immunodeficient mice, tissue processing for multiple downstream genomic applications, and multi-omics profiling approaches. We detail methods for embryonic stem cell culture, teratoma generation via subcutaneous injection, tissue dissection and processing for chromatin immunoprecipitation followed by sequencing (ChIP-Seq), RNA sequencing (RNA-Seq), single-cell multiome profiling combining chromatin accessibility (ATAC-Seq) and gene expression (scRNA-Seq), and histological analysis using hematoxylin and eosin (H&E) staining. Additionally, we provide bioinformatics workflows for analyzing the resulting genomic datasets to characterize the epigenetic and transcriptional landscapes of teratoma-derived tissues. These methods enable comprehensive molecular characterization of developmental processes and provide valuable resources for stem cell biologists studying pluripotency, differentiation, and early embryonic development.

Teratoma

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Global maintenance of histone post-translational modifications during the transition into anoxia in embryos of the annual killifish Austrofundulus limnaeus.

Many organisms have adapted to survive anoxic or hypoxic environments, but the epigenetic responses involved in this successful stress response are not well described in most species. Embryos of the annual killifish Austrofundulus limnaeus have the greatest tolerance to anoxia of all vertebrates, making them a powerful model to study the cellular mechanisms necessary for anoxia tolerance. However, the global histone landscape of this species has never been quantified or explored in relation to stress tolerance. Liquid chromatography-mass spectrometry and a Python bioinformatics workflow were used to identify histones and their post-translational modifications. This pipeline resulted in the detection of 252 unique biologically relevant histone post-translational modifications (hPTMs) (unimod + residue). These PTMs represent 16 types of biologically relevant hPTMs present during both anoxia and normoxia in Wourms' stage 36 embryos. This hPTM library presents an exciting opportunity to study histone modifications across development and in response to environmental stressors. No significant changes in PTM or histone abundance were observed between anoxic and normoxic embryos, suggesting that 24 h of anoxia is not sufficient to induce epigenetic or histone isoform changes at the organismal level. This result is inconsistent with data presented for similar stresses in mammalian cells and thus stabilization of the hPTM landscape may be an adaptation that supports anoxia tolerance.

anoxia

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

SARS-CoV-2

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

DNA sequencing for microbial surveillance in cystic fibrosis airways: advances, challenges, and clinical translation.

SUMMARYDNA sequencing has revolutionized microbial surveillance in cystic fibrosis (CF), transforming pathogen identification from culture-dependent to total microbial community identification using molecular-based approaches. Techniques such as 16S rRNA gene sequencing have uncovered the complexity of the CF airway microbiome, while shotgun metagenomics, metatranscriptomics, and viromics now provide strain-level, functional, and viral insights beyond bacterial identification. Despite these advances, key technical and logistical challenges remain, including the processing of high-viscosity sputum samples, overwhelming host DNA contamination, managing large data sets, and the integration of complex bioinformatic outputs into clinical workflows. Emerging innovations such as host DNA depletion protocols, targeted enrichment panels, and adaptive sampling on Oxford Nanopore platforms are helping to overcome these barriers, improving microbial recovery and sequencing efficiency. As cystic fibrosis transmembrane conductance regulator (CFTR) modulator therapies are changing the lives of people with cystic fibrosis (pwCF), sequencing offers an unprecedented opportunity to track potential microbial adaptation in response. This review investigates current advances, limitations, and translational opportunities in DNA sequencing for CF airway microbiome surveillance, highlighting how these technologies can help reshape research and clinical microbiology in the post-modulator era.

Cystic Fibrosis

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418 bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658 bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's ⍴ = 0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's ⍴ = 0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p > 0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2 > 0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals

Backtracking Cell Phylogenies in the Human Brain with Somatic Mosaic Variants.

Somatic mosaic variants, and especially somatic single nucleotide variants (sSNVs), occur in progenitor cells in the developing human brain frequently enough to provide permanent, unique, and cumulative markers of cell divisions and clones. Here, we describe an experimental workflow to perform lineage studies in the human brain using somatic variants. The workflow consists in two major steps: (1) sSNV calling through whole-genome sequencing (WGS) of bulk (non-single-cell) DNA extracted from human fresh-frozen tissue biopsies, and (2) sSNV validation and cell phylogeny deciphering through single nuclei whole-genome amplification (WGA) followed by targeted sequencing of sSNV loci.

Humans

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning