PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Short-read sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Coexistence of carbapenemase and hypervirulence-associated genes among Klebsiella pneumoniae high-risk clones in Hungary.

INTRODUCTION: Strains of Klebsiella pneumoniae carrying hypervirulence and carbapenemase genes represent a rapidly emerging global public health threat. Our study aimed to comprehensively characterise the genomics of hypervirulence-associated and carbapenemase genes carrying K. pneumoniae (hv(a)CpKp) isolates in Hungary. MATERIALS AND METHODS: Between January 2022 and April 2024, 89 aerobactin (iucA-D/iutA)-positive non-duplicate carbapenemase-producing K. pneumoniae isolates from 15 Hungarian healthcare institutes underwent short-read (Illumina, MiSeq, NextSeq) whole-genome sequencing, followed by detailed plasmid analysis using long-read sequencing (Nanopore, MinION) in a representative subset of 32 strains. RESULTS: Most isolates (79/89) belonged to the high-risk clone ST147. Hypervirulence-associated (hva) genes-including rmpA/rmpA2, peg344, shiF, iucA-D, and iutA-were universally present, and 59 isolates possessed chromosomally integrated yersiniabactin loci. Most isolates (87/89) carried the bla NDM-1 carbapenemase gene. Hypervirulence-associated genes were most frequently (29/32) associated with IncHI1B/IncFIB(Mar) plasmids. Notably, we identified plasmids carrying both hva and carbapenemase genes-designated as hybrid plasmids-in 13 of 32 strains. The bla NDM-1 was linked to the IS26 transposase and was present in conserved, identical cassettes on all bla NDM-1-carrying plasmids. DISCUSSION/CONCLUSION: Our study identified hv(a)CpKp strains, particularly the ST147 clone, circulating in Hungary. Our findings highlight the need for routine virulence gene monitoring and continuous genomic and plasmid-based surveillance to mitigate the clinical and epidemiological impact of emerging hv(a)CpKp lineages.

Klebsiella pneumoniae↗

scnanoseq: an nf-core pipeline for Oxford Nanopore single-cell RNA-sequencing.

MOTIVATION: Recent advancements in long-read single-cell RNA sequencing (scRNA-seq) have facilitated the quantification of full-length transcripts and isoforms at the single-cell level. Historically, long-read data would need to be complemented with short-read single-cell data in order to overcome the higher sequencing errors to correctly identify cellular barcodes and unique molecular identifiers. Improvements in Oxford Nanopore sequencing, and development of novel computational methods have removed this requirement. Though these methods now exist, the limited availability of modular and portable workflows remains a challenge. RESULTS: Here, we present, nf-core/scnanoseq, a secondary analysis pipeline for long-read single-cell and single-nuclei RNA that delivers gene and transcript-level quantification. The scnanoseq pipeline is implemented using Nextflow and is built upon the nf-core framework, enabling portability across computational environments, scalability and reproducibility of results across pipeline runs. The nf-core/scnanoseq workflow follows best practices for analyzing single-cell and single-nuclei data, performing barcode detection and correction, genome and transcriptome read alignment, unique molecular identifier deduplication, gene and transcript quantification, and extensive quality control reporting. AVAILABILITY AND IMPLEMENTATION: The source code, and detailed documentation are freely available at https://github.com/nf-core/scnanoseq and https://nf-co.re/scnanoseq under the MIT License. Documentation for the version of nf-core/scnanoseq used for this paper, including default parameters and descriptions of output files are available at https://nf-co.re/scnanoseq/1.1.0.

Single-Cell Analysis↗

Genomic Diversity and Extended-Spectrum β-Lactamase Gene Contexts of Community Resident-Carried Escherichia coli in Ecuador.

Community carriage of extended-spectrum β-lactamase (ESBL)-producing Escherichia coli represents an important reservoir of antimicrobial resistance. However, the genomic diversity and population structure of ESBL-producing E. coli circulating in community settings remain poorly characterized. This study aimed to characterize ESBL-producing E. coli isolated from fecal samples of residents in Ecuador, with an emphasis on the diversity and genomic context of ESBL genes. ESBL-producing E. coli was isolated from fecal samples obtained from 55 residents using MacConkey agar supplemented with cefotaxime. Whole-genome sequencing of the isolates was performed using a hybrid approach combining long- and short-read platforms. Plasmids and β-lactamase genes were identified using DFAST and PlasmidFinder. Bacterial identification and antimicrobial susceptibility testing were conducted by MALDI-TOF MS and the broth microdilution method, respectively. ESBL-producing E. coli were isolated from 35 of 55 fecal samples (63.6%). Complete circular genomes were obtained from 31 isolates. All isolates harbored bla CTX-M genes, predominantly belonging to the bla CTX-M-1 group, whereas 65.7% carried bla TEM, mainly bla TEM-1, and related variants. Although β-lactamase genes were predominantly plasmid-borne, chromosomal integration was detected in 40% of the isolates. Notably, 87.5% of the isolates harbored IncF plasmids with multiple replicons. Conserved IS26-flanked transposons carrying bla CTX-M and bla TEM were frequently identified in the plasmids. Phylogenetic analysis revealed substantial genomic diversity across seven phylogroups, together with closely related isolates detected within and between households. These findings provide high-resolution genomic insights into the ESBL determinants circulating in community residents and reveal region-specific patterns of ESBL genomic diversity.

CTX-M β-lactamases↗

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S↗

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157↗

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software↗

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R. auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156 kbp, 85 genes) and mitogenome (1.18 Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69 Gbp, with 94.1% complete BUSCO genes found and 35 482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus↗

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗

Global diversity and evolution of Salmonella enterica serovar Panama: a genomic epidemiology study.

BACKGROUND: Non-typhoidal Salmonella is a globally important bacterial pathogen, typically associated with foodborne gastrointestinal infection. Some non-typhoidal Salmonella serovars can also colonise typically sterile sites in people to cause invasive non-typhoidal Salmonella disease. Salmonella enterica serovar Panama is responsible for a substantial number of cases of human bloodstream infection, but despite its global dissemination, numerous outbreaks, and a reported association with invasive non-typhoidal Salmonella disease, S enterica serovar Panama (S Panama) is understudied. We aimed to describe the genomic epidemiology and evolutionary history of S Panama to provide a vital baseline of understanding for this globally important serovar. METHODS: In this genomic epidemiology study, we analysed S Panama genomes derived from historical collections, national surveillance datasets, and publicly available epidemiological and whole-genome sequencing data which span the years 1931-2019. Maximum likelihood and Bayesian phylodynamic approaches were used to investigate population structure and evolutionary history and to infer geotemporal dissemination. A combination of different bioinformatic approaches with short-read and long-read data were used to characterise geographical and clade-specific trends in antimicrobial resistance (AMR) and genetic markers for invasiveness. FINDINGS: We analysed 836 S Panama genomes, of which 559 (67%) were sequenced as part of this study. The collection represents all inhabited continents and includes isolates collected between 1931 and 2019. We identified the presence of four geographically linked S Panama clades (C1 [ie, the Latin America and the Caribbean clade; n=338], C2 [ie, the European clade; n=124], C3 [ie, the Martinique clade; n=131], and C4 [ie, the Asia and Oceania clade; n=104]) and regional trends in AMR profiles. Most isolates (715 [86%] of 836) were pan-susceptible to antibiotics and belonged to clades circulating in Latin America and the Caribbean (64%, n=458). Most antibiotic-resistant isolates in our collection (113 [93%] of 121) fell within clades C4 (ie, the Asia and Oceania clade) and C2 (ie, the European clade), the latter of which had the highest invasiveness index values based on the conservation of 196 extraintestinal predictor genes. INTERPRETATION: This first large-scale phylogenetic analysis of S Panama has revealed important information about the population structure, AMR, global ecology, and genetic markers of invasiveness of the identified genomic subtypes. Our findings provide an important baseline for understanding S Panama infection. The presence of multidrug-resistant clades with elevated invasiveness index values should be monitored through ongoing surveillance, as such clades could pose an increased public health risk. FUNDING: UK Research and Innovation Global Challenges Research Fund and Biotechnology and Biological Sciences Research Council, UK Medical Research Council, Wellcome Trust, John Lennon Memorial Scholarship, Institut Pasteur, Santé publique France, Fondation Le Roch-Les Mousquetaires, Investissement d'Avenir Programme, and Australian National Health and Medical Research Council.

Humans↗

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans↗

Long-read sequencing reveals widespread novel splicing and neojunction-derived neoantigens in nasopharyngeal carcinoma.

The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; &#x223c;44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.

Humans↗

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames↗

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics↗

Global emergence and transmission dynamics of carbapenemase-producing Citrobacter freundii sequence type 22 high-risk international clone: a retrospective, genomic, epidemiological study.

BACKGROUND: Carbapenemase-producing Citrobacter (CPC) species have recently been recognised as emerging pathogens associated with nosocomial infections in humans. The increased rate of Citrobacter freundii infections is a public health concern and there is a paucity of genomic data regarding its global transmission dynamics. We aimed to characterise the genetic features of CPC species, and their associated carbapenemase-encoding plasmids, obtained from hospitalised patients in China and from publicly available global data, with a particular focus on high-risk clones. METHODS: This was a retrospective, genomic epidemiological study of CPC species obtained from a tertiary hospital in Zhejiang Province, China, from March 5, 2013, to March 5, 2023. We used antimicrobial susceptibility testing, short-read and long-read whole-genome sequencing, phylogenomic analysis, and plasmid structure analysis. A global dataset of complete plasmid sequences encoding blaKPC, blaNDM, and blaIMP was constructed from the National Center for Biotechnology Information (NCBI) RefSeq database to provide insights into their diversity and distribution. All carbapenemase-producing Citrobacter freundii genomes from the NCBI GenBank database were incorporated in the comparative genomic analyses. Bayesian phylogeographical analysis and growth rate assays were carried out to characterise the high-risk C freundii sequence type (ST) 22 clone. FINDINGS: 1724 Citrobacter species isolates were collected from diverse clinical specimens, with 48 identified as CPC species. Citrobacter koseri (22 [46%] of 48) and C freundii (20 [42%]) were the predominant CPC species. Comparative analysis found C freundii carried significantly higher median numbers of plasmid replicons (5&#xb7;0 [IQR 3&#xb7;3-6&#xb7;0] vs 2&#xb7;0 [2&#xb7;0-3&#xb7;0]; p<0&#xb7;0001) and acquired antimicrobial resistance genes (12&#xb7;0 [7&#xb7;3-15&#xb7;8] vs 3&#xb7;0 [3&#xb7;0-5&#xb7;3]; p<0&#xb7;0001) than did C koseri. Molecular characterisation identified Inc-type plasmids, In823::Kl.pn.I3/In1589-like/In837-like integrons, Tn6296/Tn125/Tn5060 transposons, and insertion sequences (eg, IS26, IS3000, IS5, ISAba125, ISCR1), collectively facilitating the dissemination of carbapenemase genes. Global analysis of 3126 carbapenemase-encoding plasmids found epidemic plasmids with broad host ranges and global diversity. Phylogenetic investigation of predominant carbapenemase-encoding plasmids showed their persistence across geographical regions, temporal spans, and Enterobacterales species, exhibiting high genetic similarity to our clinical plasmids. A phylogenetic tree of 726 global carbapenemase-producing C freundii genomes showed that ST22 (227 [31&#xb7;3%]) represents the predominant multidrug-resistant clone across community, health-care, and environmental niches. Transmission across continents contributes to the global predominance of the ST22 clone, which carries a high load of resistance genes (median 15&#xb7;0 [IQR 11&#xb7;0-17&#xb7;0] vs 12&#xb7;0 [3&#xb7;0-16&#xb7;0]; p<0&#xb7;0001) and enhanced plasmid maintenance capacity (median replicons 5&#xb7;0 [IQR 4&#xb7;0-7&#xb7;0] vs 4&#xb7;0 [3&#xb7;0-6&#xb7;0]; p<0&#xb7;0001) relative to non-ST22 clones. INTERPRETATION: Our study provides evidence to suggest that Citrobacter species are emerging carriers of carbapenem-resistance genes. These findings provide insight into the population structure of CPC species and highlight C freundii ST22 as a prominent high-risk international clone. FUNDING: National Natural Science Foundation of China, National Health Commission Scientific Research Fund-Zhejiang Provincial Major Health Science and Technology Plan Project, Zhejiang Province Natural Science Foundation Project, Outstanding Youth Foundation of Jiangsu Province of China, the Priority Academic Program Development of Jiangsu Higher Education Institutions, and Postgraduate Research and Practice Innovation Program of Jiangsu Province.

Citrobacter freundii↗

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis↗

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6&#x2009;Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia↗

Genome size estimation from long read overlaps.

MOTIVATION: Accurate genome size estimation is an important component of genomic analyses such as assembly and coverage calculation, though existing tools are primarily optimized for short-read data. RESULTS: We present LRGE, a novel tool that uses read-to-read overlap information to estimate genome size in a reference-free manner. LRGE calculates per-read genome size estimates by analysing the expected number of overlaps for each read, considering read lengths and a minimum overlap threshold. The final size is taken as the median of these estimates, ensuring robustness to outliers such as reads with no overlaps. Additionally, LRGE provides an expected confidence range for the estimate. We validate LRGE on a large, diverse bacterial dataset and confirm it generalizes to eukaryotic datasets. On bacterial genomes, LRGE outperforms k-mer-based methods in both accuracy and computational efficiency and produces genome size estimates comparable to those from assembly-based approaches, like Raven, while using significantly less computational resources. AVAILABILITY AND IMPLEMENTATION: Our method, LRGE (Long Read-based Genome size Estimation from overlaps), is implemented in Rust and is available as a precompiled binary for most architectures, a Bioconda package, a prebuilt container image, and a crates.io package as a binary (lrge) or library (liblrge). The source code is available at https://github.com/mbhall88/lrge and an archive at https://doi.org/10.5281/zenodo.17183812 under an MIT license.

Genome Size↗