PubMed HealthSearch

SEARCH · PubMed Health

Results for “Accessory genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Comparative genomics of ESKAPE pathogen species: Integrating pan-genome architecture, antimicrobial resistance, and virulence factor repertoires.

BACKGROUND: ESKAPE pathogens are major causes of hospital-acquired infections and are characterized by extensive antimicrobial resistance (AMR) and diverse virulence mechanisms. Although species-specific pan-genome studies have revealed substantial genomic diversity, the relationships among genome plasticity, resistance burden, and virulence remain incompletely understood across the ESKAPE complex. METHODS: We analyzed 120 high-quality genomes representing six single-species ESKAPE groups (20 genomes per species). Genome quality was assessed using CheckM2. Species-specific pan-genomes were constructed with Roary, AMR genes were identified using AMRFinderPlus, and virulence factors were detected against the VFDB database using DIAMOND. AMR genes were mapped to core and accessory genome compartments through integration of Prokka annotations and Roary outputs. Statistical associations were evaluated using Fisher's exact tests and correlation analyses, with false discovery rate correction applied within each test family. Core-genome maximum-likelihood phylogenies were reconstructed to provide an evolutionary framework. RESULTS: Pan-genome sizes ranged from 4720 to 17,272 genes, with Enterobacter and Pseudomonas possessing the largest accessory genomes. Multidrug resistance (MDR; resistance to ≥3 antimicrobial classes) was detected in 93.3% of strains. After false discovery rate correction, AMR genes remained significantly enriched in the accessory genomes of Enterobacter, Enterococcus, Klebsiella, and Staphylococcus, whereas Acinetobacter and Pseudomonas did not show significant enrichment in either genome compartment. Within-species analyses identified significant positive associations between accessory genome size and AMR class burden in Staphylococcus, Enterococcus, and Enterobacter, whereas the moderate Pearson correlation observed in Pseudomonas was not significant after FDR correction. Virulence factor repertoires varied markedly among species, with Pseudomonas exhibiting the highest burden and Enterococcus the lowest. CONCLUSIONS: ESKAPE pathogens display distinct patterns of resistance and virulence. Accessory genome expansion was associated with higher AMR burden in several species, whereas other species showed no significant association between accessory genome size and AMR burden and no significant enrichment of AMR genes in either genome compartment, highlighting the species-specific nature of AMR evolution.

Virulence Factors

Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species.

SUMMARY: Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety in methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. Using Mycobacterium tuberculosis, a clonally evolving bacterium with a small accessory genome, as a model system, we systematically evaluated sources of variability in pan-genome estimates. Our analysis revealed that differences in assembly type (short-read versus hybrid), annotation pipeline, and pan-genome software, significantly impact predictions of core and accessory genome size. Extending our analysis to two additional bacterial species, Escherichia coli and Staphylococcus aureus, we observed consistent tool-dependent biases but species-specific patterns in pan-genome variability. Our findings highlight the importance of integrating nucleotide- and protein-level analyses to improve the reliability and reproducibility of pan-genome studies across diverse bacterial populations. AVAILABILITY AND IMPLEMENTATION: Panqc is freely available under an MIT license at https://github.com/maxgmarin/panqc.

Genome, Bacterial

A reusable model of pangenome selection informs optimal surveillance strategies over vaccine introductions.

BACKGROUND: The human pathogen Streptococcus pneumoniae is a major cause of disease, including pneumonia and meningitis. The introduction of Pneumococcal Conjugate Vaccines (PCVs) initially reduced the burden of disease through a reduction of colonisation by vaccine-targeted serotypes. However, since PCVs only target a proportion of pneumococcal serotypes, they shift intraspecific competition, eventually allowing non-targeted types to 'replace' vaccine types. Understanding the host and pathogen factors causing replacement is important for future vaccine development. Mechanistic understanding of vaccine replacement dynamics is crucial for forecasting and optimisation of genomic surveillance strategies to evaluate realised vaccine effectiveness. METHODS: We developed a mathematical model of the genomic and demographic factors which explain vaccine replacement, used this model to replicate serotype-frequency changes, and investigated cost-effective genomic surveillance strategies. We extended a forward-time model based on the Wright-Fisher model, developing a user-friendly model framework that describes the post-vaccine dynamics of S. pneumoniae populations. Our model describes vaccine replacement as a function of vaccine impact, immigration of new strains, and negative frequency-dependent selection (NFDS) on the accessory genome content. RESULTS: We used our model to study vaccine replacement in newly sequenced genomic surveillance data from Kathmandu (Nepal), and existing data from Massachusetts (US) and Southampton (UK), with distinct surveillance strategies. We showed that the model with NFDS better replicates replacement dynamics than a null model without NFDS, and that NFDS likely only acts on part of the S. pneumoniae accessory genome. We found consistent estimates for vaccination effectiveness across the different study locations and region-specific genes under NFDS, highlighting the importance of conducting genomic surveillance in each country of interest. By simulating data from the model, we showed that an optimal surveillance strategy prioritises per-sampling sample size over sampling frequency for small sampling budgets. CONCLUSIONS: Our model can be used to predict vaccine replacement dynamics after PCV introduction, and can be easily reapplied to analyse new data from vaccine introductions or new regions. Our model is available in the R package Stubentiger (Studying Balancing Evolution (NFDS) To Investigate Genome Replacement) on GitHub https://github.com/bacpop/Stubentiger .

Streptococcus pneumoniae

Phenotypic and phylogenomic characterization of Lactococcus garvieae isolates from rainbow trout (Oncorhynchus mykiss) in Türkiye.

Lactococcosis is an important bacterial disease of farmed fish and causes substantial economic losses in rainbow trout (Oncorhynchus mykiss) aquaculture. In this study, Lactococcus garvieae isolates recovered from rainbow trout farms in Türkiye were characterized using phenotypic, molecular, and phylogenomic methods. Among 32 presumptive Lactococcus isolates recovered from 127 dead rainbow trout, four were confirmed as L. garvieae and exhibited identical biochemical characteristics, Pulsed Field Gel Electrophoresis (PFGE) profiles, and broad growth tolerance across different pH, salinity, and temperature conditions. All isolates were presumptively classified as resistant to ciprofloxacin and florfenicol, while remaining susceptible to tetracycline and penicillin. Based on the AMR profiles, strain LG2, which exhibited the most susceptible antimicrobial profile among the isolates, was selected for whole-genome sequencing (WGS). WGS of the representative isolate LG2 generated a single 2,214,687-bp chromosomal contig with 38.5% GC content and 99.0% BUSCO completeness. In silico PCR assigned LG2 to serotype I, and the genome contained an intact capsule-associated cps/kps locus. The chromosomal lsa(D) determinant and an mdt(A)-like efflux-associated gene were detected, whereas no plasmid replicons or acquired quinolone or florfenicol resistance genes were identified, indicating discordance between the phenotypic and genomic AMR results. Taxonomic verification of 236 publicly available Lactococcus assemblies yielded 41 verified public L. garvieae genomes, which, together with LG2, formed a 42-genome within-species dataset. LG2 was most closely related to the Turkish isolate OS-37, sharing 99.96% ANI and differing by three core SNPs; both belonged to ST109, whereas the other Turkish isolates belonged to ST139. cgMLST identified a conserved genomic backbone, while pan-genome analysis identified 5,655 gene clusters and an open pan-genome characterized by a large cloud-gene fraction. These findings demonstrate the importance of species verification in Lactococcus population genomics and reveal substantial accessory-genome diversity within L. garvieae. The genomic features of LG2 provide a basis for future pathogenicity and immunogenicity studies, although experimental validation is required. Overall, these findings highlight the importance of local genomic surveillance for understanding L. garvieae population structure and provide a genomic framework for future region-specific vaccine research.

Animals

Comparative genomic analysis reveals distinct population structure in Legionella anisa.

Legionella anisa has been frequently isolated from engineered water systems; however, its population structure remains understudied compared to Legionella pneumophila. Here, we generated complete genome sequences for four L. anisa isolates recovered from a healthcare facility in Rimouski, Canada. Further the population structure of this species was investigated by performing comparative genomic analyses of the genomes generated in this study together with publicly available L. anisa genomes. Genome-wide phylogenetic analysis revealed the presence of three distinct clades separated by substantial genetic divergence (∼500 SNP), with the Rimouski isolates forming a tightly clustered group, suggesting a clonal lineage. Comparative pangenome analysis indicated moderate core genome conservation accompanied by a highly variable accessory genome (∼50%). The isolates characterized in this study harbored multiple plasmids encoding genes associated with conjugation, heavy metal resistance, and other stress-related functions, suggesting potential roles in environmental persistence. Previous studies have shown that L. anisa can proliferate within protozoan host cells, although outcomes vary depending on the host species. Our isolates showed efficient proliferation within Acanthamoeba castellanii, but not within Vermamoeba vermiformis, under the conditions tested. Together, these findings underscore the genomic diversity of this understudied Legionella species and provide a framework for future investigations regarding environmental persistence and potential pathogenicity.

Legionella anisa, Whole genome sequencing

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

Predicting natural variation in the yeast phenotypic landscape with machine learning.

Most organismal traits result from the complex interplay of many genetic and environmental factors, making their prediction difficult. Here, we used machine learning (ML) models to explore phenotype predictions for 223 traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains isolated worldwide. We benchmarked a ML pipeline with multiple linear and non-linear models to predict phenotypes from genotypes and gene expression, and determined gradient boosting machines as the best-performing model. Gene function disruption scores and gene presence/absence emerged as best predictors, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The prediction accuracy broadly varied among phenotypes, with stress resistance being easier to predict compared to growth across nutrients. ML identified relevant genomic features linked to phenotypes, including high-impact variants with established relationships to phenotypes, despite these being rare in the population. Near-perfect accuracies were achieved when other phenomics data mostly in similar conditions were used, suggesting that useful information can be conveyed across phenotypes. Overall, our study underscores the power of ML to interpret the functional outcome of genetic variants.

Genetic Variation

Genomic Insights Into the Multimetal Resilience and Biofilm-Templated Nanorod Biosynthesis of Stenotrophomonas bentonitica BII-R7: Bioremediation and Green Nanotechnology Implications.

While microbial metal reduction is widely documented, the genomic determinants that govern the morphological transition from disordered phases to structured nanocrystals remain elusive. Here, we present an integrative study of Stenotrophomonas bentonitica BII-R7, a strain exhibiting exceptional metal resistance and the unique capacity to synthesize crystalline trigonal selenium (t-Se) nanorods. Comparative pangenomic analysis of 38 Stenotrophomonas strains revealed that BII-R7 possesses a notably large accessory genome of 2311 exclusive singletons. We identify a specialized genomic toolkit, absent in all related strains, comprising key metal resistance determinants (e.g., copB, copF, and czcA) alongside extracellular remodelling enzymes (Wzyligase and GH92-glycosyl hydrolase). This unique repertoire confers BII-R7 with significantly higher Cu and Ni tolerance compared to related Stenotrophomonas species, which we hypothesize is fundamental for maintaining metabolic activity in polymetallic environments. RT-qPCR and functional assays confirm that these singletons are not only upregulated under metal stress (e.g., czcA: 42.2-fold) but are also consistent with a critical role in maintaining biofilm resilience. Crucially, we propose a mechanistic model where this unique genetic repertoire governs the assembly of a compositionally distinctive Extracellular Polymeric Substance (EPS). Using a three-state (biofilm, planktonic, EPS-depleted) experiment, we provide direct phenotypic evidence that an intact EPS matrix is required for the efficient transition from amorphous nanospheres to highly ordered crystalline nanorods, and we propose that it acts as a molecular template directing the anisotropic growth of selenium. By bridging genomics and bionanotechnology, this work positions BII-R7 as a promising candidate for sustainable green synthesis and bioremediation, while defining the targeted gene-knockout and complementation experiments now required to establish direct causal roles for the candidate determinants.

Stenotrophomonas

Comparative genomics of the monophasic variant of Salmonella Typhimurium: analysis of Colombian genomes and their relationship with international lineages.

The monophasic variant of Salmonella enterica serovar Typhimurium (STVM) represents a growing threat to global public health owing to its wide dissemination, capacity to adapt to multiple hosts, and antimicrobial resistance. In this study, 98 STVM isolates recovered in Colombia (57 from humans and 41 from pig farms and abattoirs) were genomically characterized between 2015 and 2022 and compared with 102 representative genomes of international lineages by whole-genome sequencing (WGS) and phylogenomic analysis. Phylogenomic analysis revealed the existence of two well-defined endemic lineages in Colombia (Clusters 1 and 2), arising from independent introduction events and subsequent local stabilization. Both lineages comprise isolates of human and swine origin without clear phylogenetic separation by host species, suggesting active zoonotic cocirculation and closely integrated interspecies transmission dynamics. Marked differences were observed in the accessory genome, including the differential presence of prophages (e.g., Gifsy-2, Fels-2, SW9), virulence plasmids, and resistance profiles. The Colombian lineages exhibited a high frequency of the pSTV plasmid (85%, n = 84/98) and a substantial burden of resistance determinants to quinolones (such as qnrB19, 74.5%; gyrA S83F mutation, 19.4%), phenicols (floR), tetracyclines (tetA, tetB), β-lactams (blaTEM-1B), and heavy metals. In contrast, the Colombian genomes clustered with the European ST34 lineage lacked pSTV but retained resistance and heavy metal operons. These findings demonstrate that international and endemic lineages coexist in Colombia with independent evolutionary trajectories, underscoring the need to strengthen genomic surveillance under the "One Health" approach to anticipate emerging threats and develop integrated control strategies.IMPORTANCEThe monophasic variant of Salmonella Typhimurium (STVM) has emerged as a predominant serovar in both humans and swine internationally. In Colombia, a fundamental question driving this study was whether local isolates belonged to international lineages or represented endemic strains. This study provides the first comprehensive genomic characterization demonstrating that two Colombian endemic lineages circulate simultaneously between humans and pigs without phylogenetic separation by host species, confirming active zoonotic transmission. The results demonstrate the coexistence of both lineages, each with distinctive repertoires of mobile genetic elements and specific antimicrobial resistance profiles. Understanding these transmission dynamics and evolutionary patterns is crucial for public health, as it demonstrates how zoonotic pathogens can establish locally adapted lineages with distinct resistance patterns. The genomic evidence of sustained interspecies circulation highlights the critical need for integrated surveillance strategies under the "One Health" framework. This will enable anticipating emerging threats, tracing transmission routes, and developing targeted interventions in food production systems.

One Health

Phage-encoded sRNA counteracts xenogeneic silencing in pathogenic E. coli.

Horizontal gene transfer introduces foreign DNA that can disrupt cellular processes and is therefore subject to xenogeneic silencing by nucleoid-associated proteins such as H-NS and Hha. In Enterohaemorrhagic Escherichia coli (EHEC), prophages make up a large fraction of the accessory genome and encode many virulence factors, yet their expression must overcome this silencing. We identify a prophage-encoded small RNA (sRNA), HnrS, that functions as an anti-silencing factor by targeting the H-NS paralogue Hha. HnrS is a short (66-nt) sRNA that is enriched in the locus of enterocyte effacement (LEE⁺) E. coli strains and present in up to nine copies in EHEC and Enteropathogenic Escherichia coli (EPEC) genomes. HnrS base-pairs with the hha ribosome-binding site to inhibit translation, thereby modulating Hha-H-NS repression of virulence loci including the LEE type III secretion system. Loss of HnrS alters motility, T3SS expression, and a subset of Hha-regulated genes. These findings reveal an RNA-based counter-silencing strategy encoded by prophage to relieve xenogenic silencing.

Escherichia coli Proteins

Genome-Wide Association Study of Accessory Atrioventricular Pathways.

IMPORTANCE: Understanding of the genetics of accessory atrioventricular pathways (APs) and affiliated arrhythmias is limited. OBJECTIVE: To investigate the genetics of APs and affiliated arrhythmias. DESIGN, SETTING, AND PARTICIPANTS: This was a genome-wide association study (GWAS) of APs, defined by International Classification of Diseases (ICD) codes and/or confirmed by electrophysiology (EP) study. Genome-wide significant AP variants were tested for association with AP-affiliated arrhythmias: paroxysmal supraventricular tachycardia (PSVT), atrial fibrillation (AF), ventricular tachycardia, and cardiac arrest. AP variants were also tested in data on other heart diseases and measures of cardiac physiology. Individuals with APs and control individuals from Iceland (deCODE Genetics), Denmark (Copenhagen Hospital Biobank, Danish Blood Donor Study, and SupraGen/the Danish General Suburban Population Study [GESUS]), the US (Intermountain Healthcare), and the United Kingdom (UK Biobank) were included. Time of phenotype data collection ranged from January 1983 to December 2022. Data were analyzed from August 2022 to January 2024. EXPOSURES: Sequence variants. MAIN OUTCOMES AND MEASURES: Genome-wide significant association of sequence variants with APs. RESULTS: The GWAS included 2310 individuals with APs (median [IQR] age, 43 [28-57] years; 1252 [54.2%] male and 1058 [45.8%] female) and 1 206 977 control individuals (median [IQR] year of birth, 1955 [1945-1970]; 632 888 [52.4%] female and 574 089 [47.6%] male). Of the individuals with APs, 909 had been confirmed in EP study. Three common missense variants were associated with APs, in the genes CCDC141 (p.Arg935Trp: adjusted odds ratio [aOR], 1.37; 95% CI, 1.24-1.52, and p.Ala141Val: aOR, 1.55; 95% CI 1.34-1.80) and SCN10A (p.Ala1073Val: OR, 1.22; 95% CI, 1.15-1.30). The 3 variants associated with PSVT and the SCN10A variant associated with AF, supporting an effect on AP-affiliated arrhythmias. All 3 AP risk alleles were associated with higher heart rate and shorter PR interval, and have reported associations with chronotropic response. CONCLUSIONS AND RELEVANCE: Associations were found between sequence variants and APs that were also associated with risk of PSVT, and thus likely atrioventricular reentrant tachycardia, but had allele-specific associations with AF and conduction disorders. Genetic variation in the modulation of heart rate, chronotropic response, and atrial or atrioventricular node conduction velocity may play a role in the risk of AP-affiliated arrhythmias. Further research into CCDC141 could provide insights for antiarrhythmic therapeutic targeting in the presence of an AP.

Humans

Genetic and Microbial Analysis of Invasiveness for Escherichia coli Strains Associated With Inflammatory Bowel Disease.

BACKGROUND & AIMS: The adherent-invasive Escherichia coli (AIEC) pathotype is implicated in inflammatory bowel disease (IBD) pathogenesis. AIEC strains are currently defined by phenotypic measurement of their pathogenicity, including invasion of epithelial cells. This broad definition, combined with the genetic diversity of AIEC across patients with IBD, has complicated the identification of virulence determinants. We sought to quantify the invasion phenotype of clinical isolates from patients with IBD and identify the genetic basis for their invasion into epithelial cells. METHODS: A pangenome with core and accessory genes (genotype) was assembled using whole genome sequencing of 168 E coli samples isolated from 13 patients with IBD. A modified assay for invasion of epithelial cells (phenotype) was established with consideration of antibiotic resistance phenotypes. Isolate genotype was correlated to invasiveness phenotype to identify genetic factors that cosegregate with invasion. RESULTS: Pangenome-wide comparisons of E coli clinical isolates identified accessory genes that can cosegregate with invasion phenotype. These correlations found the acquisition of antibiotic resistance genes in clinical isolates compromised the traditional gentamicin protection assays used to quantify invasion. Therefore, an alternate assay, based on amikacin resistance, identified genes cosegregating with invasion. These genes encode an arylsulfatase, a glycoside hydrolase, and genetic islands carrying propanediol utilization and sulfoquinovose metabolism pathways. CONCLUSIONS: This study highlights the importance of incorporating antibiotic resistance screening for invasion assays used in AIEC identification. Accurately screened invasion phenotypes identified accessory genome elements among E coli IBD isolates that correlate with their ability to invade epithelial cells. These results help explain why single genetic markers for the AIEC phylotype are challenging to identify.

Humans

Pseudomonas aeruginosa adaptation and persistence in the aspergilloma microbiome revealed by integrated multi-omics.

Chronic pulmonary aspergillosis involves the formation of a fungal ball (aspergilloma) in lung cavities. Pseudomonas aeruginosa commonly co-colonizes these lesions; however, the in vivo mechanisms underlying its persistence are unknown. Using a multi-omics approach on resected aspergillomas, we defined the genomic, transcriptional, and metabolic adaptations of P. aeruginosa within this polymicrobial niche. We reconstructed high-quality P. aeruginosa genomes and identified a conserved core genome, along with accessory genes for secondary metabolism, virulence, and antimicrobial resistance. Phylogenomics revealed heterogeneous evolutionary paths among co-colonizing strains. Metatranscriptomics showed stark physiological heterogeneity, from metabolically aggressive to stress-adapted states. High expression of phenazine, quorum-sensing (PQS), siderophore, and secretion-system operons was corroborated by metabolomic detection of phenazine-1-carboxylic acid and 2-heptylquinolin-4(1H)-one, confirming active bacterial antagonism in vivo. Concurrent Aspergillus fumigatus transcriptomics revealed the activation of oxidative stress responses, secondary metabolism (eg fumagillin), and iron scavenging, demonstrating reciprocal competition. Host transcriptomics revealed patient-specific immune signatures that correlated with the metabolic activity of the co-colonizers. This work provides an integrated systems-level analysis of the tri-kingdom aspergilloma ecosystem. P. aeruginosa persistence is driven by genomic plasticity and context-dependent expression of competitive pathways, shaped within a chronic inflammatory environment. These findings redefine aspergillomas as active polymicrobial consortia, establishing a framework for targeting resilient microbial communities in chronic lung disease.

Multiomics

Horizontal transfer of accessory chromosomes in fungi - a regulated process for exchange of genetic material?

Horizontal transfer of entire chromosomes has been reported in several fungal pathogens, often significantly impacting the fitness of the recipient fungus. All documented instances of horizontal chromosome transfers (HCTs) showed a marked propensity for accessory chromosomes, consistently involving the transfer of an accessory chromosome while other chromosomes were seldom, if ever, co-transferred. The mechanisms underlying HCTs, as well as the factors regulating the specificity of HCTs for accessory chromosomes, remain unclear. In this perspective, we provide an overview of the observed propensity in reported cases of horizontal chromosome transfers. We hypothesize the existence of a signal that distinguishes mobile, i.e., horizontally transferred, accessory chromosomes from the rest of the donor genome. Recent findings in Metarhizium robertsii and Magnaporthe oryzae, suggest that a mobile accessory chromosome may contain putative histones and/or histone modifiers, which could generate such a signal. Based on this, we propose that mobile accessory chromosomes may encode the machinery required for their own horizontal transmission, implying that HCT could be a regulated process. Finally, we present evidence of substantial differences in codon usage bias between core and accessory chromosomes in 14 out of 19 analysed fungal species and strains. Such differences in codon usage bias could indicate past horizontal transfers of these accessory chromosomes. Interestingly, HCT was previously unknown for many of these species, suggesting that the horizontal transfer of accessory chromosomes may be more widespread than previously thought, and therefore an important factor in fungal genome evolution.

Gene Transfer, Horizontal

Evolutionary dynamics of HIV-1 recombinants: analysis of contemporary and historical viral populations in East Africa.

BACKGROUND: Understanding the genetic evolution of HIV-1 Transmitted/Founder (T/F) virus is crucial for developing effective treatment and prevention strategies due to its rapid mutation and recombination rates. METHODS: This study compared the genetic diversity of 24 contemporary T/F viruses collected between 2016 and 2021 in Uganda and Kenya with 29 historical T/F sequences sampled between 2006 and 2011. RESULTS: Subtype analysis based on near-full-length (NFL) HIV-1 T/F genomes revealed that 57.1% (12/21) of contemporary viruses were recombinants, predominantly involving Subtype A1, D, and increasing Subtype C, with 33.3% (7/21) being A1D recombinants (A1 > D) and 19% (4/21) classified as complex recombinants involving three or more subtypes. Historical viruses showed a similar overall proportion (69%) but were mainly A1D mosaics (D > A1) with recombination confined primarily to the envelope region. In contrast, contemporary viruses shifted towards more complex recombinant patterns affecting additional genomic regions, including pol and accessory genes. Phylogenetic analysis demonstrated that contemporary viruses clustered into distinct, well-supported (98% bootstrap) sub-branches, suggesting divergency attributed to an imbalance in their proportions of subtype A1 and D sequences as well as a different content of A1 and D segments in the A1/D mosaic recombinants. CONCLUSIONS: These findings underscore the dynamic and shifting nature of HIV-1 genetic diversity in East Africa, highlighting the need for continuous molecular surveillance and region-specific treatment guidelines.

HIV-1

Comparative performance of portable DNA extraction protocols and bioinformatics workflows for rapid detection of gram-negative bacteria and antimicrobial resistance using Oxford Nanopore sequencing.

Oxford Nanopore Technology (ONT) enables rapid, portable pathogen identification and antimicrobial resistance (AMR) detection, but the reliability of downstream genomic analyses is highly dependent on DNA extraction quality, particularly in resource-limited settings. This study comparatively evaluated four portable bacterial DNA extraction protocols derived from three commercial kits to determine their impact on nanopore sequencing performance, bioinformatics workflow completion, and field deployability. Six gram-negative bacterial isolates (Escherichia coli, n = 4; Pseudomonas sp., n = 1; and Salmonella sp., n = 1) were processed using four extraction protocols: SwiftX DNA, SwiftX DNA with proteinase K (ProtK), SwiftX ParaBact, and NucleoSpin Microbial. Twenty-four resulting DNA extracts were sequenced on a single multiplexed MinION R10.4.1 flow cell. Sequencing data were analyzed using validated Galaxy-based generic and species-specific pipelines. Workflow completion was defined as successful progression through quality control, assembly, virulence, plasmid, and AMR detection modules. DNA purity varied substantially by extraction protocol and was strongly associated with successful workflow completion (Kruskal-Wallis, P = 0.0006). Accordingly, NucleoSpin Microbial achieved 100% workflow completion, and SwiftX ParaBact achieved 83%, while both SwiftX DNA-based protocols failed to complete full workflows. Importantly, key AMR genes required to classify isolates as multidrug-resistant were consistently detected using both NucleoSpin Microbial and SwiftX ParaBact extractions. However, NucleoSpin Microbial assemblies showed significantly higher contiguity and enabled a broader, more complete detection of virulence factors, pathogenicity islands, plasmid replicons, and accessory AMR genes, reflecting enhanced genomic resolution.IMPORTANCERapid whole-genome sequencing is increasingly used to detect antimicrobial resistance and guide public health responses, but its reliability depends strongly on how bacterial DNA is extracted. In this study, we have shown that DNA extraction method choice has a major impact on Oxford Nanopore sequencing performance across clinically relevant gram-negative bacteria. While silica column-based extraction maximized genomic completeness and analytical depth, paramagnetic bead-based reverse purification offered superior portability with sufficient resolution for frontline AMR surveillance. These findings highlight a practical trade-off between field deployability and high-resolution genomic characterization in low-resource settings.

DNA extraction

PanForest: predicting genes in genomes using random forests.

MOTIVATION: The presence or absence of some genes in a genome can influence whether other genes are likely to be present or absent. Understanding these gene co-occurrence and avoidance patterns reveals fundamental principles of genome organization, with applications ranging from evolutionary reconstruction to rational design of synthetic genomes. RESULTS: PanForest, presented here, uses random forest classifiers to predict the presence and absence of genes in genomes from the set of other genes present. Performance statistics output by PanForest reveal how predictable each gene's presence or absence is, based on the presence or absence of other genes in the genome. Further, PanForest produces statistics indicating the importance of each gene in predicting the presence or absence of each other gene. The PanForest software can run serially or in parallel, thereby facilitating the analysis of pangenomes at Network of Life scale.A pangenome of 12 741 accessory genes in 1000 Escherichia coli genomes was analysed in around 5 h using eight processors. To demonstrate PanForest's utility, we present a case study and show that certain genes associated with resistance to antimicrobial drugs reliably predict the presence or absence of other genes associated with resistance to the same drug. Further, we highlight several associations between those genes and others not known to be associated with antimicrobial resistance (AMR), or associated with resistance to other drugs. We envisage PanForest's use in studies from multiple disciplines concerning the dynamics of gene distributions in pangenomes ranging from biomedical science and synthetic biology to molecular ecology. AVAILABILITY AND IMPLEMENTATION: The software if freely available with a full manual and can be found with at www.github.com/alanbeavan/PanForest DOI: https://doi.org/10.5281/zenodo.17865482.

Software

Conserved assembly architecture of the essential herpesvirus packaging accessory factor.

To create a new wave of infectious virions, all herpesviruses require an accessory factor of unknown function to package their viral genomes into nascent capsids. Here, we present cryo-EM structures of the packaging accessory factor from the α-herpesvirus herpes simplex virus type 1 (HSV-1, UL32) and the β-herpesvirus human cytomegalovirus (HCMV, UL52). Unlike homologs from the γ-herpesviruses, neither UL32 nor UL52 form stable homopentameric rings. UL52 forms incomplete pentameric rings lacking one or two protomers. UL32 does not form stable higher-order species, but stabilization through chemical crosslinking revealed a novel quaternary structure where three pentameric rings assemble into a "tripentamer." Our results reveal that herpesvirus packaging accessory factors adopt distinct oligomeric states but are constrained to pentameric symmetry. Assembly of protomers into a ring creates a positively charged central channel that we show is critical for infectious virus production in HSV-1. Taken together, our study points to a structurally conserved, essential function of packaging accessory factors across the Herpesviridae.

Virus Assembly