PubMed HealthSearch

SEARCH · PubMed Health

Results for “host prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

The mouse gut microbiota responds to predator odor and predicts host behavior.

Chronic stressors can alter the mammalian gut microbiota in ways that mediate host stress responses, but the impacts of acute stressors on these interactions are less well understood. Here, we show that brief exposure of wild-derived mice to predator odor altered gut-microbiota composition, which in turn predicted host behavior. We investigated the individual and combined effects of 15-minute exposures to synthetic fox fecal odor and 30 days of chronic social isolation, an established chronic stressor. Using ethological assays, visceral adipose tissue transcriptomics, and genome-resolved metagenomics, we found that predator-odor exposure significantly affected mouse behavior, gene expression, and gut microbiota. Predator odor-responsive bacteria were associated with the expression of genes involved in anti-microbial defense, and host behavioral responses were predicted by random forest models trained on gut-microbiota profiles. These findings indicate interactions between the gut microbiota and wild-mouse responses to the threat of predation, an ecologically relevant acute stressor.

Journal Article

Order among chaos: High throughput MYCroplanters can distinguish interacting drivers of host infection in a highly stochastic system.

The likelihood that a host will be susceptible to infection is influenced by the interaction of diverse biotic and abiotic factors. As a result, substantial experimental replication and scalability are required to identify the contributions of and interactions between the host, the environment, and biotic factors such as the microbiome. For example, pathogen infection success is known to vary by host genotype, bacterial strain identity and dose, and pathogen dose. Elucidating the interactions between these factors in vivo has been challenging because testing combinations of these variables quickly becomes experimentally intractable. Here, we describe a novel high throughput plant growth system (MYCroplanters) to test how multiple host, non-pathogenic bacteria, and pathogen variables predict host health. Using an Arabidopsis-Pseudomonas host-microbe model, we found that host genotype and bacterial strain order of arrival predict host susceptibility to infection, but pathogen and non-pathogenic bacterial dose can overwhelm these effects. Host susceptibility to infection is therefore driven by complex interactions between multiple factors that can both mask and compensate for each other. However, regardless of host or inoculation conditions, the ratio of pathogen to non-pathogen emerged as a consistent correlate of disease. Our results demonstrate that high-throughput tools like MYCroplanters can isolate interacting drivers of host susceptibility to disease. Increasing the scale at which we can screen drivers of disease, such as microbiome community structure, will facilitate both disease predictions and treatments for medicine and agricultural applications.

Arabidopsis

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

Conserved protein folds underpin the diversification of secreted proteins in a fungal pathogen.

BACKGROUND: During host colonization, fungal plant pathogens secrete effector-like proteins that alter host cell physiology and target plant-associated microbes. However, rapid evolution and low sequence conservation hinder the study and characterization of these proteins. The fungus Zymoseptoria passerinii infects Hordeum spp. and includes lineages adapted to wild and domesticated barley. To date, the evolution of effector-like proteins in this species has not been addressed. RESULTS: We combined multiple structure-based and network analyses to unravel the secretome of Z. passerinii. We first compared AlphaFold2 and ESMFold predictions to establish the baseline for structural analyses. We identified 72 structural clusters in the secretome, revealing fold-level relationships across divergent sequences. We showed that effector-like proteins with predicted host immune-interfering functions evolved from a limited group of protein folds, whereas proteins with predicted antimicrobial properties were distributed across fold groups. Physicochemical comparisons indicate that putative antimicrobial effectors predominantly emerged through amino acid replacements on common effector-enriched scaffolds in Z. passerinii, reconfiguring surface charge and electrostatics. We analyzed intra- and interspecific variation in selected effector-enriched families by comparing Z. passerinii proteins and homologs across the genus Zymoseptoria. We describe constrained core folds, with local variation in loop and surface-exposed regions, consistent with fold stability while still enabling protein diversification. We further report that putative antimicrobial effector homologs are broadly distributed across the genus despite sequence divergence. CONCLUSIONS: The secretome of Z. passerinii is organized around common structural folds that support diverse biological roles, including host manipulation and host-associated microbial interactions. Conserved scaffolds combined with surface and physicochemical variation likely contribute to rapid adaptive evolution of effector-like proteins in Z. passerinii.

Fungal Proteins

Dental wastewater reveals a hidden reservoir of oral bacteriophage diversity.

Bacteriophages (phages) are being explored as alternatives or complements to antibiotics because of their ability to selectively kill bacterial pathogens. However, phages that infect many oral bacteria remain undiscovered. Here, we discovered that dental wastewater harbors previously underexplored phage diversity. Viral particles concentrated from dental wastewater displayed diverse morphologies, including abundant filamentous phage-like particles. Deep long-read metagenomic sequencing of concentrated viral particles generated 7.4 billion bases of sequence data and yielded 255 medium- to high-quality viral operational taxonomic units (vOTUs), including 46 predicted complete genomes. Comparison with large phage databases revealed that 63 of these 255 vOTUs had no detectable match, indicating that extensive sequencing of dental wastewater substantially expands the number of potential bacteriophages associated with the human oral microbiome. Host prediction linked many vOTUs to oral-associated bacterial taxa, including species with few or no previously reported phages, such as Porphyromonas gingivalis, Tannerella forsythia, and Candidatus Saccharibacteria. Functional annotation identified diverse genes associated with antiphage defense systems within a subset of vOTUs, suggesting that oral phages may contribute to the movement of genes encoding bacterial immune functions within the oral microbiome. Together, these findings expand the known oral phageome and show that dental wastewater contains a largely untapped diversity of phages.IMPORTANCEThe human oral cavity contains a diverse microbial community, but the bacteriophages (phages) that infect many oral bacteria remain poorly characterized. This gap limits our understanding of how phages shape oral microbial communities. Here, we show that dental wastewater is an underexplored source of oral phage diversity. Deep long-read metagenomic sequencing revealed 255 medium- to high-quality phage operational taxonomic units, many of which are not present in existing oral phage databases. These genomes include predicted phages of periodontal disease-associated bacteria and other oral taxa with few or no known phages. Dental wastewater therefore expands the known human oral phageome and reveals candidate phages linked to bacteria associated with oral health and disease.

Bacteriophages

Dynamics of gut bacteriophage in diversity outbred mice studied over lifespan and during extreme caloric restriction.

BACKGROUND: The majority of bacteria in the vertebrate gut harbor integrated bacterial viruses ("bacteriophages" or "phages"; integrated phage are termed "prophages"). To probe phage replication strategies in the mammalian gut microbiome, we investigated phage activity in a large longitudinal study of diversity outbred mice (913 animals) undergoing extreme dietary restriction with detailed phenotypic characterization across lifespan. RESULTS: We assembled 54,119 candidate DNA viral genomes from 2997 longitudinal metagenomes, forming 6462 viral operational taxonomic units (vOTUs). Over 85% of vOTUs annotated as novel. Viruses annotated predominantly as prophages in the Caudoviricetes class. We detected no eukaryotic DNA viruses, and none of the strictly lytic Crassvirales order that is abundant in human gut. The most prevalent phages had the widest predicted host ranges. The relative abundance of most phages was highly correlated to that of their inferred host bacteria, suggesting quiescent prophages dominate viral metagenomes, consistent with "piggyback-the-winner" dynamics. After accounting for close phage-bacterial covariation, we did identify a subset of phages changing in relative abundance and prevalence relative to their hosts in response to dietary restriction and aging. In particular, phages with larger genomes become less common in diets with restricted calories, potentially reflecting a higher fitness cost to their host. Generalist phages were enriched for a gene encoding a single-strand DNA binding protein which is reportedly involved in DNA repair and protection from nucleases encoded by host cells. Lytic phages became more common with aging, and we observed a reduction in phage richness with age, both findings previously observed in human cohorts. CONCLUSION: These studies enrich our understanding of DNA phage dynamics in gut while emphasizing the predominance of "piggyback-the-winner" strategies.

Animals

NRG-P0074 Viral Sample RU1 from Unclassified Mosigvirus Genomic Characterization and Host Range Analysis.

BACKGROUND: Machine learning models for phage-host range prediction and design require comprehensive training data on phage genomes and host ranges to predict phage-host interactions effectively. MATERIALS AND METHODS: This study characterizes phage sample NRG-P0074 viral sample RU1 from unclassified Mosigvirus, originally isolated by the Betty Kutter. The complete genome of NRG-P0074 was sequenced, annotated, and analyzed using various bioinformatic tools. Host range analysis was conducted using the Escherichia coli Reference (ECOR) Library and nine Escherichia coli (E. coli) K12 strains (Keio Knockout Collection) with single nonessential gene deletions. RESULTS: The genome of NRG-P0074 spans 168,357 base pairs with a guanine-cytosine (GC) content of 37.5%. NRG-P0074 exhibited permissiveness in 15.28% of the ECOR isolates and all 9 Keio knockout strains. Comparative genomic analysis revealed that NRG-P0074 is closely related to E. coli phage a20. Its genome is comprised of 270 coding sequences, 153 known genes, 16 terminators, 3 ribosomal-binding sites, 0 tRNAs, and 117 hypothetical proteins. CONCLUSIONS: This research provides valuable data for developing machine learning models to predict phage-host interactions, aiding the development of targeted phage therapies against antibiotic-resistant bacteria.

ECOR Library

Host-aware Identification of Intrinsic Gene Expression Biopart Parameters using Combinatorial Libraries.

Model-based design in synthetic biology is limited because bioparts are typically characterised by relative metrics that vary across genetic and physiological contexts. To address this, we introduce a host-aware framework for quantitatively characterising bioparts in combinatorial libraries of plasmid-based constitutive expression constructs. The approach integrates a digital twin of Escherichia coli, conditioned on measured growth rate, with model-in-the-loop parameter identification to separate biopart-associated properties from host-dependent effects. Using structured combinatorial libraries, we identify mechanistically interpretable, transferable parameters for plasmid origins, promoters and ribosome binding sites. In particular, we define an intrinsic translation initiation capacity that captures the dominant RBS-associated contribution to translation while context-dependent expression emerges from host physiology and local sequence context. The resulting parameterisation accurately predicts protein synthesis across physiological conditions, supports incremental library expansion, and reveals localised failures of modularity, providing a scalable foundation for predictive host-aware design in synthetic biology.

Escherichia coli

Functional capacities drive recruitment of bacteria into plant root microbiota.

Root-associated microbiomes are shaped by the plant, yet vary across environments and hosts, challenging prediction and engineering. Here, to uncover principles of bacterial selection at the root-soil interface, we applied a systems-level approach using reconstitution studies with communities of isolates from Arabidopsis, barley and Lotus grown in soil. Functional divergence among the microbiota of the host plants reflected distinct strategies: in Arabidopsis and barley, recruitment was primarily shaped by inoculum, while Lotus root environment favoured fewer, functionally diverse isolates, akin to a 'Swiss army knife' strategy. Despite taxonomic variability, root microbiomes encoded overlapping functions. Across major taxa, isolates with broad but distinct functional repertoires within their families were consistently more abundant. Using a genome-to-function framework that is function centric, taxonomically inclusive and host-context aware, we identified 266 functions enriched across all root microbiomes. This functional backbone emerged as a core signature of plant-associated bacteria, providing a solid foundation for microbiome engineering in agriculture.

Journal Article

The role of mobile genetic elements in adaptation of the microbiota to the dynamic human gut ecosystem.

The human intestinal microbiota is a dynamic ecosystem shaped by extensive horizontal gene transfer, particularly in individuals from industrialized populations. In this review, we discuss recent advances in our understanding of how mobile genetic elements (MGEs) contribute to microbial ecology and evolution in this diverse community, focusing on MGEs carrying fitness-conferring genes. Bacteroidales species can colonize individuals for decades and serve as major hubs for MGE exchange. Most MGEs are highly variable across individuals and geographies. Occasionally, conserved MGEs can spread across geography and lifestyles. Functional characterizations of MGEs reveal their roles in antibiotic resistance, interbacterial antagonism, biofilm formation, immune evasion, and nutrient acquisition, among others. Substantive progress in our understanding of MGEs in the gut microbiome offers promising avenues for therapeutic microbiome interventions. However, major challenges remain in functional prediction, host-MGE linkage, and experimental characterization.

Humans

Interkingdom remodeling of the intestinal bacteriome and virome during Toxoplasma gondii infection in rats.

Toxoplasma gondii infection is associated with intestinal microbiome disruption, but its effects on genome-resolved bacterial populations, the gut virome, and bacteriome-virome relationships remain poorly understood. Using previously generated shotgun metagenomic datasets from 36 intestinal samples collected from 18 Sprague-Dawley rats across control, acute, and chronic infection groups, we reconstructed 294 quality-filtered, non-redundant bacterial metagenome-assembled genomes (MAGs) and identified 899 medium-to-high-quality viral operational taxonomic units (vOTUs) from assembled metagenomic contigs. Infection was associated with reduced bacterial richness in the small intestine during both acute and chronic stages and lower Shannon diversity during chronic infection. In contrast, large-intestinal α-diversity remained stable despite significant compositional reorganization. Taxonomic changes included increased Lactobacillus intestinalis, Limosilactobacillus reuteri, and Prevotella sp900547005, together with decreased Rothia sp002492045 and Akkermansia muciniphila. Functional profiling revealed region- and stage-specific changes in predicted bacterial metabolic potential, including reduced energy-related pathways and carbohydrate-active enzyme abundance. The virome also showed significant compositional changes in both intestinal regions. Quimbyviridae and Podoviridae_crAss-like viruses decreased in the small intestine during chronic infection, while Quimbyviridae, Flandersviridae, and Podoviridae_crAss-like viruses showed stage-specific decreases in the large intestine. Predicted bacterial hosts were assigned to 48.39% of vOTUs, with Lachnospiraceae and Ruminococcaceae being the most frequently linked families. Trans-kingdom networks further revealed region-specific positive and negative abundance correlations between bacterial and viral taxa. These findings extend previous microbiota-metabolome observations by integrating genome-resolved bacteriome analysis with contig-based virome profiling, providing a foundation for future mechanistic studies of toxoplasmosis-associated microbiome remodeling.

Gut virome

Bacteria and phage consortia modulate cecal SCFA production and host metabolism to enhance feed efficiency in ducks.

BACKGROUND: The gut microbiota influences poultry health, nutrition, feed efficiency (FE), and overall productivity. However, the relationship between gut microbes, including bacteria and phages, and FE in ducks remains underexplored. To address this, we integrated cecal 16S amplicon, metagenome, microbiota-derived short-chain fatty acids (SCFAs) profiling, liver transcriptome, and serum metabolome data to illustrate the contribution of the gut microbiome (bacteria and viruses) to duck FE. RESULTS: We reconstructed viral genomes and prokaryotic metagenome-assembled genomes (MAGs) and annotated their genes using comprehensive databases. Prokaryotic hosts of viruses were also predicted to understand virus-host dynamics within the gut ecosystem. Our results revealed that high-FE ducks have higher concentration of propionate and butyrate in cecum compared with low-FE ducks. The metagenome sequencing revealed distinct cecal microbiota profiles between two groups, with increased relative abundance of representative SCFA producers, especially Paraprevotella sp905215575 and Bacteroides sp944322345, and enhanced SCFA-biosynthesis pathways in high-FE ducks. Virome genome assembly identified two phages encoding auxiliary metabolic genes (AMGs) involved in pyruvate metabolism, enhancing nutrient availability for host bacteria to produce SCFAs (e.g., temperate phage-encoded pyruvate phosphate dikinase) or exploiting host central metabolic pathways for viral replication (e.g., lytic phage-encoded formate C-acetyltransferase). Furthermore, these representative SCFA-producing bacteria and phage consortia were associated with serum metabolites (including L-histidine and 4-hydroxydecanedioylcarnitine) linked to duck FE. CONCLUSION: Collectively, these findings provide novel insights into the gut microbial factors regulating FE in ducks, offering potential strategies to optimize poultry nutrition and productivity. Video Abstract.

Animals

No evidence of fine-scale local adaptation of winter moths to variable tree phenology.

Spatial variation in plant phenology can impose strong selective pressures on herbivorous insects whose fitness relies on synchrony with host plants, promoting local adaptation to host timing. Winter moths (Operophtera brumata) have been shown to synchronize egg hatching with host budburst, but whether this reflects local adaptation remains unclear. We used three complementary approaches to assess small-scale local adaptation of winter moths to oak phenology in Wytham Woods, UK, a 385-hectare woodland with repeatable variation in individual oak budburst phenology. We experimentally investigated whether host tree phenology predicts hatch timing using common gardens across multiple temperatures, evaluated fitness benefits of synchrony using translocations, and assessed population structure and gene-environment associations using whole-genome sequencing. We found no support for local adaptation to individual trees. Common garden experiments revealed systematic differences in hatch timing which were unrelated to host budburst, while translocations indicated no fitness consequences of asynchrony. Genetic analyses showed no detectable population structure or association with budburst timing. Local adaptation to host phenology therefore appears not to arise on individual trees but may instead occur at broader spatial scales. Understanding the scale of local adaptation is essential for predicting how insect-plant synchrony will respond to environmental change across heterogeneous landscapes.

Animals

Surface architecture of the bacterial envelope determines phage adsorption route in pathogenic Escherichia coli O157:H7.

UNLABELLED: The outermost surface layers of Gram-negative bacteria determine phage access to terminal receptors, yet their genetic basis has been mapped almost exclusively in laboratory strains that lack them. Here we apply genome-wide RB-TnSeq fitness profiling to four Escherichia coli O157:H7 strains from distinct phylogenetic clades sharing the O157 O-antigen, using 38 phages with terminal receptors previously mapped in E. coli K-12 strain. RB-TnSeq fitness landscapes across all four pathogenic backgrounds were mostly similar, and dominated by surface-associated loci, including the gfc-etk group 4 capsule operon, O-antigen biosynthesis genes, LPS core assembly genes and outer membrane proteins. Disruption of gfc-etk abolished infection in 11 genetically diverse myoviruses, establishing the O-antigen capsule as a widespread required primary recognition substrate. O-antigen loci generated two classes of fitness score patterns. For 10 phages, disruption increased infectivity, indicating it is a barrier to receptor access; for 3 others, disruption abolished infectivity, demonstrating it can also be a primary recognition substrate. Outer membrane protein receptor identity was conserved across laboratory and pathogenic backgrounds, with the same proteins recognized in both K-12 and O157:H7, while glycan layer state determines whether these receptors are reached. These results demonstrate that outer surface glycan layers can act as primary and optional recognition substrates for phage infection, or as physical barriers preventing terminal receptor access. Extending the ability to probe phage-targeted receptors beyond outer membrane proteins provides a framework for incorporating glycan layer state into predictive models of phage-host interactions. IMPORTANCE: Bacteriophage-based interventions for controlling Escherichia coli O157:H7, a major foodborne pathogen responsible for tens of thousands of illnesses annually in the United States, require a mechanistic understanding of the factors governing strain-level susceptibility. Predictive frameworks developed in laboratory model strains lacking O-antigen and capsular polysaccharides can map the terminal protein receptors that phages bind, but are currently limited in their ability to determine whether those receptors are accessible in pathogenic isolates carrying full outer surface complexity. This study provides the first genome-scale, functional genetic map of phage susceptibility determinants in O157:H7 and demonstrates that the state of the outer surface layers, specifically the O-antigen and the gfc-etk capsule, determines whether phages can reach conserved terminal receptors. This finding explains differences in phage susceptibility between strains sharing nearly identical gene content, and identifies the molecular layers that must be characterized to predict phage host interaction in pathogenic E. coli backgrounds.

Journal Article

Transposable elements create distinct genomic niches for effector evolution among Magnaporthe oryzae lineages.

BACKGROUND: Plant-pathogen interactions are characterized by evolutionary arms races. At the molecular level, fungal effectors can target important plant functions, while plants evolve to improve effector recognition. Rapid evolution in genes encoding effectors can be facilitated by transposable elements (TEs). In Magnaporthe oryzae, the causal agent of blast disease in several cereals and grasses, TEs play important roles in chromosomal evolution as well as the gain or loss of effector genes in host specialized lineages. However, a global understanding of TE dynamics driving effector evolution at population scale and across lineages is lacking. RESULTS: Here, we focus on 16 AVR effector loci assessed across a global sampling of 11 reference genomes and 447 newly generated draft genome assemblies from publicly available short-read sequencing data across all major M. oryzae lineages and outgroups. We classified each effector based on evidence for duplication, deletion and translocation processes among lineages. Next, we determined AVR gain and loss dynamics across lineages allowing for a broad categorization of effector dynamics. Each AVR was integrated in a distinct genomic niche determined by the TE activity profile contributing to the diversification at the locus. We quantified TE contributions to effector niches and found that TE identity helped diversify AVR loci. We used the large genomic dataset to recapitulate the evolution of the rice blast AVR1-CO39 locus. CONCLUSIONS: Taken together, our work demonstrates how TE dynamics are an integral component of M. oryzae effector evolution, likely facilitating escape from host recognition. In-depth tracking of effector loci is a valuable tool to predict the durability of host resistance.

Ascomycota