PubMed HealthSearch

SEARCH · PubMed Health

Results for “Consensus Sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Enhanced transgene expression from single-stranded AAV vectors in human cells in vitro and in murine hepatocytes in vivo.

We identified that distal 10 nucleotides in the D-sequence in AAV2 inverted terminal repeat (ITR) share partial sequence homology to 1/2 binding site of glucocorticoid receptor-binding element (GRE). Here, we describe that (1) purified GR binds to AAV2 D-sequence, and the D-sequence competes with GR binding to its cognate binding site; (2) dexamethasone-mediated activation of GR pathway significantly increases the transduction efficiency of AAV2 vectors in human cells; (3) human osteosarcoma cells, U2OS, which lack expression of GR, are poorly transduced by AAV2 vectors, but stable transfection with a GR expression plasmid restores vector-mediated transgene expression; (4) replacement of the distal 10 nucleotides in the D-sequence of the AAV2 ITR with a full-length GRE consensus sequence significantly enhances transgene expression in human cells in vitro and in murine hepatocytes in vivo; and (5) none of the ITRs in AAV1, AAV3, AAV4, AAV5, and AAV6 genomes contains the GRE 1/2 binding site, and insertion of a full-length GRE consensus sequence in the AAV6-ITR also significantly enhances transgene expression from AAV6 vectors, both in vitro and in vivo. These novel vectors, termed generation Y AAV vectors, which are serotype, transgene, or promoter agnostic, should be useful in human gene therapy.

AAV vectors

Genomic characterization and phylogenetic placement of Matryoshka RNA virus 1 associated with Plasmodium vivax malaria in Africa.

Plasmodium vivax is a major cause of human malaria. It harbours Matryoshka RNA virus 1 (MaRNAV-1), a bi-segmented positive-sense RNA virus. MaRNAV-1 was first described in P. vivax and is now recognized as part of a wider group of Matryoshka viruses. These viruses also infect other haemosporidian parasites such as Leucocytozoon and Haemoproteus. The presence of MaRNAV-1 in African-origin human P. vivax, however, has not been clearly established. This study investigated whether MaRNAV-1 is present in public African-origin P. vivax transcriptomic datasets. Any viral sequences recovered were characterized using comparative genomic and phylogenetic analyses. A secondary in silico analysis targeted African-origin P. vivax RNA-seq runs from public repositories. Although the search covered Africa, only Ethiopian datasets could be confidently identified, retrieved and compiled at the time. After quality control and screening for MaRNAV-1 RNA-dependent RNA polymerase (RdRp) signals, three high-confidence runs were selected for further analysis. Reference-guided reconstruction, ORF prediction, blast-based validation and RdRp phylogenetic analysis were performed. MaRNAV-1 was identified in three Ethiopian P. vivax malaria transcriptomes. This was supported by strong segment-level mapping, near-complete coverage, high mean depth and minimal low-depth masking. The recovered genomes showed the expected bisegmented organization of MaRNAV-1. Segment I was highly conserved and encoded the canonical RdRp in all three consensus sequences. Segment II showed the conserved organization of two overlapping hypothetical ORFs in all three consensus sequences. Blast analyses confirmed close similarity to MaRNAV-1 reference sequences. Phylogenetic inference grouped the Ethiopian sequences within the broader P. vivax-associated MaRNAV-1 lineage, alongside other recognized MaRNAV lineages distinct from more divergent narna-like viruses. These findings provide genomic evidence for MaRNAV-1 in publicly available African-origin P. vivax transcriptomic datasets and add to the emerging evidence for the virus in the African malaria context.

MaRNAV

Phosphoproteomic Analysis of Cortical Tissue from Mice Lacking Both CaMKIIα and CaMKIIβ Identifies Novel In Vivo Substrates.

Ca2+/calmodulin-dependent protein kinase II (CaMKII) plays a critical role in calcium signaling. Several studies have shown that mice with single Camk2a or Camk2b gene knockouts are viable, yet exhibit distinct phenotypes, whereas the double knockout of both genes is lethal. These findings indicate that each gene can have distinct roles and that they also partially compensate for each other in yet unknown essential brain functions. In order to provide insight into potential novel CaMKII functions, we performed parallel phosphoproteomic analyses on nonstimulated cortex tissues from inducible Camk2a and Camk2b double knockout (Camk2af/f;Camk2bf/f;CAG-CreESR) mice and from wild type mice. A total of 5622 phosphorylated peptides derived from 2080 proteins were identified. Phosphorylation at serine/threonine residues in 130 proteins was downregulated in the double knockout mice, including residues in 113 proteins that have not previously been identified as potential CaMKII substrates. Comparison of amino acid sequences surrounding the downregulated phosphorylation residues provided new insights into the CaMKII-substrate consensus sequences in vivo. This data set provides an important resource for future studies examining novel roles for CaMKII in the brain.

Animals

Direct RNA nanopore sequencing of full-length coronavirus genomes provides novel insights into structural variants and enables modification analysis.

Sequence analyses of RNA virus genomes remain challenging owing to the exceptional genetic plasticity of these viruses. Because of high mutation and recombination rates, genome replication by viral RNA-dependent RNA polymerases leads to populations of closely related viruses, so-called "quasispecies." Standard (short-read) sequencing technologies are ill-suited to reconstruct large numbers of full-length haplotypes of (1) RNA virus genomes and (2) subgenome-length (sg) RNAs composed of noncontiguous genome regions. Here, we used a full-length, direct RNA sequencing (DRS) approach based on nanopores to characterize viral RNAs produced in cells infected with a human coronavirus. By using DRS, we were able to map the longest (∼26-kb) contiguous read to the viral reference genome. By combining Illumina and Oxford Nanopore sequencing, we reconstructed a highly accurate consensus sequence of the human coronavirus (HCoV)-229E genome (27.3 kb). Furthermore, by using long reads that did not require an assembly step, we were able to identify, in infected cells, diverse and novel HCoV-229E sg RNAs that remain to be characterized. Also, the DRS approach, which circumvents reverse transcription and amplification of RNA, allowed us to detect methylation sites in viral RNAs. Our work paves the way for haplotype-based analyses of viral quasispecies by showing the feasibility of intra-sample haplotype separation. Even though several technical challenges remain to be addressed to exploit the potential of the nanopore technology fully, our work illustrates that DRS may significantly advance genomic studies of complex virus populations, including predictions on long-range interactions in individual full-length viral RNA haplotypes.

Cell Line

Oral bacteriome in pediatric patients with malignancies prior to chemotherapy: a pilot study using full-length 16S rRNA sequencing.

OBJECTIVE: To characterize the composition, diversity, and ecological features of the oral bacteriome in pediatric patients with malignancies prior to chemotherapy initiation. METHODS: In this prospective pilot study,supragingival plaque samples were collected from 10 pediatric cancer patients prior to the initiation of chemotherapy. Bacterial genomic DNA was extracted from each sample, and the full-length 16S rRNA gene was amplified and sequenced on the PacBio Sequel II platform using circular consensus sequencing (CCS). Raw CCS reads were quality-filtered and denoised into amplicon sequence variants (ASVs) using DADA2, and taxonomic assignment was performed against the SILVA 138 reference database. Alpha diversity was assessed using the Chao1, Shannon, Simpson, and Faith's phylogenetic diversity (PD whole tree) indices, while beta diversity was evaluated through principal coordinate analysis (PCoA), and non-metric multidimensional scaling (NMDS). Microbial co-occurrence networks were constructed to characterize bacterial interactions, and functional potential was predicted using PICRUSt2, and BugBase. RESULTS: A total of 614,473 high-quality CCS reads were generated, yielding 1,697 ASVs. Alpha diversity analysis revealed substantial inter-individual variation in microbial richness and diversity among the pediatric cancer patients. The bacterial community was dominated by the phyla Firmicutes, Proteobacteria, Bacteroidota, Actinobacteriota. At the genus level, Streptococcus, Prevotella, Neisseria, and Haemophilus were the most abundant taxa. Beta diversity analysis revealed distinct clustering patterns, indicating highly individualized microbial profiles. Co-occurrence network analysis identified several keystone taxa and potential pathogenic associations within the supragingival plaque community. Functional prediction indicated that the dominant metabolic pathways were related to amino acid metabolism, carbohydrate metabolism, and membrane transport. CONCLUSION: These preliminary findings reveal a taxonomically diverse, highly individualized pre-chemotherapy oral bacteriome, providing foundational baseline profiles to guide future longitudinal investigations of chemotherapy-induced dysbiosis and personalized interventions.

Humans

First nationwide full-genome characterisation of human-derived Andes virus in Chile: a retrospective genomic epidemiology study.

BACKGROUND: Andes virus (ANDV) is the only hantavirus known to transmit between humans and causes hantavirus cardiopulmonary syndrome in Chile and Argentina. In Chile, ANDV genomic diversity remains incompletely characterised. This study aimed to characterise the genetic diversity, geographical structure, and molecular signatures of ANDV using human clinical samples collected over a 13-year period (2011-24). METHODS: We conducted a retrospective genomic epidemiology study of ANDV infections in Chile. Clinical samples from patients with confirmed ANDV, collected between March 9, 2011, and June 27, 2024, were analysed and sequenced. Clinical and epidemiological data were obtained from diagnostic laboratories and surveillance programmes. Consensus sequences for the S, M, and L segments were generated, and genetic clustering and divergence were assessed using phylogenetic inference and variant calling. FINDINGS: We analysed clinical samples from 58 infected individuals and identified two major genomic variants of ANDV with distinct geographical distributions, defined by regionally structured patterns of nucleotide and amino acid substitutions across the S, M, and L segments: ANDV Chi-North (central Chile) and ANDV-South (southern Chile). No consistent clustering by clinical severity was observed, and no recurrent non-synonymous substitutions were uniquely associated with severe disease. Substitutions previously associated with person-to-person transmission in outbreaks in Argentina were not consistently observed in Chilean sequences, including in four person-to-person transmission cases. Although some substitutions described in ANDV-like viruses were present in the Chi-North lineage, this lineage remained phylogenetically distinct and geographically restricted to central Chile. INTERPRETATION: To our knowledge, this study provides the first nationwide genomic characterisation of human-derived ANDV in Chile. The identification of geographically structured variants indicates that ANDV diversity in Chile is driven by regional diversification rather than clinical outcome. The absence of consistent amino acid signatures associated with disease severity or person-to-person transmission suggests that these phenotypes are unlikely to be explained by viral genetic variation alone. These findings refine current understanding of ANDV evolution and highlight the need for continued integrated genomic surveillance in endemic regions. FUNDING: Agencia Nacional de Investigación y Desarrollo de Chile and National Institutes of Health.

Humans

Cooperation of transposable elements to endow global networks of initiators of hybrid assembly pathways of endogenous multiprotein complexes.

Mechanisms governing initiation steps of the assembly of endogenous multi-protein complexes (EMC) remain incompletely understood. Here, multiple lines of observations are reported describing the function-aligned initiation sequence of hybrid assembly pathways (HAP) of EMC. The first step of HAP-guided chain reactions of protein-protein interactions (PPI) of EMC assemblies constitutes the creation of cell type-specific pools of hetero and homo dimers. The molecular anatomy of HAP was elucidated by defining qualitative and quantitative characteristics of protein binding to a compendium of 200,393 distinct genomic regulatory elements (GRE), including 49,667 sequences representing control sets of genomic loci as well as 150,726 GRE of different evolutionary origins. The consensus sequence of HAP actions consists of: a) Initiation on genomic DNA of the formation of metastable hetero- and homodimers of EMCs' protein constituents; b) Release of dimers from DNA templates for delivery to the EMC assembly compartments; c) Assembly of defined EMC by sequential on demand addition of proteins to preformed dimers serving as attractors of EMC-specific ensembles of monomers. Chromosome-naïve DNA scaffolds facilitating creation of intracellular dimer pools engage networks of ~700 transcription factors (TFs), 534 of which manifest region-specific patterns of significantly enriched expression in 1358 brain regions. HAP initiators appear to operate within nucleosome-depleted islands of transposable elements (TE) - derived sequences within heterochromatin. PPI assembly lines of EMCs operate in 2 concurrent modes: TF-TF PPI cascade and PPI HUB protein cascade. Regardless of the number of DNA-bound initiator TFs (ranging from one to 716 TFs), both modes of operations reached the equilibrium at the PPI constituents saturation levels of ~245 proteins for TF-TF PPI modes and of ~351 proteins for PPI HUB protein modes. Distinct panels of DNA-bound initiator TFs and proteins of PPI cascade ensembles are enriched in either defined sets of neuroanatomical structures (TF-TF mode) or among structural-functional constituents of synapses (HUB proteins mode). Thus, these bifurcated cascades appear biologically congruent: TF-TF constituents map to transcriptional signatures of hundreds of brain regions, whereas HUB constituents map to synaptogenesis and synaptic structures, suggesting the unified logic of genomic functions coordinating region identity and connectivity. Evidence-supported examples of default operations of PPI-guided assemblies of hetero- and homodimers of Yamanaka factors, neurogenesis constituents, and protein components of postsynaptic density of excitatory and inhibitory synaptogenesis are reported with detailed analytical focus on human Claustrum. The foundational set of observations reported in this contribution should facilitate experimental and theoretical explorations of TE-seeded genomic codes for initiators of PPI chain reactions of protein dimerization creating pools of attractors to guide and accelerate the EMC assemblies.

Humans

Whole-genome sequencing of adenovirus 41 directly from wastewater using nested overlapping PCR and MinION.

Human adenovirus F41 (HAdV-F41) is one of the leading causes of children's acute gastroenteritis and was recently linked to an outbreak of severe acute hepatitis of unknown etiology among children during 2021 to 2022. While most evidence is based on clinical data, wastewater-based epidemiology offers a community-level approach to monitoring circulating strains and enhancing outbreak preparedness. In this study, we developed an overlapping amplicon-based whole-genome sequencing approach to directly detect HAdV-F41 from archived wastewater samples, using nested PCR with 13 primer sets. Archived wastewater samples were collected between 2021 and 2022 from three treatment plants in Seattle, USA. The viral load ranged from 1.2 × 103 to 8.4 × 103 genome copies per liter. The Oxford Nanopore platform was used for whole-genome sequencing. Complete or partial (>84%) HAdV-F41 genomes were recovered from wastewater samples, with mean coverage depths ranging from 10³ to 10⁵. The consensus sequences showed more than 99% similarity to reference genomes in the NCBI database. The phylogenetic analysis revealed that 2 sequences clustered within lineage 2a and 11 within lineage 2b, reflecting that at least two sub-lineages were circulating in the community at that time. Our results demonstrate that the overlapping amplicon-based whole-genome sequencing approach using the Oxford Nanopore platform reliably recovers HAdV-F41 genomes from wastewater. This method offers high-resolution genomic surveillance of circulating, clinically relevant HAdV-F41, supporting wastewater-based epidemiology as a valuable tool for detecting emerging variants and strengthening the early warning system for future disease outbreaks.IMPORTANCEHuman adenovirus F41 is a primary cause of childhood gastroenteritis and has been linked to recent outbreaks of severe acute hepatitis in children, yet community-level genomic surveillance of this virus remains limited. This study shows that wastewater can be used to recover nearly complete HAdV-F41 genomes through a targeted overlapping-amplicon sequencing strategy on the Oxford Nanopore platform. By applying this method to archived wastewater samples, we detected the simultaneous circulation of multiple viral lineages in a large city. These findings extend wastewater-based epidemiology beyond SARS-CoV-2 and emphasize its importance for monitoring clinically significant enteric viruses. The method described here offers a scalable tool for tracking viral evolution in communities and enhancing early warning systems for future outbreaks.

Wastewater

Complete genome characterization, phylogenetic analysis, and capsid P2 variation of a goose astrovirus genotype 2 isolate from Guizhou, China.

Goose astrovirus genotype 2 (GAstV-2) is associated with gout and renal disease in goslings, but its occurrence in Guizhou Province remains poorly documented. We isolated a GAstV-2 strain, designated GZJP2024, from goslings with visceral gout on a farm in Jinping County, Guizhou Province, China. PCR detected GAstV-2 but not goose parvovirus, goose reovirus, Tembusu virus, fowl adenovirus, or goose astrovirus genotype 1. Serial passage in goose embryos produced mortality and hemorrhagic lesions during the third passage. Whole-genome sequencing yielded a 7,251-nt genome containing three overlapping open reading frames (ORF1a, ORF1b, and ORF2). Sequence identity and phylogenetic analyses assigned GZJP2024 to the GAstV-2 lineage. ORF1b was the most conserved coding region, whereas ORF2 was more variable. Comparison with consensus sequences from representative GAstV-2 strains identified five amino acid substitutions in ORF1a and ten in ORF2. Four ORF2 substitutions (E456D, L540Q, S608T, and A614T) occurred in the capsid P2 domain and overlapped or neighbored predicted B-cell epitope-rich regions. Template-based mapping placed E456D, S608T, and A614T on exposed regions of a spike-like capsid structure. These findings document a GAstV-2 isolate from a gout-affected goose farm in Guizhou and provide sequence data for future regional surveillance.

Capsid P2

Detecting SARS-CoV-2 cryptic lineages using publicly available whole genome wastewater sequencing data.

Beginning in early 2021, unique and highly divergent lineages of SARS-CoV-2 were sporadically found in wastewater sewersheds using a sequencing strategy focused on amplifying the most rapidly evolving region of SARS-CoV-2, the receptor binding domain (RBD). Because these RBD sequences did not match known circulating strains and their source was not known, we termed them "cryptic lineages". To date, more than 20 cryptic lineages have been identified using the RBD-focused sequencing strategy. Here, we identified and characterized additional cryptic lineages from SARS-CoV-2 wastewater sequences submitted to NCBI's Sequence Read Archives (SRA). Wastewater sequence datasets were screened for individual sequence reads that contained combinations of mutations frequently found in cryptic lineages but not contemporary circulating lineages. Using this method, we identified 18 cryptic lineages that appeared in multiple (2-81) samples from the same sewershed, including 12 that were not previously reported. Partial consensus sequences were generated for each cryptic lineage by extracting and mapping sequences containing cryptic-specific mutations. Surprisingly, seven of the mutations that appeared convergently in cryptic lineages were reversions to sequences that were highly conserved in SARS-CoV-2-related enteric bat Sarbecoviruses. The apparent reversion to bat Sarbecovirus sequences is consistent with the notion that SARS-CoV-2 adaptation to replicate efficiently in respiratory tissues preceded the COVID-19 pandemic.

SARS-CoV-2

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

Long-Read Sequencing of the MUC1 VNTR: Genomic Variation, Mutational Landscape, and Its Impact on ADTKD Diagnosis and Progression.

BACKGROUND: ADTKD-MUC1 is caused by frameshift mutations in MUC1 gene that produce a frameshifted protein (MUC1fs) toxic to kidney cells. The gene's variable number of tandem repeats (VNTR), with high GC content, makes it largely inaccessible to standard sequencing. As a result, both the reference sequence and natural variation in this region remain poorly defined, complicating mutation detection and data interpretation. Standard methods also fail to pinpoint the exact VNTR unit affected, limiting insight into mutation mechanisms and genotype-phenotype correlations. METHODS: We employed Single Molecule, Real-Time (SMRT) sequencing and characterized the genomic sequence of MUC1 in 300 individuals including 279 individuals from 143 families suspected of having ADTKD-MUC1. We compared these results to those obtained using the CLIA-approved mass spectrometry-based probe extension (PE) assay, which specifically detect the most prevalent 59dupC mutation. We correlated the structural features of the MUC1 VNTR with the rate of kidney function decline in affected individuals. RESULTS: We identified MUC1 consensus sequences for 205 unique VNTR alleles, with 9 distinct types of frameshift mutations present on 52 distinct mutated VNTR alleles. MUC1 frameshift mutations were identified in 71 of 143 families (50%) with suspected ADTKD, comprising 135 genetically affected individuals (48%). The SMRT assay exhibited complete concordance and revealed that the PE assay is capable of detecting frameshift mutations in approximately 85% of affected families. The constellation of VNTR structures supports a genotype-progression model, in which fast progressors exhibit a significantly lower number of repeat units on the wild-type allele and a higher number of repeats on the mutation-bearing allele, including an increased number of frameshifted repeat units. CONCLUSIONS: SMRT sequencing outperforms current diagnostic methods for ADTKD-MUC1 and reveals the prognostic value of VNTR structures. Although their contribution to disease progression is modest (~6% variance explained), it remains biologically and clinically meaningful.

Autosomal Dominant Tubulointerstitial Kidney Disea

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

SARS-CoV-2

Determining genotype and antimicrobial resistance of Salmonella Typhi in environmental samples by amplicon sequencing.

BACKGROUND: Estimates of the burden of typhoid fever due to Salmonella enterica serovar Typhi (S. Typhi) rely on data from clinical surveillance, which is rarely done in low income settings and is also limited by the poor sensitivity of the assays used and the reliance on health seeking by patients. Environmental surveillance for S. Typhi shed by symptomatic and asymptomatic individuals in wastewater offers a sensitive surveillance tool that could help to inform burden estimates. Sequencing S. Typhi direct from wastewater concentrates has the potential to identify circulating genotypes and associated antimicrobial resistance (AMR) genes, supporting public health interventions such as vaccination and antimicrobial usage. METHODOLOGY AND PRINCIPAL FINDINGS: We designed a multiplex targeted amplicon sequencing protocol for genotyping and determining AMR in S. Typhi from wastewater samples, targeting SNPs that identify genotypes of interest and both chromosomal and plasmid-borne AMR. PCR products were sequenced using the Oxford Nanopore Technologies (ONT) MinION, and genotypes and AMR identified using the GenoTyphi program. We tested this approach on samples from south India from both hospital outflow and wastewater collected from the community. All samples tested were suspected to be positive for S. Typhi following quantitative PCR for ttr, tviB, and staG gene targets. Out of 110 samples tested we were able to determine a genotype and/or AMR for 8. All samples that gave a genotype call suggested a genotype consistent with those found in clinical cases in India during the same time period and produced consensus sequences that clustered with S. Typhi when included in a phylogenetic tree. CONCLUSIONS: In this study, we provide proof of concept data for amplicon sequencing of S. Typhi in wastewater which with further optimisation could be used to complement clinical surveillance data or provide data on S. Typhi presence in the absence of clinical surveillance. This information can inform public health interventions, and the concept could be applied to other pathogens of interest for genotyping from environmental surveillance samples.

Salmonella typhi

The Arabidopsis TIRome informs the design of artificial TIR (Toll/interleukin-1 receptor) domain proteins.

The TIR (Toll/interleukin-1 receptor) domain is an ancient protein module that functions in immune and cell death responses across the Tree of Life. TIR domains encoded by plants and prokaryotes function as enzymes to produce diverse small molecule immune signals. Plant genomes can encode hundreds of TIR-domain containing proteins-many of which confer important agricultural disease resistance as TIR-NLR (nucleotide-binding, leucine-rich repeat) immune receptors. Despite their importance, how natural variation influences TIR enzymatic output and immunity-associated cell death is largely unexplored. We assayed a complete collection of the TIR domains of Arabidopsis thaliana Col-0 (the "AtTIRome") to explore variation in TIR metabolite production and cell death signaling. Roughly half of the AtTIRome triggered cell death in transient assays. Artificial TIR proteins designed based on consensus sequences of the AtTIRome's cell death phenotypic classes revealed polymorphisms controlling variation in TIR cell death elicitation and metabolite production. Structure-function analyses of artificial TIRs revealed that natural variation in the "BB-loop", a flexible region overlying the catalytic pocket, determines differences in function across Arabidopsis TIR-containing proteins. We further demonstrate that artificial TIRs are functional on an NLR chassis and that BB-loop variation can tune the activity of a natural TIR-NLR protein. These findings shed light on the diversity of TIR outputs and reveal methods to design and engineer TIR-based immune receptors.

Arabidopsis

First chromosome-level genome assembly of the colonial chordate model Botryllus schlosseri (Tunicata).

BACKGROUND: Botryllus schlosseri (Tunicata) is a colonial, laboratory model tunicate recognized for its remarkable developmental diversity, its regenerative abilities, and its peculiar genetically determined allorecognition system governed by a polymorphic locus controlling chimerism and cell parasitism. RESULTS: We report the first chromosome-level genome assembly of B. schlosseri subclade A1. By integrating long and short reads with Hi-C scaffolding, we produced both a phased diploid genome assembly and a conventional collapsed consensus sequence of 533 Mb. Of this total length, 96% belonged to 16 chromosome-scale scaffolds, with a BUSCO completeness score of 91.4%. We then compared our assembly with other high-quality tunicate genomes, revealing some synteny conservation but also extensive genomic rearrangements and a general loss of colinearity. CONCLUSIONS: The chromosome-level resolution of this assembly enhances our understanding of genome organization in colonial modular organisms. Comparative analyses highlight the dynamic nature of tunicate genomes, with conserved macrosynteny yet extensive microsyntenic rearrangements and scrambling, underscoring their rapid evolutionary trajectory. This high-quality genome assembly provides a valuable resource for exploring the unique biological features of colonial chordates, including their exceptional regenerative abilities and complex allorecognition system.

Animals

Functional minigenome system reveals polymerase features of swine orthopneumovirus.

Swine orthopneumovirus (SOV), a recently identified porcine pneumovirus, has been detected in pig farms worldwide; however, its pathogenicity and molecular biology remain poorly understood. To facilitate the study of SOV replication and transcription, we developed a functional minigenome system based on consensus sequences from multiple strains of SOV and related pneumoviruses. Here, we constructed and optimized this system in BSRT7/5 cells, revealing that the RNA-dependent RNA polymerase (RdRp) activity depends on a conserved protein phosphatase 1 (PP1) binding site within the phosphoprotein P, as a single F131A substitution markedly reduced polymerase function. Additionally, we identified and characterized the M2-1 binding site on P, which is essential for viral transcription. These findings provide new insights into SOV polymerase complex requirements and establish a foundation for reverse genetics approaches to rescue infectious viruses, advancing our understanding of SOV biology and its potential role in porcine respiratory disease.IMPORTANCERecently, a newly identified porcine pneumovirus, swine orthopneumovirus (SOV), was detected in pig farms in different countries. Although detected mainly in sick animals, this virus has not been isolated yet and its pathogenicity remains to be determined. We started by setting up a minigenome system with a view to develop reverse genetics and rescue infectious virions. This minigenome system was used to study the functioning of the SOV RNA polymerase and compared it with RSV. Although some similarities exist between SOV and RSV, the RdRp of RSV cannot rescue the SOV minigenome. SOV seems to belong to another genus/genogroup of pneumoviruses, which includes PVM and the canine pneumovirus. Our functional minigenome paves the way for reverse genetics of SOV and determination of its pathogenicity in different host species.

Swine Diseases

Transcription Start Regions in PTU-intergenic regions drive cell cycle-dependent transcriptional activation events in Leishmania donovani.

Leishmania displays an unconventional mode of transcription, with long clusters of genes being transcribed polycistronically from Transcription Start Regions (TSRs), being processed into monocistronic units prior to translation. It has long been believed that transcription is constitutive: failure to identify consensus sequences across TSRs (except a GT-rich motif supporting transcription in Trypanosoma brucei) and absence of canonical eukaryotic transcription factors led to the conclusion that regulation is primarily post-transcriptional, with epigenetics playing a role in triggering transcription initiation. This study stems from our previous findings identifying a few genes to be activated in a cell cycle-dependent manner. Using nuclear run-on assays to analyze nascent transcripts of two chromosomes, chromosomes 2 and 14, we find that while most genes are constitutively transcribed, a subset of genes gets activated at specific cell cycle stages. Reporter assays reveal that this transcriptional activation is driven by the regions immediately upstream of the genes. Sequence analyses of these TSRs lying in polycistronic intergenic regions (PIRs) uncovered a 10-mer GT-rich motif, in synchrony with earlier findings in T. brucei identifying a GT-rich motif at bidirectional TSRs. We also identify a second 25-mer motif at these TSRs, and deletion analyses find this motif to be critical for regulating gene expression. The findings of this study reveal that transcriptional events in these unicellular parasites are more complex than believed thus far: not all transcriptional events are constitutive, polycistronic transcription is not the only mode of transcription, and cis-acting sequence elements regulate at least some transcriptional events in these parasites.IMPORTANCEEndemic to 90 countries, Leishmania parasites cause a spectrum of diseases called Leishmaniases. No vaccines for human use are available to date, and the drugs currently used to treat the disease are expensive, have toxic side effects, and have complex administration regimens, with emerging drug resistance compounding problems. Researchers continue to investigate Leishmania cellular processes, with the hope of uncovering new therapeutic target sites. Gene regulation in these parasites is unusual, being modulated by various mechanisms, including epigenetic modifications, gene dosage, and post-transcriptional processing. Transcription is typically polycistronic and constitutive, initiating from Transcription Start Regions (TSRs) lying upstream of the first gene in the polycistronic transcription unit (PTU). The work presented here reveals that a subset of genes is transcribed monocistronically in a cell cycle-dependent manner from Transcription Start Regions lying in the PTU-intergenic regions (PIRs), underscoring the complexities of gene regulation in these parasites.

Leishmania donovani