PubMed HealthSearch

SEARCH · PubMed Health

Results for “Long-read sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics

Evaluating detection of Histophilus somni immunoglobulin-binding protein A DR2 Fic: A species-specific gene target for recombinase polymerase amplification relative to long-read sequencing of respiratory samples from feedlot calves.

Histophilosis is an important cause of morbidity and mortality as well as antimicrobial use in feedlot cattle across North America. Detection of Histophilus somni by culture is challenging, and there is no standardized tool for distinguishing isolates that carry virulence factors most likely to contribute to disease. The DR2 repeat of H. somni-associated virulence factor 'immunoglobulin-binding protein A' (ibpA DR2) harbors a Fic domain that mediates host cell cytotoxicity and is essential for histophilosis. For rapid detection of ibpA DR2 in extracted DNA, we developed a real-time recombinase polymerase amplification (RPA) assay with a runtime of 24&#xa0;min at 39&#xa0;&#xb0;C. DNA from H. somni-RPA-positive respiratory swabs (n&#xa0;=&#xa0;73) was screened for ibpA DR2 using the novel RPA assay and long-read metagenomic sequencing, as well as nanopore whole-genome sequencing (WGS) of H. somni isolated from the same samples. IbpA DR2 was identified in 71% and 70% of tested samples using RPA and WGS, respectively, and in &#x2264;41% of samples using metagenomic sequencing. The likelihood of detection by RPA did not differ (OR 1.1, 95% CI (0.42, 2.9), P&#xa0;>&#xa0;0.99) from WGS; however, agreement between these assays was only fair (&#x3ba;&#xa0;=&#xa0;0.31). Conversely, RPA (OR 3.4, 95% CI (1.6, 8.2)) and WGS (OR 8.0, 95% CI (2.4, 42)) were more likely (P&#xa0;<&#xa0;0.001) to detect ibpA DR2 than metagenomic sequencing, likely reflecting limited coverage of H. somni by metagenomics. This study demonstrated that RPA and long-read WGS detected ibpA DR2 with similar frequencies in extracted DNA and H. somni isolates, respectively. Further testing of non-target isolates confirmed the analytical specificity of ibpA DR2 to H. somni. Further investigation of the diagnostic validity for RPA-based ibpA DR2 detection is required in a larger cohort of field samples, as a rapid screening tool for H. somni most likely to contribute to disease.

Animals

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common &#x223c;4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome

Conference report: the third Bacterial Genome Sequencing Pan-European Network conference.

The third Bacterial Genome Sequencing Pan-European Network conference, held in Engelberg, Switzerland (12-15 January 2026), brought together experts from six European countries to discuss the implementation of bacterial genome sequencing in clinical microbiology and public health. Key themes included regulatory frameworks (In Vitro Diagnostic Regulation, General Data Protection Regulation), standardization, quality control, data sharing, economic evaluation, and the integration of artificial intelligence and long-read sequencing into diagnostic workflows. Across presentations, panel discussions, and workshops, participants emphasized that successful implementation of genome sequencing requires more than technical capacity: it depends on robust validation, sustainable funding, interoperable data standards, ethical governance, and interdisciplinary collaboration. The meeting highlighted that sequencing should remain question-driven and clinically meaningful, balancing cost, turnaround time, and public health impact. Overall, the conference reinforced the need for coordinated European efforts to advance responsible, standardized, and sustainable genomic surveillance and diagnostics.

bacterial genome sequencing

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

From regionalization to homogenization: Nationwide metagenomic assessment of priority pathogens and the resistome in Polish hospital wastewater.

Hospital wastewater (HWW) is a critical hotspot for the dissemination of antibiotic resistance genes (ARGs) and pathogens. This study provides the first comprehensive metagenomic characterization of HWW across Poland, analyzing 64 medical facilities across two seasons via Nanopore long-read sequencing (total of 128 HWW samples). The HWW microbiome was mostly dominated by Proteobacteria, Bacteroidota, and Firmicutes. Multivariate analysis confirmed a significant seasonal shift in the resistome. Winter samples exhibited geographic regionalization, with localized hotspots of specific ARGs, including vancomycin resistance (operon van) and carbapenemase genes (blaOXA, blaNDM). Conversely, summer samples showed a significant trend toward nationwide homogenization, characterized by a uniform distribution of ESBL genes (blaTEM, blaCTX-M) and multidrug resistance (MDR) determinants, alongside the persistence of localized clinical hotspots. Klebsiella pneumoniae emerged as a central network hub, particularly in summer, showing strong correlations with ESBLs. Quantitative genomic co-occurrence analysis revealed a functional division within dominant taxa: while environmental species like Acinetobacter johnsonii comprised the general background microbiome, clinical pathogens such as Acinetobacter baumannii served as primary vectors, showing frequent associations with high-risk ARGs. Environmental and opportunistic bacteria, such as Aeromonas spp. and Citrobacter spp., were identified as putative 'bridge hosts' associated with mobile resistance determinants and potentially contributing to HGT. The findings indicate that seasonal factors, such as increased temperature and sub-inhibitory antibiotic concentrations, may contribute to the transition from regionalized to homogenized resistance profiles, demonstrating that background resistome convergence can coexist with point-source clinical outbreaks. This seasonal "blurring" of regional boundaries positions HWW as an active vector for large-scale antimicrobial resistance (AMR) dissemination. These results underscore the urgent need for nationwide metagenomic surveillance and advanced wastewater treatment strategies within the "One Health" framework to mitigate the environmental spread of WHO priority pathogens.

Acinetobacter baumannii

Insights into the regulation of the HOTAIR proximal promoter.

HOTAIR (HOX transcript antisense RNA) is a HOXC-cluster long intervening non-coding RNA (lincRNA) whose cancer relevance is tightly coupled to how its transcription is wired into hormone, hypoxia, inflammatory, and developmental signaling. HOTAIR is known to associate with cancer cell proliferation, motility, tumor invasion, and metastasis. The present mini-review focuses on the regulatory architecture and mechanistic complexity of HOTAIR transcriptional regulation, with emphasis on three organizing principles. First, we consider the impact of promoter choice between a canonical proximal promoter (P1), which supports the 2.2-2.4 kb transcript, and an alternative upstream promoter/TSS (P2), which contributes to context-dependent transcription initiation. Second, we examine the long-distance enhancer-promoter communication between HOTAIR distal enhancer and P1/P2. Third, we summarize the recent epigenetic and epi-transcriptomic mechanisms involved in HOTAIR transcript initiation and elongation. A combination of these events determines isoform-specific transcription to govern cell-type-, context-, and cancer specific modulation of HOTAIR expression that promotes tumor formation and cancer progression. Finally, the review proposes how large-scale RNA datasets, long-read sequencing, and isoform-specific studies can refine our understanding of this versatile lincRNA's regulation.

Humans

A chromosome-level, haplotype-resolved genome assembly for the barn owl, Tyto alba.

Recent advances in long-read sequencing have enabled near telomere-to-telomere (T2T) assemblies across diverse taxa. However, avian genomes remain challenging due to numerous microchromosomes, small, typically < 20Mb, DNA molecules that are gene-, GC-, and repeat-rich. As a consequence, microchromosomes are often missing from genome assemblies. Here, we present a chromosome-level, haplotype-resolved genome assembly for the Western barn owl (Tyto alba). Using a trio-binning strategy with Illumina parental reads combined with PacBio HiFi and Oxford Nanopore Technologies data, we generated two phased contig sets. These were scaffolded into 40 linkage groups using a linkage map. Comparative analyses identified unplaced HiFi scaffolds corresponding to microchromosomes, which we integrated into six additional microchromosomes using long reads information. The two assemblies present 46 chromosomes, matching the karyotype of the species. They exhibit strong synteny between parental haplotypes, except for a &#x223c;38 Mb complex region on chromosome 7 containing nested inversions. This high-quality reference provides a haplotype-resolved and chromosome-level genome for Strigiformes, enabling fine-scale studies of structural variation and avian genome evolution.

Tyto alba

Interplay between the role of DNA methylation in regulating gene expression and TE-silencing in a reptilian methylome.

DNA methylation is a major component of eukaryotic genomes with an important role in the defence against transposable elements, to transcriptionally silence their activity and prevent transposition. DNA methylation also plays a major role in the regulation of gene expression. This dual role can come into conflict, where DNA methylation in gene regulatory regions becomes perturbed due to transposable element transposition, leading to disruption of gene expression. Here, we describe how this conflict is reflected in DNA methylation patterns in the sand lizard genome where there is recent transposable element activity. Using long-read sequencing technology we show that CpG islands in gene transcriptional start sites are typically hypomethylated and associated with higher gene expression. Outside transcriptional start sites, a majority of CpG islands overlapped transposable elements and were associated with hypermethylation, consistent with a host-defence role in suppressing transposition activity. We identify 605 instances where transcriptional start sites were associated with transposable elements (4.3% of all genes). These instances were far rarer in conjunction with a CpG island, when methylation signatures would be in conflict. Transposable elements were found to be closer to and at higher density the more hypermethylated a transcriptional start site was, suggesting strong selection against selfish genetic elements transposing into hypomethylated transcriptional start sites.

CpG islands

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and&#x2009;~&#x2009;2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia

Reference genome of the Californian trapdoor spider Aptostichus stephencolberti Bond 2008 (Araneae: Mygalomorphae: Euctenizidae).

We present a reference genome assembly for the trapdoor spider Aptostichus stephencolberti. This species, described in 2008, is endemic to the highly fragmented coastal dune habitats of Northern California from Monterey to the San Francisco Bay Area. Trapdoor spiders are ideal taxa for landscape scale genomic studies owing to their extreme site fidelity and limited dispersal capabilities; these same characteristics make them prone to extinction. Genomic studies of species like A. stephencolberti can reveal novel areas of endemism and high conservation value that may not be evident in species with wider ranges and greater dispersal capabilities. As part of the California Conservation Genomics Project, we constructed the A. stephencolberti reference genome from high quality long-read sequences, scaffolded with proximity ligation Omni-C data. The primary assembly comprises 551 scaffolds spanning 3.63 Gbp, a scaffold N50 of 62.2 Mbp and BUSCO completeness of 95.6%. We estimate 52 chromosomes yet find no (TTAGG)n telomer repeats. Expanding the telomeric repeat search finds an ancestral loss of the repeat from all spiders. Automated annotation using the NCBI refseq pipeline and RNAseq data from whole adults finds 14,067 genes with a BUSCO annotation completeness of 95.56%. Repeat annotation identified 77% of the genome to be interspersed repeats. This resource, the first for family Euctenizidae will facilitate future study and resulting conservation actions of A. stephencolberti and other Aptostichus sp. populations associated with the rapidly changing California coastal dune ecosystem.

Aptostichus stephencolberti

Development and evaluation of a multiplex PCR-based dual-platform targeted sequencing framework for precise differentiation of lumpy skin disease virus.

BACKGROUND: Lumpy skin disease virus (LSDV) shares over 96% genomic identity with goatpox and sheeppox viruses, presenting severe diagnostic challenges due to cross-reactivity. METHODS: To address this bottleneck, we established a targeted sequencing framework integrating multiplex PCR with short-read and long-read platforms. By sequentially screening target pathogens, identifying low-homology genes, and designing short and gradient long-fragment primer pools, we evaluated these dual-platform panels using highly homologous poxvirus samples. RESULTS: The short-read panel stably detected target viruses at inputs as low as 5.26 &#xd7;101 copies/&#x3bc;L. Under strict alignment criteria, LSDV mapping rates reached 42.91%, suppressing non-target signals to 3.05%. The Nanopore-Targeted Sequencing (NTS) long-amplicon strategy successfully eliminated homologous interference. By applying length-dependent diagnostic thresholds (&#x2265; 100 reads for short amplicons; &#x2265; 50 reads for long amplicons), precise species-level identification was achieved, maintaining near-zero cross-reads (0-5) in ultra-long regions. Crucially, the field-deployable NTS workflow enabled complete detection in approximately 4 h. CONCLUSION: This complementary strategy seamlessly meets both laboratory demands for high-sensitivity enrichment and frontline requirements for rapid typing, providing a reliable tool for LSDV surveillance, mutation tracking, and outbreak control.

Capripoxvirus differentiation

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Genetic Screening of Colombian Patients With Early-Onset Parkinson Disease.

BACKGROUND AND OBJECTIVES: Early-onset Parkinson disease (EOPD), defined as symptom onset before 50 years of age, accounts for approximately 10% of patients and is suggested to have a greater genetic component than typical late-onset forms of the disease. Recessive variants in PRKN, PINK1, and DJ-1, are the most common genetic cause of EOPD, however, most studies are in patients of white ancestry. This study aims to analyze genetic variants in PRKN, PINK1, and DJ-1 in Colombian patients to help address the gap in EOPD genetic research of South American populations. METHODS: We analyzed 43 unrelated patients with EOPD using Sanger sequencing for the PRKN, PINK1, and DJ-1 genes and employed multiplex ligation-dependent probe amplification to detect copy number variants. Additionally, long-read whole-genome sequencing was conducted on 3 unresolved patients with age at onset before 30 years of age (long-read sequencing [LRS] patient A-C). RESULTS: We identified known pathogenic single-nucleotide variants and copy number variants in the PRKN gene accounting for 2 patients' disease (4.6% of patients). We observed 2 pathogenic variants in PRKN (c.155delA; p.N52Mfs*29 and c.1083+1G>A) in patient 1, who reported an age at onset of 16 years. We further detected a homozygous duplication of PRKN exons 5-6 in an additional patient, age at onset of 18 years. DISCUSSION: Our study helps characterize genetic contributors to EOPD in Colombian patients, demonstrating genetic forms (PRKN, PINK1, and DJ-1) are rare. Our results highlight a need to include diverse populations in research to improve genetic understanding of disease.

Journal Article

Serotypic and Genomic Diversity of Vibrio anguillarum in Rainbow Trout Farms in Turkey: Implications for Vibriosis Control and Vaccine Candidate Selection.

Outbreaks of vibriosis caused by Vibrio anguillarum are a persistent constraint on rainbow trout (Oncorhynchus mykiss) aquaculture. However, information on the population structure of field strains in Turkey has been lacking. Here, we report the first systematic serotypic, proteomic, and genomic characterization of 23 V. anguillarum isolates collected over 10&#x2009;years from rainbow trout farms located in six major aquaculture regions of Turkey. Serological analyses based on microagglutination, supported by ELISA characterization of hyperimmune sera, identified a clear predominance of serotype O1, whereas isolate V12 exhibited a non-agglutinating, atypical O-antigen profile. Protein profiling (SDS-PAGE) and immunoblotting showed largely conserved whole-cell protein patterns among the isolates, but distinct immunogenic bands at 14, 18, and 40&#x2009;kDa were detected in isolates V18 and V21. Long-read whole-genome sequencing revealed that most Turkish isolates grouped within the global O1 clade, while V12, V25, and V28 isolates occupied more distant branches. Comparative genomics demonstrated a conserved core virulence gene set (RTX toxins, siderophore and iron-uptake systems, motility and adhesion factors, Type VI secretion system), with strain-dependent variation in accessory loci such as anguibactin and T6SS-I. Experimental infections of rainbow trout demonstrated significant differences in virulence among isolates (p&#x2009;<&#x2009;0.05), with the V18 isolate showing high, the V15 intermediate, and the V12 low-mortality rates. By elucidating the relationship among the serotype, immunogenic protein profiles, virulence gene repertoires, and in&#xa0;vivo pathogenicity, this study provides a comprehensive overview of the antigenic and genomic diversity of Vibrio anguillarum isolates from Turkey. Notably, the identification of V18 and V21 as promising candidate strains for further vaccine evaluation, characterized by high virulence and unique immunogenic features, provides a scientific foundation for the development of serotype-specific vaccination strategies to mitigate vibriosis-associated losses in aquaculture.

Animals

Novel bacterial hosts and mobile genetic structure of tet(X) variants in tetracycline-contaminated aquatic environment uncovered by culture and long-read metagenomics.

Clinically important tigecycline (3rd-generation tetracycline) resistance tet(X) variants were inferred to have evolutionarily originated from environmental bacteria, and have been recognized among environment, human and animals. However, genetic basis for environmental proliferation and dissemination of tet(X) variants remains ambiguous. This study profiled tet(X) variants at gene, contig, isolate, and community levels in environmental community subjected to long-term stepwise increasing oxytetracycline (1st-generation tetracycline) or tigecycline pressure using long-term microcosm experiments, quantitative PCR, bacterial isolation, whole-genome sequencing, and Nanopore-based long-read metagenomics. We confirmed that both oxytetracycline and tigecycline enriched the abundance of tetracycline resistance genes especially oxytetracycline-enriched tet(X3). Unexpectedly diverse bacterial hosts and genetic structure of tet(X)-positive mobile elements in the environment microbiome were identified using bacterial isolation and long-read Nanopore metagenomics. Pseudomonas defluvii was first reported to carry tet(X3) in the chromosome, forming IS26-tet(X3)-res-ISCR2 circular intermediate to transfer between different DNA molecules. Database mining revealed similar mobile segments have prevailed among animal-derived Acinetobacter species. Unlike the widely reported ISCR2-mediated transfer of tet(X6), we identified a novel mobile multidrug transposon TnAs3 where tet(X6) and class 1 integron co-transferred as its passenger region. Mobile tet(X2)-ere(D)-aadS-erm(F)-blaOXA-347 segment was annotated in Runella, and co-occurrences of tet(X2) and ere(D), aadS, blaOXA-347 were also found in Flavobacterium, Arsenicibacter, Chryseobacterium and Pedobacter. Overall, tetracycline-contaminated aquatic microbiome harboured diverse mobile tet(X)-positive segments which have not yet been acquired by clinical pathogens, and thus served as the genetic pool of tet(X) variants together with indigenous bacterial hosts, especially the newly reported Pseudomonas defluvii. Reducing pollution of older-generation tetracyclines would be a proactive way to mitigate environmental evolution and possible clinical effects of tet(X) variants.

Metagenomics

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus