PubMed HealthSearch

Biomedical subjects

Heng Li

Publications and source records attributed to Heng Li.

10 recordsLinked to original sources

Silent carriage of tigecycline- and carbapenem-resistant Klebsiella quasipneumoniae co-harbouring tetX(4) and blaNDM-1 in healthy individuals in China.

OBJECTIVES: To investigate the occurrence and genomic characteristics of tigecycline- and carbapenem-resistant Klebsiella quasipneumoniae isolated from healthy individuals in China. METHODS: During a nationwide screening programme, faecal samples from healthy community individuals were cultured for carbapenem-resistant Enterobacterales. Antimicrobial susceptibility testing, whole-genome sequencing and conjugation assays were performed for two K. quasipneumoniae isolates co-harbouring tet(X4) and blaNDM-1. RESULTS: Two K. quasipneumoniae strains carrying tet(X4) and blaNDM-1 were recovered from healthy individuals without recent hospitalisation or antibiotic exposure. Both isolates showed resistance to tigecycline (>8 μg/mL) and carbapenems (≥4 μg/mL). Genomic analysis demonstrated close clonal relatedness between the isolates (ST6460-1LV, KL151). The blaNDM-1 gene was located on an IncX3 plasmid associated with mobile genetic elements, whereas tet(X4) was carried by an IncX1 plasmid. Conjugation assays confirmed successful transfer of both resistance genes, which frequently co-transferred into recipient strains. CONCLUSIONS: To our knowledge, this is the first report of K. quasipneumoniae co-harbouring tet(X4) and blaNDM-1 in healthy individuals in China. The findings indicate that healthy community populations may serve as a hidden reservoir for last-resort antimicrobial resistance genes and underscore the importance of community-based surveillance within a One Health framework.

Klebsiella quasipneumoniae

Large-scale genomic analysis places Chinese CC398 as a persistent human-associated MSSA lineage apart from the dominant global LA-MRSA clade.

Staphylococcus aureus clonal complex (CC)398 has emerged as a dominant livestock-associated methicillin-resistant S. aureus (LA-MRSA) lineage worldwide; however, its evolutionary trajectory and regional diversification remain incompletely understood. We developed a core-genome multilocus sequence typing (cgMLST) scheme with hierarchical clustering and applied it to over 30,000 S. aureus genomes, revealing frequent cross-border transmission of CC398. Subsequent time-calibrated phylogenetic analysis placed the most recent common ancestor at 1942 (95% CI: 1939-1945), with the human-to-livestock host jump around 1969 (95% CI: 1968-1972). Chinese CC398 exhibits a distinct trajectory: unlike the LA-MRSA lineages dominating Europe and North America, Chinese isolates are predominantly human-associated methicillin-susceptible S. aureus (HA-MSSA), forming unique East Asia-specific phylogroups (SAP1, SAP2, and AP1-AP3), with distinct resistance and virulence profiles. The LA lineage remains limited in China, with multinational mixed clusters emerging only after 2019. Analysis of global transmission networks revealed a significant correlation between LA-CC398 spread and international trade in fresh swine products, while no such correlation was observed for the human-associated lineage. Beyond the established lineage markers tet(M) and scn, our analysis identified additional differentially distributed genes, including cadC-a chromosomal cadmium resistance regulator-as a novel HA-lineage-enriched gene whose functional role in host adaptation remains to be determined. This study reveals that CC398 followed fundamentally different evolutionary paths in China versus Western countries, challenging a one-size-fits-all model of its dissemination.IMPORTANCEThis study illustrates how large-scale microbial genomics can resolve the evolutionary origins and regional diversification of bacterial pathogens. By applying a novel cgMLST scheme to over 30,000 S. aureus genomes, we show that CC398 followed fundamentally different evolutionary paths in China versus Western countries-challenging the prevailing model of uniform global dissemination-and that livestock-associated MRSA expansion is closely linked to international trade in fresh pork products. These findings highlight the need for integrated surveillance across human, animal, and trade interfaces to anticipate the emergence and spread of zoonotic pathogens.

Staphylococcus aureus

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

FuFiHLA: a tool for full-field HLA typing from long-read data.

MOTIVATION: Allele typing for Human Leukocyte Antigen (HLA) genes has many important clinical applications. Popular short-read typing can only accurately distinguish alleles at the coding sequence level, which potentially limit our understanding of the effect of variants in non-coding region. Long read data has been proved to be useful in typing HLA alleles in full resolution, but only a few tools are publicly available and with significant limitations in practical application. RESULTS: We developed FuFiHLA, a lightweight open-source software, to type HLA alleles. Currently it supports typing alleles of six HLA genes (HLA-A, HLA-B, HLA-C, HLA-DRB1, HLA-DQA1, and HLA-DQB1) from long reads. Evaluation using 233 PacBio HiFi WGS samples from HPRC shows that FuFiHLA achieves 99.6% accuracy in the full field allele typing and QV as 51.8 for consensus allele sequence construction. Additional testing on four Nanopore R10 reads demonstrates slightly reduced accuracy in the fourth field. AVAILABILITY: FuFiHLA is available at https://github.com/jingqing-hu/FuFiHLA under MIT License.

Humans

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

High-resolution metagenome assembly for modern long reads with myloasm.

Long-read metagenome assembly promises complete genomic recovery from microbiomes. However, the complexity of metagenomes poses challenges. We present myloasm, a metagenome assembler for PacBio HiFi and Oxford Nanopore Technologies (ONT) R10.4 long reads. Myloasm uses polymorphic k-mers to construct a high-resolution string graph and then leverages differential abundance for graph simplification. On real-world ONT metagenomes, myloasm assembled three times more complete circular contigs than the next-best assembler. Myloasm can make ONT and HiFi comparable for assembly: for a jointly sequenced gut metagenome, myloasm with ONT assembled more complete circular genomes than any assembler with HiFi. Myloasm recovers previously inaccessible within-species diversity; we recovered six complete Prevotella copri single-contig genomes from a gut metagenome and eight complete TM7 (Saccharibacteria) contigs with > 93% similarity from an oral metagenome. With this improved resolution, we resolved two 98% similar ermF antibiotic resistance genes spreading through distinct strain-specific mobile genetic elements in a human gut.

Journal Article

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article

Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species.

SUMMARY: Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety in methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. Using Mycobacterium tuberculosis, a clonally evolving bacterium with a small accessory genome, as a model system, we systematically evaluated sources of variability in pan-genome estimates. Our analysis revealed that differences in assembly type (short-read versus hybrid), annotation pipeline, and pan-genome software, significantly impact predictions of core and accessory genome size. Extending our analysis to two additional bacterial species, Escherichia coli and Staphylococcus aureus, we observed consistent tool-dependent biases but species-specific patterns in pan-genome variability. Our findings highlight the importance of integrating nucleotide- and protein-level analyses to improve the reliability and reproducibility of pan-genome studies across diverse bacterial populations. AVAILABILITY AND IMPLEMENTATION: Panqc is freely available under an MIT license at https://github.com/maxgmarin/panqc.

Genome, Bacterial

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere