PubMed HealthSearch

Biomedical subjects

Qian Qin

Publications and source records attributed to Qian Qin.

3 recordsLinked to original sources

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Strong phylogenetic signal from chloroplast genomes of three Barringtonia species provides the first genomic resources for their conservation.

BACKGROUND: The genus Barringtonia (Lecythidaceae) is a vital component of tropical coastal forests and mangrove ecosystems. Among its members, B. racemosa and B. fusicarpa are classified as Endangered and Vulnerable, respectively, due to habitat degradation and anthropogenic pressures, underscoring the urgent need for genetic studies to guide conservation. Chloroplast (cp.) genomes serve as essential resources for phylogenetic reconstruction and conservation genetics. However, the scarcity of cp. genome data for Barringtonia has limited comprehensive evolutionary and conservation-oriented investigations. RESULTS: We assembled and annotated the first complete cp. genomes of B. racemosa, B. fusicarpa, and B. acutangula. All three genomes exhibit the typical quadripartite structure, ranging from 158,959 bp (B. racemosa) to 159,837 bp (B. acutangula), and contain 132 genes (87 protein-coding, 37 tRNA, 8 rRNA) with a GC content of 36.68%-36.86%. Collinearity and IR boundary analyses revealed high structural conservation without large-scale rearrangements. Interspecific sequence-level variations were detected in simple sequence repeats (SSRs) and long repeats. Nucleotide diversity (π) analysis identified highly polymorphic regions, including rpl20 (π = 0.080), rpoA (π = 0.064), rps3 (π = 0.063), and ndhF (π = 0.060), which represent promising molecular markers for population genetics within the genus. Codon-based selection analyses (Ka/Ks) showed that all protein-coding genes are under strong purifying selection (mean Ka/Ks 0.32-0.37), with no evidence of positive selection. Pairwise genetic distances (p-distances) among Barringtonia species are extremely low (mean 0.0046), while distances to the related genus Bertholletia are ~ 6-fold higher, supporting their generic distinction. CONCLUSIONS: Phylogenetic analysis robustly supports Barringtonia as a monophyletic clade (bootstrap = 100%), with B. racemosa and B. fusicarpa forming a sister lineage to B. acutangula. This study provides the first high-quality cp. genome resources for the two threatened Barringtonia species, revealing strong structural and sequence conservation but no direct chloroplast genomic correlates of endangerment. The identified polymorphic regions and repeat markers lay a foundation for future population genetics, phylogeographic studies, and conservation-oriented genetic management of these ecologically important coastal plants.

Genome, Chloroplast

FuFiHLA: a tool for full-field HLA typing from long-read data.

MOTIVATION: Allele typing for Human Leukocyte Antigen (HLA) genes has many important clinical applications. Popular short-read typing can only accurately distinguish alleles at the coding sequence level, which potentially limit our understanding of the effect of variants in non-coding region. Long read data has been proved to be useful in typing HLA alleles in full resolution, but only a few tools are publicly available and with significant limitations in practical application. RESULTS: We developed FuFiHLA, a lightweight open-source software, to type HLA alleles. Currently it supports typing alleles of six HLA genes (HLA-A, HLA-B, HLA-C, HLA-DRB1, HLA-DQA1, and HLA-DQB1) from long reads. Evaluation using 233 PacBio HiFi WGS samples from HPRC shows that FuFiHLA achieves 99.6% accuracy in the full field allele typing and QV as 51.8 for consensus allele sequence construction. Additional testing on four Nanopore R10 reads demonstrates slightly reduced accuracy in the fourth field. AVAILABILITY: FuFiHLA is available at https://github.com/jingqing-hu/FuFiHLA under MIT License.

Humans