PubMed HealthSearch

Biomedical subjects

Xiao Wang

Publications and source records attributed to Xiao Wang.

14 recordsLinked to original sources

Matched targeted therapy use after broad genomic profiling in advanced Non-Small cell lung cancer.

INTRODUCTION: While broad genomic profiling is increasingly used in advanced NSCLC (aNSCLC), the impact of test results on subsequent guideline-concordant targeted therapy selection remains incompletely understood. METHODS: Using a merged dataset of two large, nationwide, patient-level databases, we identified patients who were diagnosed with aNSCLC 2017-2023, had potentially actionable genomic profiling findings, and initiated systemic therapy. Patients were categorized into actionability subgroups based on contemporaneous regulatory approvals and NCCN guideline recommendations. Within each subgroup, we assessed receipt of guideline-concordant targeted therapy within 24 months, including potential underuse (non-receipt of recommended treatment) and overuse (receipt of non-recommended treatment). RESULTS: Among 6620 patients (67.4% ≥65 years, 54.6% female, 68.9% White), guideline-concordant targeted therapy use varied substantially by actionability category: 2313 (89.6%) of 2582 patients with available 1st-line on-label options received them (10.4% underuse), while 212 (67.3%) of 315 patients with available later-line on-label options received them after 1st-line (32.7% underuse). Among 441 patients with available guideline-concordant off-label options, only 122 (27.7%) received them (72.3% underuse). Conversely, 238 (8.6%) of 3282 patients received matched but guideline-discordant off-label options, representing overuse of ineffective or unestablished therapies. Smoking history, squamous histology, and high PD-L1 expression were associated with lower targeted therapy receipt. CONCLUSIONS: In this cohort study of aNSCLC care, the guideline concordance of targeted therapy use varied by clinical actionability of molecular testing results. Underuse was more common in patients with later-line and off-label targeted therapy options. Patients with classical smoking-related risk profiles were substantially less likely to receive targeted therapy even when actionable alterations were identified.

Journal Article

TargetQC: A targeted quality control framework for clinical genomic testing.

Reliable genetic testing depends on accurate assessment of sequencing quality in clinically relevant genomic regions that directly influence variant interpretation. We developed TargetQC, a flexible quality control framework that supports user-defined gene sets, coverage thresholds, and variant sets for evaluating sequencing performance across exome sequencing (ES) and genome sequencing (GS) platforms. TargetQC assesses exon and gene coverage, identifies regions meeting predefined coverage thresholds, evaluates variant detection accuracy, and measures sequencing quality at pathogenic variant sites. We applied TargetQC to the reference sample NA12878 and 665 clinical samples across five ES platforms and one GS platform. ES-VendorB and ES-VendorE achieved the most complete coverage of OMIM coding regions in NA12878, whereas ES-VendorD and ES-VendorE showed the highest coverage compliance in clinical samples. ES-VendorB and GS demonstrated the highest variant detection accuracy. TargetQC provides a practical framework for benchmarking sequencing performance and informing platform selection in clinical genomics.

exome sequencing

Multidimensional GWAS analyses on longitudinal phenotypes reveal candidate genes regulating multi-stage egg production traits in Wannan yellow chicken.

Egg production performance directly determines the economic viability of indigenous chicken breeding. However, the genetic regulation of multi-stage egg production traits remains difficult to characterize due to their complex and dynamic nature. Here, we integrated a multidimensional GWAS framework, including single-trait GWAS, multi-trait GWAS (MTAG), and longitudinal trajectory-based GWAS (TrajGWAS), to identify stage-specific and shared genetic effects underlying egg production traits in Wannan yellow chickens (WNY). Whole-genome sequencing of 354 WNY hens (10× depth) and quality control yielded 14,253,816 SNPs for analysis. Selective sweep analyses comparing red jungle fowl, commercial layers, and WNY identified a genomic region containing IGF1 under significant selection pressure. Single-trait GWAS identified SNPs 4_57990480 (BMPR1B) and 17_370912 (LOC112531479) associated with egg production across three laying stages (21-30, 31-40, and 21-40 weeks). MTAG further identified loci 8_4336468 (FASLG) and 21_654726 (CHD5) with shared effects across the laying period, whereas TrajGWAS revealed longitudinal associations involving PRKG1 and identified dynamic loci associated with clutch traits, including GRID1. For clutch traits, stage-specific loci were detected for average clutch size (ACS) and maximum clutch size (MCS), including SNP 8_8542036 at 21-30 weeks, PROK1 at 31-40 weeks, and CUL5, ALKBH8 across the entire laying period. These results demonstrate that integrating complementary GWAS strategies improves the resolution of genetic architecture underlying egg production traits by capturing trait-specific, shared, and stage-dependent genetic effects. The identified GWAS loci and selective-sweep candidate regions provide insights into the genetic architecture of egg production traits and breed differentiation.

Egg production

Large-scale whole-genome sequencing reveals the landscape and health implications of de novo mutations.

De novo mutations (DNMs) are an important source of congenital diseases. With delayed parenthood and assisted reproductive technology (ART) use increasing, it is essential to elucidate how these reproductive factors influence DNMs and whether resulting mutations influence offspring health. Here we performed whole-genome sequencing of 24,030 individuals from 7,851 parent-offspring families, identifying 390,924 de novo single-nucleotide variants (dnSNVs). Paternal and maternal aging exhibited distinct mutational patterns, with maternal DNM accumulation accelerating at advanced ages. Increased paternal dnSNVs partially accounted for the association between advanced parental age and shorter gestational duration. Moreover, ART showed age-independent, procedure-specific effects: intracytoplasmic sperm injection (ICSI) and ovarian stimulation were associated with increased paternal and maternal dnSNVs, respectively, and ICSI-associated paternal dnSNVs also partially accounted for the association between ICSI and shorter gestational duration. In vitro embryo manipulation was associated with increased early post-zygotic mosaic mutations, particularly C > A substitutions linked to delayed neurocognitive development at 1 year. Collectively, these findings advance understanding of the determinants and consequences of de novo mutagenesis.

Journal Article

The role of KIAA1467 in breast cancer: insights from pan-cancer and single-cell sequencing analysis.

BACKGROUND: Improving the response rate of single-agent immune checkpoint blockade (ICB) urgently requires the discovery of new therapeutic targets for combinatorial regimens. Analyses of tumor microenvironment (TME)-associated biomarkers have verified that KIAA1467 drives the formation of an immune-excluded, non-inflamed TME in breast cancer (BRCA). This study systematically explores the expression pattern, prognostic value, immune regulatory function, biological effects, and drug resistance relevance of FAM234B (also known as KIAA1467) in BRCA. METHODS: We performed pan-cancer survival analysis using The Cancer Genome Atlas (TCGA) datasets. Multi-omics bioinformatics analyses were conducted to evaluate KIAA1467 expression across malignancies. Single-cell RNA sequencing (scRNA-seq) data from GSE176078 was utilized to localize KIAA1467 expression at the cellular level. Immunohistochemistry and western blot assays validated KIAA1467 expression in BRCA clinical specimens. Correlation analyses were implemented to assess relationships between KIAA1467 expression, clinicopathological features, immune modulators, tumor-infiltrating immune cells, and p53 mutation status. Functional enrichment analysis uncovered relevant signaling pathways. Bioinformatic half maximal inhibitory concentration (IC50) prediction and in vitro cellular experiments were applied to evaluate associations between KIAA1467 and chemotherapeutic drug sensitivity. RESULTS: TCGA pan-cancer survival analysis demonstrated that elevated KIAA1467 expression significantly predicted shortened overall survival in BRCA and multiple other tumor types. KIAA1467 displayed distinct expression patterns across cancers, with prominent upregulation in BRCA. scRNA-seq confirmed enriched KIAA1467 expression within BRCA cells, and its upregulation in BRCA tissues was further verified by immunohistochemistry and western blot. High KIAA1467 expression was positively correlated with advanced tumor grade and lymphatic metastasis. KIAA1467 showed negative correlations with most immune modulators and core immune checkpoint molecules, as well as tumor-infiltrating immune cells in the TME, implying its potential function in tumor immune evasion. Low KIAA1467 expression was tightly linked to p53 mutations. Enrichment analysis indicated participation of KIAA1467 in epithelial-mesenchymal transition, apoptosis and cell cycle arrest. Furthermore, high KIAA1467 expression corresponded to higher estimated IC50 values of cisplatin, gefitinib, paclitaxel and gemcitabine, consistent with reduced chemosensitivity observed in vitro. CONCLUSIONS: This study reveals the multifaceted oncogenic role of KIAA1467 in BRCA. KIAA1467 participates in remodeling an immunosuppressive TME, correlates with malignant progression and chemoresistance, and may serve as a promising candidate target to optimize ICB-based combination therapy for BRCA. These findings offer new perspectives for the clinical treatment and comprehensive management of BRCA.

KIAA1467

Whole-genome evolutionary dynamics of human parainfluenza virus type 3 in Shanghai, China, 2016-2024.

• Fifty whole-genome sequencing revealed co-circulating HPIV-3 C3 sub-lineages C3f and C3a in Shanghai, China. • Whole-genome phylogeny dated the HPIV-3 tMRCA to ∼1925.6 and revealed two post-1990 demographic expansions. • Recombination signals detected in the HN gene and other regions may lead to discordance in partial-gene phylogenies. • The L gene showed the highest variability and harbored the largest number of putative positively selected sites.

Letter

Metabolomic differences in the Ophiura sarsii complex from the Yellow Sea Cold Water Mass and Bering Sea Cold Pool.

Metabolomics provides a functional readout of cellular physiology and can reveal metabolite-level differences associated with environmental and evolutionary contexts. Here, we used GC-MS- and LC-MS-based metabolomics to characterize metabolic profiles of the Ophiura sarsii complex from the Yellow Sea Cold Water Mass (YSCWM) and the Bering Sea Cold Pool (BSCP). This metabolomics analysis identified 398 LC-MS/MS and 87 GC-MS/MS differential metabolites (DEMs). Marked metabolic differences were observed between the two taxa, involving antioxidant-related metabolites, central carbon-related intermediates, osmolyte-associated compounds, and membrane lipid components. O. sarsii vadicola from the YSCWM showed higher levels of glutathione, glucose, citric acid, D-ribulose 5-phosphate, and unsaturated lipid-related metabolites, indicating differences in antioxidant-related and energy-associated metabolic profiles. By contrast, O. sarsii from the BSCP was characterized by higher levels of sugar alcohols, particularly myo-inositol, together with differences in membrane lipid-associated metabolites. These results provide metabolomics-based evidence for metabolite-level physiological differences between two members of the O. sarsii complex sampled from the Yellow Sea Cold Water Mass and the Bering Sea Cold Pool, while the relative contributions of lineage divergence and site-specific environmental variation remain to be tested experimentally.

Metabolomics

Efficient prime editing in vivo and in vitro using lipid nanoparticles.

Prime editing is a versatile clinical genome editing method that enables precise substitutions, small insertions and deletions at specified locations in the genomes of living systems including human cells. Although non-viral lipid nanoparticle (LNP) delivery of RNA in vivo has become a preferred method for gene editing in animals and patients, its application to complex, three-component prime editing systems has yielded low editing efficiencies. Here we developed a systematic prime editing LNP (PE-LNP) optimization platform that addresses key bottlenecks in cargo design that limit editing efficiency. This generalizable workflow yielded PE-LNPs that can achieve 49% average in vivo prime editing in the bulk mouse liver with a single dose of 2 mg kg-1. We applied our workflow to the correction of PAH R408W, a cause of phenylketonuria, in a mouse model and achieved prime editing efficiencies and serum phenylalanine levels anticipated to be curative. We also show that PE-LNPs minimize off-target editing compared with DNA delivery methods, induce only transient elevation of liver enzymes and can be dosed repeatedly to improve editing efficiencies. These PE-LNP systems provide an attractive alternative to viral delivery by offering transient expression that minimizes off-target editing, no observed long-term toxicity and high levels of non-viral in vivo liver prime editing.

Animals

Synthetic community derived from the root core microbes of a desert shrub Caragana korshinskii enhances wheat drought tolerance.

BACKGROUND: Drought, intensified by climate change, poses a mounting threat to global food security by severely constraining crop productivity. While microbial inoculants offer promise for drought tolerance, their poor adaptability remains insufficient for extremely water-deficient environments. Desert plants host unique drought-adapted microbiomes that remain largely unexplored for agricultural applications. RESULTS: Here, we investigated the microbial community of the desert shrub Caragana korshinskii and identified a core set of drought-responsive strains. A synthetic microbial community (SynCom) derived from these strains significantly improved wheat growth under drought stress. Metagenomic analyses revealed that microbial functions related to biofilm formation, quorum sensing, and carbon metabolism were enriched, with Pseudomonas identified as a key functional taxon. Guided by inter-strain interactions in biofilm assembly, we streamlined the consortium into a five-member synthetic community, where quorum-sensing signals promoted community-wide biofilm formation. Community biofilm production improved strain colonization and conferred greater drought tolerance compared to monocultures. In plants, mechanistic investigations indicated that the simplified SynCom inoculation universally upregulated MAPK and jasmonic acid signaling pathways. Furthermore, carbohydrate metabolic pathways such as starch and sucrose metabolism were specifically activated, suggesting a multi-level mechanism underlying SynCom-mediated drought tolerance. CONCLUSIONS: These findings demonstrate that SynCom constructed on the endophytic flora of desert plants can significantly enhance crop drought tolerance. Our work highlights the pivotal role of community biofilm synthesis in facilitating root colonization and activating a multidimensional drought tolerance network in plants. This study not only gives an ecological perspective on desert microbiome adaptations but also offers a strategic framework for developing effective microbial inoculants for arid-region agriculture. Video Abstract.

Caragana

Molecular Biomarker Testing Patterns and Turnaround Time in US Patients With Advanced Non-Small Cell Lung Cancer.

BACKGROUND: Patients with advanced non-small cell lung cancer (aNSCLC) are recommended to undergo molecular testing for targetable genomic alterations. However, as high-throughput methods are increasingly used, long test turnaround time (TAT) may lead to lower receipt of appropriate targeted therapy. Guidelines recommend a 2-week TAT for ALK and EGFR testing, 2 prevalent pathogenic alterations with highly effective targeted therapies. PATIENTS AND METHODS: Using an electronic health record-derived, deidentified database, we conducted a retrospective cohort study of patients with aNSCLC diagnosed between 2011 and 2023 who received testing for ≥1 of 8 molecular markers. We assessed the number of biomarkers tested per patient, testing modality, and TAT (defined as the interval between specimen collection and result date) over time. We also evaluated patients with ALK/EGFR-altered aNSCLC who initiated early nontargeted treatment prior to test result availability, examining associations with TAT and clinical outcomes. RESULTS: The study sample comprised 33,945 patients, with a mean age of 68.2 years; 49.4% were female, 58.3% were White, and 83.4% reported a history of smoking. From 2011 to 2023, the mean number of biomarkers tested per patient (range, 2.0-6.8) and the use of next-generation sequencing (NGS) increased, whereas the mean TAT converged to 3 weeks. Fewer than half of the patients with ALK/EGFR-altered aNSCLC had a TAT of ≤2 weeks, and 1 in 8 initiated early nontargeted treatment. Longer TAT was associated with early nontargeted treatment when analyzed as both a continuous variable (odds ratio, 1.83 per week) and a binary variable (TAT >2 vs ≤2 weeks; odds ratio, 6.02). Early treatment was associated with worse median progression-free survival (9 vs 11 months) in patients with ALK/EGFR-altered aNSCLC. CONCLUSIONS: Biomarker testing and NGS use have increased over time in US patients with aNSCLC. TAT has plateaued and remains longer than recommended in consensus guidelines. Longer TAT was associated with early nontargeted therapy in patients with ALK+/EGFR+ aNSCLC, leading to suboptimal first-line treatment and poorer clinical outcomes.

Humans

Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC.

Three-dimensional nuclear DNA architecture comprises well-studied intra-chromosomal (cis) folding and less characterized inter-chromosomal (trans) interfaces. Current predictive models of 3D genome folding can effectively infer pairwise cis-chromatin interactions from the primary DNA sequence but generally ignore trans contacts. There is an unmet need for robust models of trans-genome organization that provide insights into their underlying principles and functional relevance. We present TwinC, an interpretable convolutional neural network model that reliably predicts trans contacts measurable through proximity ligation-dependent (in situ and intact Hi-C) and independent (DNA SPRITE) genome-wide chromatin conformation assays. . TwinC uses a paired sequence design from replicate Hi-C experiments to learn single base pair relevance in trans interactions across two stretches of DNA. The method achieves high predictive accuracy (AUROC=0.80) on a cross-chromosomal test set from in situ and intact Hi-C experiments in heart tissue. Furthermore, we train TwinC using in situ Hi-C data from the widely used GM12878 cell line and validate its performance with orthogonal DNA SPRITE assay in the same cell type. Mechanistically, the neural network learns the importance of compartments, chromatin accessibility, clustered transcription factor binding and G-quadruplexes in forming trans contacts. In summary, TwinC models and interprets trans genome architecture, shedding light on this poorly understood aspect of gene regulation.

Journal Article

Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq data.

We develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated.

Humans

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article