PubMed HealthSearch

SEARCH · PubMed Health

Results for “deletion (InDel)”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases

Enhanced exonuclease-Cas9 systems promote multiple nucleotide deletions with higher efficiency and broader targeting scope in plants.

CRISPR-Cas9 is a widely used platform for plant genome editing, but its outcomes are typically dominated by small insertions and deletions (indels). Such limited mutation profiles restrict its utility in functional studies of non-coding RNAs and regulatory elements, such as microRNAs (miRNAs), untranslated regions (UTRs), and promoter sequences, where larger sequence disruptions are often required. Here, we developed enhanced exonuclease-Cas9 platforms, termed multiple nucleotide deletion Cas9 (MND-Cas9) systems, for efficient generation of large deletions in rice. By screening four exonucleases (RecJ, T5, TREX2, and SbcB), we established MND-Cas9v1 systems based on TREX2 or SbcB that produced substantially larger deletions without reducing editing efficiency. Further optimization with an inserted DNA-binding domain (DBD) between Cas9 and exonuclease yielded MND-Cas9v2, which simultaneously enhanced efficiency and deletion size. To expand PAM compatibility, we introduced PAM-relaxed Cas9-NG and SpG variants, generating MND-Cas9-NG/SpGv2 systems with broader targeting scope and superior performance compared to their parental nucleases. Finally, we demonstrated the utility of these systems in two applications: MND-Cas9v2 efficiently knocked out the miRNA gene OsMIR530, producing larger seeds, and generated extended deletions in the 3'UTR of OsGhd2, which upregulated its expression and increased grain size. These results demonstrate that MND-Cas9 systems enable high-efficiency generation of extended deletions and facilitate functional analyses of non-coding RNAs and regulatory sequences. Overall, this work establishes a versatile and expandable exonuclease-Cas9 platform that substantially broadens the mutational spectrum and application potential of CRISPR-Cas9 for plant genome engineering.

CRISPR-Cas Systems

Incorporating indel channels into average-case analysis of seed-chain-extend.

MOTIVATION: Given a sequence s1 of n letters drawn independently and identically (i.i.d.) from an alphabet of size &#x3c3; and a mutated substring s2 of length m<n, we want to recover the mutation history that generated s2 from s1. Many modern sequence aligners for this task use seed-chain-extend with k-mer seeds. Previously, Shaw and Yu showed linear-gap cost chaining can produce a chain with 1-O(1m) recoverability, the proportion of the mutation history that is recovered, in O(mn2.43&#x3b8;&#x2009;log&#x2009;n) expected time for seed-chain-extend (assuming pre-seeded reference), where &#x3b8;<0.206 is the mutation rate under a substitution-only channel and s1 is uniformly random. A gap remains between theory and practice, as real genomes include insertions and deletions (indels). RESULTS: We introduce mathematical machinery to deal with the two new obstacles introduced by indel channels: the dependence of neighbouring anchors and the presence of anchors that are only partially correct. We prove that expected recoverability of an optimal chain is &#x2265;1-O(1m) and expected runtime is O(mn3.15&#xb7;&#x3b8;T&#x2009;log&#x2009;n), given the total mutation rate &#x3b8;T=&#x3b8;i+&#x3b8;d+&#x3b8;s (sum of substitution, insertion, and deletion rates) is &#x3b8;T&#x2264;0.159. We thus narrow (but not close) the gap between theory and practice. AVAILABILITY AND IMPLEMENTATION: https://github.com/Lazarus42/seed_chainer_indels.

INDEL Mutation

Dual-dimensional profiling of host genomic variations and HPV integration in PD-L1-stratified cervical cancer via Oxford Nanopore Technology.

BACKGROUND: The integration of human papillomavirus (HPV) DNA into the host genome is a key step in the development of HPV-associated cervical cancer (CC). However, the genomic characteristics of host genomic variations and HPV integration within the context of programmed death-ligand 1 (PD-L1) expression stratification have not been systematically investigated. METHODS: Whole-genome sequencing was performed using Oxford Nanopore Technology (ONT) on six samples (three from the high PD-L1 expression group and three from the low PD-L1 expression group). The characteristics of host genomic variations under different PD-L1 expression stratifications were explored, including structural variations (SV), copy number variations (CNV), single nucleotide polymorphisms (SNP), and insertion-deletions (Indel). Subsequently, the distribution features of HPV integration sites were analyzed, different integration types were identified, and pathway analysis was conducted. RESULTS: Whole-genome SV analysis revealed that the total number of SVs and the composition of mutation types were similar between the high and low PD-L1 expression groups, with insertions (INS) and deletions (DEL) predominating in both. These variations were primarily enriched in intergenic regions and introns. In the low PD-L1 expression group, integration events were observed at multiple chromosomal loci, with the most frequent integration occurring in the KLF5 gene region on chromosome 13. No frequently integrated loci were identified in the high PD-L1 expression group. Additionally, four distinct HPV integration breakpoint patterns were preliminarily identified and analyzed. CONCLUSION: PD-L1 expression stratification did not significantly alter the overall genomic instability of the host. However, differences were observed in the distribution patterns of HPV integration sites. These findings provide new insights into the genomic heterogeneity of CC under different PD-L1 expression backgrounds and may lay the groundwork for future research exploring stratified immunotherapy based on HPV integration features.

Humans

Dissecting genetic architecture and improving machine learning&#x2011;based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC&#xa0;>&#xa0;0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

Development of recombinant inbred lines and QTL analysis of plant height and fruit shape-related traits in Cucurbita pepo L.

UNLABELLED: Zucchini (Cucurbita pepo subsp. pepo) stands as an economically vital crop in China. In zucchini breeding, plant architectural patterns and fruit morphological characteristics serve as pivotal traits. In this study, we employed quantitative trait locus (QTL) analysis using recombinant inbred lines (RILs) derived from two distinct inbred lines, JinGL (subsp. ovifera) and HM-S2 (subsp. pepo), in conjunction with a high-density genetic map. Our investigation focused on ten QTLs associated with six horticulturally significant traits, including hypocotyl length (HL), plant height (PH), and four fruit-related traits: fruit length (FL), fruit diameter (FD), fruit shape index (FSI), and fruit weight (FW). The QTLs governing HL and PH were mapped to Chr03/LG10 and named qhl3.1 and qph3.1, respectively. The candidate gene Cp4.1LG10g05910/CpDw for qph3.1 was successfully identified. Additionally, three novel QTLs related to fruit size and shape were discovered. Among them, qfsi8.1/qfl8.1, demarcated by Marker238258 and Marker240069 on Chromosome 08/Linkage group 17 (Chr08/LG17), is a new major QTL regulating the fruit shape of zucchini. Through genomic insertion-deletion (InDel) and qRT-PCR analyses, we predicted genes within the qfsi8.1/qfl8.1 candidate interval, uncovering Cp4.1LG17g02030/CpIAA12 and Cp4.1LG17g02010/CpCalB as potential candidate genes. We developed molecular markers tightly linked to qph3.1 and qfl8.1 and validated them in 171 and 224 Cucurbita pepo germplasms, achieving accuracy rates of 96% and 100%, respectively. This study deepens our understanding of the genetic basis of key traits and provides valuable references for molecular breeding in Cucurbita pepo. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s11032-025-01592-y.

Cucurbita pepo

Targeted multiplex gene knockouts in Lemna minor using CRISPR/Cas9.

Lemna minor (commonly known as duckweed) is a fast-growing aquatic plant recognized as a promising green bioreactor for recombinant protein production. Its rapid proliferation, high protein yield, environmental adaptability, and edibility make it highly attractive for biotechnological applications. It is essential to develop and expand genetic tools tailored to this species to maximize these advantages and further unlock its biotechnological potential. A key strategy for achieving this goal is the implementation of advanced genome editing technologies, such as the CRISPR/Cas9 system. Although multiplex CRISPR/Cas9 gene editing has previously been successfully applied in Lemna aequinoctialis, the capability of the endogenous plant tRNA processing system for multiplex editing in L. minor using the polycistronic tRNA-sgRNA (PTG)/Cas9 system has not yet been explored. In this study, a PTG construct was engineered to include four sgRNAs designed to simultaneously target two plant-specific glycosyltransferase genes: &#x3b1;-1,3-fucosyltransferase (FucT) and &#x3b2;-1,2-xylosyltransferase (XylT). As anticipated, the PTG-Cas9 system successfully induced frameshift mutations, characterized by insertions and deletions (indels), in regenerated L. minor plants derived from transformed calli. Validation via PCR and RT-PCR analysis, followed by sequencing of the target loci, confirmed the presence of indels at the target sites. Furthermore, western blot analyses utilizing antibodies specific to XylT and FucT in two homozygous lines (lines 44 and 217) revealed truncated XylT proteins in both lines. Moreover, an in-frame FucT protein was detected in line 217, whereas FucT expression was absent in line 44. This study marked the first successful demonstration of PTG-Cas9 system for multiplex genome editing in L. minor, paving the way for advanced genetic engineering in this species.

CRISPR-Cas Systems

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28&#xd7;, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS&#xa0;&#x2248;&#xa0;1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (H&#x2092;&#xa0;&#x2248;&#xa0;0.71) and nucleotide diversity (&#x3c0;&#xa0;&#x2248;&#xa0;7&#xa0;&#xd7;&#xa0;10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

ONCOLINER: A new solution for monitoring, improving, and harmonizing somatic variant calling across genomic oncology centers.

The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.

Humans

Genomic and genetic dissection underlying seedling drought resilience in oats.

Drought threatens global crop yields, and common oat, a vital nutritional source for food and feed, is particularly constrained in the semi&#x2011;arid regions where it is widely cultivated. Here, we report two high-quality genome assemblies for drought-resilient (Borris37) and drought-sensitive (XymC06) oat accessions with distinct seedling survival rates and genome sizes of 10.92&#x2009;Gb and 10.96&#x2009;Gb, and construct comprehensive landscapes of insertion&#x2011;deletions (InDels) and structural variants (SVs). Integrating population-level genomic, transcriptomic and phenotypic (seedling survival rate), we demonstrate that InDels and SVs underpin divergent drought resilience and identify 52 candidate genes associated with drought resistance whose expression is significantly modulated by these variants. Borris37 accumulates 36 favorable alleles of these genes. An InDel in the AsNF-YB3 promoter enhances binding to AsARF1, upregulating AsNF&#x2011;YB3 under drought, and overexpression of AsNF&#x2011;YB3 reduces ROS accumulation. Our findings provide resources and targets for drought&#x2011;resistance breeding in oat, thereby supporting global food security.

Drought Resistance

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 &#xd7; depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article

DeepSomatic: Accurate somatic small variant discovery for multiple sequencing technologies.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies now offer potential advantages in terms of repeat mapping and variant phasing. We present DeepSomatic, a deep learning method for detecting somatic SNVs and insertions and deletions (indels) from both short-read and long-read data, with modes for whole-genome and exome sequencing, and able to run on tumor-normal, tumor-only, and with FFPE-prepared samples. To help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available a dataset of five matched tumor-normal cell line pairs sequenced with Illumina, PacBio HiFi, and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples and technologies (short-read and long-read), DeepSomatic consistently outperforms existing callers, particularly for indels.

Journal Article

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Identification of Genetic Variations in HLA Region for Kidney Functions.

HLA allelic polymorphisms are associated with a variety of kidney-related traits in different populations. Although Taiwanese-specific genetic variants associated with kidney function have been reported, the role of HLA alleles is unclear. In this study, the association&#xa0;between eGFR and genetic variations in the HLA region was explored in a cohort of 59,448 Taiwanese subjects. A total of 448 genetic variations in the HLA region are significantly associated with eGFR. HLA-C*03 is associated with decreased eGFR, while HLA-DQA1*03, HLA-DQB1*03:03 and HLA-DQB1*03:03:02 demonstrated protective effects. Moreover, amino acid changes on HLA-C and HLA-DRB1 are significantly associated with eGFR. Finally, the eGFR-associated single nucleotide variations (SNVs) and insertions and deletions (indels) are enriched in the HLA-DQB1 gene. After conditional analysis, we identified two independent signals, including rs2853941, rs3830060. In summary, this study highlights the role of HLA-C, HLA-DQA1, HLA-DQB1 and HLA-DRB1 variations in kidney function in the Taiwan Han Chinese population.

Adult

Mutagenic Impact and Evolutionary Influence of Chemoradiotherapy in Hematologic Malignancies.

UNLABELLED: Ionizing radiotherapy (RT) is a widely used treatment strategy for malignancies. In solid tumors, RT-induced double-strand breaks lead to the accumulation of insertion-deletions (indels; ID), and their repair by nonhomologous end joining has been linked to the ID8 mutational signature in surviving cells. However, the extent of RT-induced mutagenesis in hematologic malignancies and its impact on their mutational profiles and interplay with commonly used chemotherapies has not yet been explored. In this study, we interrogated 580 whole-genome sequence (WGS) samples from patients with large B-cell lymphoma, multiple myeloma, and myeloid neoplasms and identified ID8 only in relapsed disease. Yet ID8 was detected after exposure to both RT and mutagenic chemotherapy (i.e., platinum and melphalan). Using WGS of single-cell colonies derived from treated lymphoma cells, we revealed a dose-response relationship between RT and platinum and ID8. Finally, using ID8 as a genomic barcode, we demonstrate that a single RT-surviving cell may seed distant relapse. SIGNIFICANCE: RT and the ID8 indel signature are related, but their genomic impact on hematologic malignancies is unclear. Leveraging WGS, we linked ID8 to both RT and mutagenic chemotherapy and validated that platinum can induce ID8. We used ID8 as a genomic barcode to reveal that RT-resistant cells may seed systemic relapse.

Humans