PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “INDEL Mutation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals↗

Analysis of retrotransposon structural diversity uncovers properties and propensities in angiosperm genome evolution.

Analysis of LTR retrotransposon structures in five diploid angiosperm genomes uncovered very different relative levels of different types of genomic diversity. All species exhibited recent LTR retrotransposon mobility and also high rates of DNA removal by unequal homologous recombination and illegitimate recombination. The larger plant genomes contained many LTR retrotransposon families with >10,000 copies per haploid genome, whereas the smaller genomes contained few or no LTR retrotransposon families with >1,000 copies, suggesting that this differential potential for retroelement amplification is a primary factor in angiosperm genome size variation. The average ratios of transition to transversion mutations (Ts/Tv) in diverging LTRs were >1.5 for each species studied, suggesting that these elements are mostly 5-methylated at cytosines in an epigenetically silenced state. However, the diploid wheat Triticum monococcum and barley have unusually low Ts/Tv values (respectively, 1.9 and 1.6) compared with maize (3.9), medicago (3.6), and lotus (2.5), suggesting that this silencing is less complete in the two Triticeae. Such characteristics as the ratios of point mutations to indels (insertions and deletions) and the relative efficiencies of DNA removal by unequal homologous recombination compared with illegitimate recombination were highly variable between species. These latter variations did not correlate with genome size or phylogenetic relatedness, indicating that they frequently change during the evolutionary descent of plant lineages. In sum, the results indicate that the different sizes, contents, and structures of angiosperm genomes are outcomes of the same suite of mechanistic processes, but acting with different relative efficiencies in different plant lineages.

Animals↗

NanoFilter: enhancing phasing performance by utilizing highly consistent INDELs and SNVs in nanopore sequencing.

MOTIVATION: Nanopore sequencing data offer longer reads compared to other technologies, which is beneficial for phasing and genome assembly. INDELs provide valuable haplotype information and have significant potential to improve phasing performance. However, accurately identifying INDELs with variant callers is challenging, and incorporating INDELs into phasing remains a complex task. To address these issues, we developed NanoFilter, a novel filtering strategy designed to filter out INDELs that contain wrong phasing information based on their consistency. RESULTS: Our assessment using Nanopore R10 simplex data shows that filtering out low-consistency INDELs increases their precision from 88.3% to 98.8%, nearly matching the precision of SNVs. In the phasing results of Margin, incorporating these filtered INDELs leads to a 12.77% increase in N50 length and fewer switch errors. Furthermore, we found that SNVs filtered by NanoFilter will enhance assembly performance. When NanoFilter is integrated into the HapDup assembly pipeline, NanoFilter reduces the Hamming error rate and increases N50 length by 7.8%. AVAILABILITY AND IMPLEMENTATION: NanoFilter is available at https://github.com/Chenshanming-repo/NanoFilter (DOI: 10.5281/zenodo.16777826) and HapDup-NanoFilter is available at https://github.com/Chenshanming-repo/HapDup-NanoFilter (DOI: 10.5281/zenodo.16777890).

Nanopore Sequencing↗

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software↗

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA↗

Assessing the readiness of Oxford Nanopore sequencing for clinical genomics applications.

Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.

Humans↗

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans↗

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms↗

[Relationship of I/D polymorphism of angiotensin converting enzyme gene with hypertension in Xinjiang Kazakh isolated group].

OBJECTIVE: To investigate whether the insertion/deletion(I/D) polymorphism in the angiotensin converting enzyme(ACE) gene is associated with essential hypertension in Xinjiang Kazakh isolated population. METHODS: The study covered 201 hypertensives and 151 normotensive controls in Xinjiang Barlikun Kazakh population. The I/D polymorphism of ACE gene was determined by polymerase chain reaction. RESULTS: The frequencies of D and I in the hypertensive group (0.44 and 0.56, respectively) were not significantly different from the controls(0.39 and 0.61, respectively, P=0.16). The frequencies of ACE genotypes of DD, ID, and II were 0.18, 0.52, 0.30 in hypertensives respectively and 0.17, 0.43, 0.40 in control group respectively. There was no significant difference in genotypes between hypertensive group and normotensive group (P=0.14). CONCLUSION: The results suggested that the I/D polymorphism of ACE gene might not be associated with hypertension in the Kazakh population of Xinjiang Barlikun area.

Asian People↗

Meta-analysis of indels causing human genetic disease: mechanisms of mutagenesis and the role of local DNA sequence complexity.

A relatively rare type of mutation causing human genetic disease is the indel, a complex lesion that appears to represent a combination of micro-deletion and micro-insertion. In the absence of meta-analytical studies of indels, the mutational mechanisms underlying indel formation remain unclear. Data from the Human Gene Mutation Database (HGMD) were therefore used to compare and contrast 211 different indels underlying genetic disease in an attempt to deduce the processes responsible for their genesis. Each indel was treated as if it were the result of a two-step insertion/deletion process and was assessed in the context of 10 base-pairs DNA sequence flanking the lesion on either side. Several indel hotspots were noted and a GTAAGT motif was found to be significantly over-represented in the vicinity of the indels studied. Previously postulated mechanisms underlying micro-deletions and micro-insertions were initially explored in terms of local DNA sequence regularity as measured by its complexity. The change in complexity consequent to a mutation was found to be indicative of the type of repeat sequence involved in mediating the event, thereby providing clues as to the underlying mutational mechanism. Complexity analysis was then employed to examine the possible intermediates through which each indel could have occurred and to propose likely mechanisms and pathways for indel generation on an individual basis. Manual analysis served to confirm that the majority of indels (>90%) are explicable in terms of a two-step process involving established mutational mechanisms. Indels equivalent to double base-pair substitutions (22% of the total) were found to be mechanistically indistinguishable from the remainder and may therefore be regarded as a special type of indel. The observed correspondence between changes in local DNA sequence complexity and the involvement of specific mutational mechanisms in the insertion/deletion process, and the ability of generated models to account for both the number and identity of the bases deleted and/or inserted, makes this approach invaluable not only for the analysis of indel formation, but also for the study of other types of complex lesion.

Base Sequence↗

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases↗

IMPACT OF FLUORESCENT DYES ON MUTATIONS IN NEXT GENERATION SEQUENCING LIBRARY GENERATION.

DNA labelling fluorescent dyes such as ethidium bromide have long been considered to be highly mutagenic during DNA replication. While recent studies have pushed back on this narrative, the intercalative nature of these dyes continues to raise the possibility that these dyes can induce mutations. The iconPCR instrument by n6tec uses fluorescent dyes to measure amplification in real time and to adjust cycling conditions. However, since this use of qPCR is preparative and not analytical, mutations introduced by fluorescent dyes would be propagated into the sequencing reaction. To address the impact of these dyes on downstream analyses, we have performed routine mutation calling as well as mutational signature analysis on samples amplified using the iconPCR in the presence of either SYBR or EvaGreen. Sequence analysis revealed very minimal impacts of dyes on the reactions, largely within the noise regimen with only subtle changes in mutation rates seen. Mutational signature analysis was unable to identify any key signatures assignable to the dyes in either substitutions or indel domains. The mutational impact of intercalating dyes during fluorescence-guided amplification is therefore minimal and can be disregarded in all but the most sensitive NGS applications.

Fluorescent Dyes↗

Deletion bias in avian introns over evolutionary timescales.

The role that introns play in the function and evolution of nuclear genomes is not fully understood. Recent models of intron evolution suggest that selection and drift may interact to maintain introns in multicellular organisms. In addition, deletion mutations are more likely to become fixed than insertion mutations. Examination of indel substitutions over macroevolutionary timescales in pigeons and doves (Aves: Columbiformes) reveals that deletion substitutions outnumber insertion substitutions by over six times. The length of indel events is variable.

Animals↗

An integrative method for accurate comparative genome mapping.

We present MAGIC, an integrative and accurate method for comparative genome mapping. Our method consists of two phases: preprocessing for identifying "maximal similar segments," and mapping for clustering and classifying these segments. MAGIC's main novelty lies in its biologically intuitive clustering approach, which aims towards both calculating reorder-free segments and identifying orthologous segments. In the process, MAGIC efficiently handles ambiguities resulting from duplications that occurred before the speciation of the considered organisms from their most recent common ancestor. We demonstrate both MAGIC's robustness and scalability: the former is asserted with respect to its initial input and with respect to its parameters' values. The latter is asserted by applying MAGIC to distantly related organisms and to large genomes. We compare MAGIC to other comparative mapping methods and provide detailed analysis of the differences between them. Our improvements allow a comprehensive study of the diversity of genetic repertoires resulting from large-scale mutations, such as indels and duplications, including explicitly transposable and phagic elements. The strength of our method is demonstrated by detailed statistics computed for each type of these large-scale mutations. MAGIC enabled us to conduct a comprehensive analysis of the different forces shaping prokaryotic genomes from different clades, and to quantify the importance of novel gene content introduced by horizontal gene transfer relative to gene duplication in bacterial genome evolution. We use these results to investigate the breakpoint distribution in several prokaryotic genomes.

Algorithms↗

Patterns and relative rates of nucleotide and insertion/deletion evolution at six chloroplast intergenic regions in new world species of the Lecythidaceae.

Insertions and deletions (indels) in chloroplast noncoding regions are common genetic markers to estimate population structure and gene flow, although relatively little is known about indel evolution among recently diverged lineages such as within plant families. Because indel events tend to occur nonrandomly along DNA sequences, recurrent mutations may generate homoplasy for indel haplotypes. This is a potential problem for population studies, because indel haplotypes may be shared among populations after recurrent mutation as well as gene flow. Furthermore, indel haplotypes may differ in fitness and therefore be subject to natural selection detectable as rate heterogeneity among lineages. Such selection could contribute to the spatial patterning of cpDNA haplotypes, greatly complicating the interpretation of cpDNA population structure. This study examined both nucleotide and indel cpDNA variation and divergence at six noncoding regions (psbB-psbH, atpB-rbcL, trnL-trnH, rpl20-5'rps12, trnS-trnG, and trnH-psbA) in 16 individuals from eight species in the Lecythidaceae and a Sapotaceae outgroup. We described patterns of cpDNA changes, assessed the level of indel homoplasy, and tested for rate heterogeneity among lineages and regions. Although regression analysis of branch lengths suggested some degree of indel homoplasy among the most divergent lineages, there was little evidence for indel homoplasy within the Lecythidaceae. Likelihood ratio tests applied to the entire phylogenetic tree revealed a consistent pattern rejecting a molecular clock. Tajima's 1D and 2D tests revealed two taxa with consistent rate heterogeneity, one showing relatively more and one relatively fewer changes than other taxa. In general, nucleotide changes showed more evidence of rate heterogeneity than did indel changes. The rate of evolution was highly variable among the six cpDNA regions examined, with the trnS-trnG and trnH-psbA regions showing as much as 10% and 15% divergence within the Lecythidaceae. Deviations from rate homogeneity in the two taxa were constant across cpDNA regions, consistent with lineage-specific rates of evolution rather than cpDNA region-specific natural selection. There is no evidence that indels are more likely than nucleotide changes to experience homoplasy within the Lecythidaceae. These results support a neutral interpretation of cpDNA indel and nucleotide variation in population studies within species such as Corythophora alta.

Base Sequence↗

Analysis of eriophyid mite rDNA internal transcribed spacer sequences reveals variable simple sequence repeats.

Ribosomal DNA internal transcribed spacers of the eriophyid mites Cecidophyopsis ribis, C. selachodon, C. spicata, C. alpina, C. aurea, C. grossulariae and Phylocoptes gracillis were amplified using PCR, cloned and sequenced. Sequences for the ITS1 of Cecidophyopsids were 92-99% homologous. Cecidophyopsis inter-specific differences were found in seventeen simple sequence repeats (vSSRs), fourteen point mutations and two indels. No intra-specific variation in vSSRs was detected. A hypothetical structure for ITS1 was obtained and vSSRs were mapped onto this. Changes in vSSRs were compensated for by changes in complementary vSSRs or through multiple point mutations. A comparison with vSSRs of other arthropods suggested that the levels of intra-specific variation in Cecidophyopsis mites was less than in organisms which do not use arrhentoky for male determination.

Animals↗

Genomic rearrangements in the CFTR gene: extensive allelic heterogeneity and diverse mutational mechanisms.

Cystic fibrosis (CF) is caused by mutations in the cystic fibrosis transmembrane conductance regulator gene (CFTR/ABCC7). Despite the extensive and enduring efforts of many CF researchers over the past 14 years, up to 30% of disease alleles still remain to be identified in some populations. It has long been suggested that gross genomic rearrangements could account for these unidentified alleles. To date, however, only a few large deletions have been found in the CFTR gene and only three have been fully characterized. Here, we report the first systematic screening of the 27 exons of the CFTR gene for large genomic rearrangements, by means of the quantitative multiplex PCR of short fluorescent fragments (QMPSF). A well-characterized cohort of 39 classical CF patients carrying at least one unidentified allele (after extensive and complete screening of the CFTR gene by both denaturing gradient gel electrophoresis and denaturing high-performance liquid chromatography) participated in this study. Using QMPSF, some 16% of the previously unidentified CF mutant alleles were identified and characterized, including five novel mutations (one large deletion and four indels). The breakpoints of these five mutations were precisely determined, enabling us to explore the underlying mechanisms of mutagenesis. Although non-homologous recombination may be invoked to explain all five complex lesions, each mutation appears to have arisen through a different mechanism. One of the indels was highly unusual in that it involved the insertion of a short 41 bp sequence with partial homology to a retrotranspositionally-competent LINE-1 element. The insertion of this ultra-short LINE-1 element (dubbed a "hyphen element") may constitute a novel type of mutation associated with human genetic disease.

Alleles↗

An indelible lineage marker for Xenopus using a mutated green fluorescent protein.

We describe the use of a DNA construct (named GFP.RN3) encoding green fluorescent protein as a lineage marker for Xenopus embryos. This offers the following advantages over other lineage markers so far used in Xenopus. When injected as synthetic mRNA, its protein emits intense fluorescence in living embryos. It is non-toxic, and the fluorescence does not bleach when viewed under 480 nm light. It is surprisingly stable, being strongly visible up to the feeding tadpole stage (5 days), and in some tissues for several weeks after mRNA injection. We also describe a construct that encodes a blue fluorescent protein. We exemplify the use of this GFP.RN3 construct for marking the lineage of individual blastomeres at the 32- to 64-cell stage, and as a marker for single transplanted blastula cells. Both procedures have revealed that the descendants of one embryonic cell can contribute single muscle cells to nearly all segmental myotomes rather than predominantly to any one myotome. An independent aim of our work has been to follow the fate of cells in which an early regulatory gene has been temporarily overexpressed. For this purpose, we co-injected GFP.RN3 mRNA and mRNA for the early Xenopus gene Eomes, and found that a high concentration of Eomes results in ectopic muscle gene activation in only the injected cells. This marker may therefore be of general value in providing long term identification of those cells in which an early gene with ephemeral expression has been overexpressed.

Animals↗