PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Single-Nucleotide Polymorphism”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

TET2 promotes monocyte inflammatory activation in asthma via ALKBH5-m6A regulation and PI3K signaling: evidence from m6A-SNP and single-cell analyses.

Asthma is a complex inflammatory airway disease with strong genetic determinants, yet the functional relevance of most asthma-associated non-coding variants remains unclear. Emerging evidence suggests that N6-methyladenosine (m6A) modification may serve as a critical epitranscriptomic link between genetic variation and immune regulation. In this study, we aimed to systematically identify functionally relevant m6A-regulated genes in asthma by integrating large-scale GWAS data, m6A-SNP annotations, and single-cell transcriptomic analyses, and to investigate their roles in monocyte-driven airway inflammation. We identified TET2 as a key m6A-regulated gene associated with both asthma and lung function, which was selectively upregulated in monocytes during asthma and accompanied by activation of inflammatory and PI3K signaling pathways. Mechanistic experiments further demonstrated that inflammatory stimulation induced ALKBH5 expression, reduced m6A modification of TET2 mRNA, and increased TET2 protein levels, thereby promoting PI3K/AKT signaling and pro-inflammatory cytokine production, whereas inhibition of TET2 or ALKBH5 attenuated these effects. Collectively, these findings demonstrate that ALKBH5-mediated m6A regulation of TET2 enhances PI3K/AKT signaling in monocytes, thereby promoting inflammatory responses in asthma. Our study establishes TET2 as a key m6A-regulated gene linking genetic susceptibility to monocyte-driven inflammation, and highlights the ALKBH5-m6A-TET2 axis as a potential therapeutic target for modulating aberrant immune responses in asthma.

Humans↗

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans↗

Comparative genomics of carbapenem-resistant Acinetobacter baumannii isolated from pediatric patients in a tertiary care hospital.

Acinetobacter baumannii is a short gram-negative bacillus, notable for its intrinsic multidrug resistance and genomic plasticity, which facilitates the acquisition of additional resistance genes via mobile genetic elements. Due to its increasing carbapenem resistance, the World Health Organization has classified it as a critical priority pathogen. This study performed a comparative genomic analysis of 20 carbapenem-resistant A. baumannii clinical strains isolated from the Hospital Infantil de México Federico Gómez (CRAB-HIMFG), alongside 11 genomes from other Mexican strains. The pangenome was determined to be open, and core genome single-nucleotide polymorphism-based analysis grouped the CRAB-HIMFG strains within CC758/IC5 and CC92/IC2. A novel sequence type (ST) in the MLST-Pasteur scheme was identified, related to STPas156, and in the MLST-Oxford scheme, associated with STOxf758 and STOxf1054. Virulence and resistance genes comprised 0.61% to 2.23% of the pangenome. Oxacillinase genes and efflux pumps primarily mediated carbapenem resistance, while virulence genes included those encoding biofilm and type IV pili. Capsule typing revealed a correlation with established international clones, IC2 and IC5. Plasmids exhibited high diversity, harboring maintenance modules and toxin-antitoxin systems, with the dissemination of resistance genes linked to insertion sequences. Biofilm formation and twitching motility were not always expressed, as they depend on additional environmental factors. Our study shows that comparative genomics is an essential tool to analyze clinically and epidemiologically significant genomes, providing critical insights into gene distribution, genomic architecture, and horizontal gene transfer mechanisms in microbial populations.IMPORTANCEIn recent years, a reported increase in the mortality rate associated with infections caused by A. baumannii, along with a rise in carbapenem resistance, poses a serious clinical challenge. The WHO considered this microorganism critical for research into alternative therapies and epidemiological surveillance. Despite advances in bioinformatics, genomic studies have yet to fully elucidate the structural rearrangements and secretion systems of A. baumannii. This knowledge gap hinders our understanding of its remarkable genomic plasticity and its ability to acquire and spread resistance and virulence genes through horizontal gene transfer.

Acinetobacter baumannii↗

Influence of FAM13A gene polymorphism and serum matrix metalloproteinases 9 and 12 on the phenotypes of chronic obstructive pulmonary disease.

PURPOSE: FAM13A as a susceptibility gene for chronic obstructive pulmonary disease(COPD).Many studies verified that FAM13A involved epithelial‒mesenchymal transition (EMT) via the TGF-β1 pathway, some accompanied by an increase in MMP levels. The present study aimed to explore the disease susceptibility of the FAM13A gene, with clinical phenotypes, and investigate the relationships between FAM13A SNP loci and the serum levels of MMP-9 and MMP-12. PATIENTS AND METHODS: We recurited 497 patients with stable COPD patients and 303 healthy controls. Data on blood tests, pulmonary function, and HRCT imaging were collected. Serum MMP-9 and MMP-12 levels were measured by ELISA. Genomic DNA was extracted, and SNPs in the FAM13A gene were detected using targeted region genotyping chips. Logistic regression analysis was performed to assess the associations between SNP loci and COPD susceptibility. Differences in pulmonary function, haematological indicators, bronchial wall thickness, and emphysema parameters among different genotypes were evaluated. Multiple linear regression analysis was used to explore the relationship between genotypes and serum MMP-12 level. RESULTS: We screened a total of 476 SNPs and identified the rs2869947 polymorphism in the FAM13A gene as significantly associated with an increased risk of COPD,Stratified analyses further revealed that this association was particularly in males and individual with BMI ≥ 24.Serum levels of MMP-9 and MMP-12 were significantly higher in COPD patients compared with healthy controls. Genotype(AA vs.GG) showed no significant association with pulmonary function severity,bronchial wall indices,hematological marker,and serum MMP-9 levels in COPD patients(P > 0.05).Compared with GG genotype, AA genotype presented significantly higher LAA-950% and serum MMP-12 levels (P = 0.049 and P = 0.023). CONCLUSION: Our findings suggest that the FAM13A SNP rs2869947 may be associated with COPD susceptibility in the Han Chinese population. The FAM13A AA genotype increased serum MMP-12 levels and correlated with emphysema phenotype.

Humans↗

Genetic Identification of Burned Human Remains: A Systematic Review.

Background/Objectives: DNA-based identification of degraded human remains represents a major challenge in forensic science, particularly in cases involving burned, fragmented, or commingled bodies. Advances in forensic genetics have expanded the analytical capabilities for such samples; however, the effectiveness of different approaches and their integration within Disaster Victim Identification (DVI) workflows remain heterogeneous. This systematic review aims to critically evaluate current evidence on DNA-based identification of degraded remains, focusing on methodological strategies, emerging genomic technologies, and DVI applications, while integrating laboratory evidence and operational forensic practice into a structured analytical framework. Methods: A systematic literature search was conducted in Scopus and Web of Science from database inception to 5 June 2026, following PRISMA 2020 guidelines. Eligible studies included original research addressing DNA analysis of degraded, thermally altered, or highly compromised human remains in forensic or DVI contexts. After a multistep screening process involving title/abstract and full-text evaluation, 37 studies were included. Data were extracted and organized into three thematic categories: (i) core DNA analysis, (ii) advanced molecular technologies, and (iii) DVI case applications. Results: The findings demonstrate that DNA recovery from degraded remains is influenced by thermal exposure, tissue type, and sampling strategy. Teeth and dense cortical bone consistently provide higher DNA yield. While autosomal STR profiling remains the primary analytical approach, its limitations in highly degraded samples are mitigated through the complementary use of mitochondrial DNA (mtDNA), Y-chromosome STRs (Y-STRs), and SNP markers, together with advanced sequencing technologies such as massively parallel sequencing (MPS). Emerging technologies, including rapid DNA systems and predictive models based on macroscopic indicators, significantly enhance efficiency and success rates. DVI studies report identification rates exceeding 90-95% when multidisciplinary and structured workflows are applied. The evidence further supports a flexible triage-based analytical strategy, in which marker selection is guided by tissue preservation and degradation level. Conclusions: DNA-based identification of degraded human remains has evolved into an adaptive, multi-level forensic process. Successful outcomes rely on the integration of optimized sampling, hierarchical genetic analysis, and coordinated DVI strategies. The findings support a triage-based framework that links tissue selection, degradation assessment, and analytical methodology to maximize identification success. Future developments should focus on predictive models, advanced genomic tools, and standardized workflows to further improve identification in challenging forensic scenarios.

Humans↗

Exome-wide association study of bleeding events in patients receiving direct oral anticoagulants.

BackgroundDirect oral anticoagulants (DOACs) are first-line medications for stroke prevention in non-valvular atrial fibrillation (AF). However, variability in drug response poses risks of hemorrhagic or thromboembolic events.ObjectivesAlthough genetic influences on DOACs safety are increasingly recognized, robust evidence directly linking specific polymorphisms to bleeding risk remains limited.DesignMulti-center observational case-control study including exome-wide association analysis of 196 non-valvular AF patients treated with rivaroxaban or apixaban, comprising 97 with bleeding complications and 99 without.MethodsDOAC plasma concentrations, urinary 6-β-hydroxycortisol and cortisol levels were measured for CYP3A4 phenotyping. Sequencing was performed on the DNBSEQ G-400 platform. Single-nucleotide variant (SNV) associations with bleeding risk were assessed using logistic regression with additive, dominant, and recessive genetic models. Polygenic risk scores (PRSs) were calculated to evaluate cumulative genetic effects.ResultsNo SNVs reached Bonferroni-corrected significance under any model. PRSs showed weak predictive ability for bleeding with apixaban. For rivaroxaban, regression indicated that ln Css min/D + 1 index increased with PRS, age, and 6-β-hydroxycortisol/cortisol ratio, but decreased with higher 6-β-hydroxycortisol and coronary heart disease presence. No statistically significant differences were found for the PharmGKB Level 3 variants rs1045642 (rivaroxaban) and rs2231142 (apixaban). Trends toward statistical significance were observed for the rs2472304-G variant in rivaroxaban users, rs6977165-C in apixaban users, and for the CYP3A4*1/*36 diplotype.ConclusionResidual equilibrium concentration of DOACs, including dose-adjusted, did not independently predict bleeding risk in non-valvular AF patients. Variants rs2472304 and rs6977165 may warrant further investigation as potential contributors to bleeding risk.

Humans↗

Polysomal Profiling Coupled to Allele-Specific Proteomics Reveals an EIF4H TranSNP Allele Possessing Higher mRNA Translation Potential.

To search for genetic sources of allele-specific mRNA translation, we leveraged heterozygous polymorphisms and variants present in the exome of HCT116 colorectal adenocarcinoma-derived cells, computing allelic fractions from both total and polysome-associated RNA from RNA-Seq data. Allelic imbalance in polysomal RNA led us to nominate 52 coding variants associated with allele-specific mRNA translation, of which 16 are nonsynonymous. To validate instances of allele-specific translation, a proteomics workflow was developed that combines label-free shotgun analysis, high-pH reversed-phase peptide fractionation, and targeted parallel reaction monitoring using isotope-labeled peptide standards. Using this approach, we provide proof-of-concept validation of the heterozygous G>A, R183H missense single-nucleotide variant rs1554710467 in the eukaryotic initiation factor 4H (EIF4H) gene. The variant is present in two EIF4H alternatively spliced variants, which showed equivalent translation efficiency in HCT116 cells but differ in abundance. The alternative peptide containing H183 was significantly more abundant than the corresponding reference peptide containing R183, consistent with the over-representation of the alternative allele in polysomal RNA in HCT116 cells. A dual-fluorescence ribosome-stalling assay confirmed the enhanced translation potential of the variant allele. The two EIF4H allelic proteins exhibited similar stability and subpolysomal localization. This study demonstrates the feasibility of using allele-specific proteomics at the endogenous protein levels by exploiting heterozygous coding variants. Overall, our approach extends the toolbox available to investigate allele-specific differences in mRNA translation potential, a relatively underexplored layer of gene expression regulation that could reveal interindividual differences in disease-relevant phenotypes.

Humans↗

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans↗

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Polymorphism, Single Nucleotide↗

Characterization of phosphorylation variants for identifying adaptive alleles in Zea.

Large-scale genome sequencing of maize wild species (teosinte) has uncovered thousands of genetic mutations, but distinguishing causal alleles from neutral variations remains a significant challenge. In this study, we conducted a comprehensive analysis of phosphorylation-associated single-nucleotide variations (pSNVs) to enhance our understanding of adaptive variations in the Zea genus. We collected 234 teosinte genomes from seven different taxa and 507 cultivated maize genomes to identify single-nucleotide variants that target phosphorylation machinery, which is crucial for plant development and environmental adaptation. Our analysis identified 33 687 pSNVs within the Zea genus and revealed a reduction in genetic conservation along with an increase in protein abundance and expression for genes harboring pSNVs. Additionally, pSNVs present stronger purifying selection pressures compared with other missense mutations. We found that maize possesses fewer pSNVs than teosinte, likely due to the effects of selection and hitchhiking. By examining the role of pSNVs related to kinase-substrate rewriting events and exhibiting evolutionary divergence jointly, our results suggest that pSNVs impact multiple traits, particularly flowering time variation between teosinte and maize. Furthermore, we documented the widespread presence of pSNVs in Arabidopsis thaliana, rice, and wheat, identifying 46 pSNVs that have convergently evolved between maize and other species. Our study provides another insight into uncovering adaptive alleles in wild species by incorporating protein signaling sites and emphasizes the potential of utilizing wild species for future crop improvement.

Zea mays↗

G4SNVHunter: An R/Bioconductor Package for Evaluating SNV-Induced Disruption of G-Quadruplex Structures Leveraging the G4Hunter Algorithm.

G-quadruplexes (G4s) are nucleic acid secondary structures with important regulatory functions. Single-nucleotide variants (SNVs), one of the most common forms of genetic variation, can potentially impact the formation of G4 structures if they occur within G4 regions. However, there is currently a lack of software tools specifically designed to assess such effects. Here, we present an R/Bioconductor package named G4SNVHunter, which enables rapid detection of variants that may disrupt G4 structures. This tool, based on the core principles of the G4Hunter algorithm, can provide precise quantitative assessment of the propensity for G4 formation within genomic sequences. Specialized experimental methods can then be designed based on the results provided by G4SNVHunter to further verify the specific functions of the affected G4 structures, facilitating deeper insights into the biological impacts of genetic variants from the perspective of G4 structures. To showcase the functionality of the G4SNVHunter package, we analyzed the Neandertal and Denisovan archaic introgressed variants detected by the Sprime software, and identified approximately 5,800 variants located within G4 regions, among which around 230 may impair G4 structure formation propensity. The source code for the G4SNVHunter package has been publicly released under the MIT license at https://github.com/rongxinzh/G4SNVHunter and https://bioconductor.org/packages/devel/bioc/html/G4SNVHunter.html.

G-Quadruplexes↗

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning↗

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals↗

AUTS2-related syndrome: Insights from a large European cohort.

PURPOSE: AUTS2-related syndrome is characterized by developmental delay, autism spectrum disorder, and intellectual disability. From alternative promoters, AUTS2 encodes 2 distinct long and short isoforms encoding a putative transcriptional activator. METHODS: Through a European collaborative study, we collected clinical and genotype data on the largest AUTS2-related syndrome cohort of 58 patients harboring genomic rearrangements or single-nucleotide variants (SNVs). RESULTS: Pathogenic SNVs were recurrently found in individuals from different countries, suggesting mutational hotspots. Independent of the underlying defect at the AUTS2 locus, we observed that autistic behavior, hyperactivity, learning difficulties, and speech delay are common features of AUTS2-related syndrome. Among patients with SNVs, individuals carrying pathogenic variants affecting both longer and shorter AUTS2 transcripts showed a recognizable phenotype with microcephaly, brachycephaly, microretrognathia, broad nasal base, and anteverted nares. Behavioral disorders were more common in patients with variants affecting only the longer isoform. Arthrogryposis and stiff movements were only observed in patients with SNVs. CONCLUSION: This study provides a comprehensive clinical characterization of AUTS2-related syndrome, reveals few genotype-phenotype correlations, and suggests that the disruption of the 2 distinct AUTS2 transcripts has a different impact on the clinical phenotype.

Humans↗

Combined somatic mutation and transcriptome analysis reveals region-specific differences in clonal architecture in human cortex.

The human cerebral cortex is specialized into regions, but little is known about how human cellular lineages shape cortical regional variation and neuronal cell-type distribution during development. Here, we map single-cell lineages of human cortical regions and neuronal subtypes using >1,000 somatic single-nucleotide variants (sSNVs) identified from deep bulk whole-genome sequencing and analyzed over 25 regions and >72,000 single cells. In the fronto-parietal cortex, sSNVs are rarely restricted, marking neuron-generating clones that disperse into neighboring regions. In contrast, the primary visual cortex harbors 30%-70% more sSNVs than the neighboring secondary visual cortex. Clones at this border exhibit more restricted dispersion, suggesting late developmental lineage segregation. Single-nucleus sSNV and whole-transcriptome analysis reveal glutamatergic neuron clones with modest regional restrictions that share low-mosaic sSNVs with some GABAergic neurons, suggesting a recent dorsal cortical progenitor. Our analysis reveals human-specific cortical lineage patterns, regional differences in clonal patterns, and late divergence of some glutamatergic/GABAergic lineages.

Humans↗

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans↗

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software↗

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans↗