PubMed HealthSearch

SEARCH · PubMed Health

Results for “SNPs”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Virulence-associated variants in Cryptococcus neoformans sequence type 93 are less likely to be associated with population structure compared to independent rare mutations.

Cryptococcus neoformans is a pathogenic yeast that is the causative agent of cryptococcal meningitis. While it is well known that the genotype of C. neoformans impacts patient outcomes, the reason for this association has not been well elucidated. In this study, we examined the relationship between two subpopulations in the sequence type 93 clade of C. neoformans: ST93A and ST93B. We found extensive linkage disequilibrium (LD) among the single nucleotide polymorphisms (SNPs) that differentiate ST93A from ST93B. We also found differences in the extent of linkage among SNPs within each subpopulation; LD was more extensive within ST93B than ST93A. SNPs associated with virulence were in long-range linkage disequilibrium with less frequency than recurrent SNPs not associated with virulence. We investigated the karyotype of ST93A and ST93B using contour-clamped gel electrophoresis and long-read sequencing and found that the extensive long-range linkage was not due to chromosomal rearrangements. Overall, we found that the two subpopulations in ST93 are driven by SNPs in LD. We additionally found that recurrent SNPs associated with virulence were less frequently evolutionarily linked and were two times more likely to be independent, congruent mutations rather than tied to phylogeny.IMPORTANCECryptococcus neoformans is an important pathogen that is widely distributed and ubiquitous in the environment. The majority of the human population has a latent, controlled infection suggesting that C. neoformans is uniquely adapted to cause infection. In spite of this, the reason C. neoformans is a pathogen remains unknown; interestingly, most environmental isolates are avirulent but are genetically very similar to disease-causing virulent isolates. Recent evidence from genome-wide association studies shows that small mutations in key virulence-associated genes are associated with the virulence of specific isolates. The data presented here provide an evolutionary framework for those small mutations. The mutations that impact disease are not being collected over long-term evolution. The mutations may instead occur independently during infection. Identifying these genes that are more likely to be mutated during infection will be fundamental for understanding C. neoformans virulence.

Cryptococcus neoformans

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180 bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve > 20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total = 84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans

FLT4 gene polymorphisms influence isolated ventricular septal defect predisposition in a Southwest China population.

BACKGROUND: Ventricular septal defect (VSD) is the most common congenital heart disease. Although a small number of genes associated with VSD have been found, the genetic factors of VSD remain unclear. In this study, we evaluated the association of 10 candidate single nucleotide polymorphisms (SNPs) with isolated VSD in a population from Southwest China. METHODS: Based on the results of 34 congenital heart disease whole-exome sequencing and 1000 Genomes databases, 10 candidate SNPs were selected. A total of 618 samples were collected from the population of Southwest China, including 285 VSD samples and 333 normal samples. Ten SNPs in the case group and the control group were identified by SNaPshot genotyping. The chi-square (&#x3c7;2) test was used to evaluate the relationship between VSD and each candidate SNP. The SNPs that had significant P value in the initial stage were further analysed using linkage disequilibrium, and haplotypes were assessed in 34 congenital heart disease whole-exome sequencing samples using Haploview software. The bins of SNPs that were in very strong linkage disequilibrium were further used to predict haplotypes by Arlequin software. ViennaRNA v2.5.1 predicted the haplotype mRNA secondary structure. We evaluated the correlation between mRNA secondary structure changes and ventricular septal defects. RESULTS: The &#x3c7;2 results showed that the allele frequency of FLT4 rs383985 (P&#x2009;=&#x2009;0.040) was different between the control group and the case group (P&#x2009;<&#x2009;0.05). FLT4 rs3736061 (r2&#x2009;=&#x2009;1), rs3736062 (r2&#x2009;=&#x2009;0.84), rs3736063 (r2&#x2009;=&#x2009;0.84) and FLT4 rs383985 were in high linkage disequilibrium (r2&#x2009;>&#x2009;0.8). Among them, rs3736061 and rs3736062 SNPs in the FLT4 gene led to synonymous variations of amino acids, but predicting the secondary structure of mRNA might change the secondary structure of mRNA and reduce the free energy. CONCLUSIONS: These findings suggest a possible molecular pathogenesis associated with isolated VSD, which warrants investigation in future studies.

Child

Relationship between mucin gene polymorphisms and different types of gallbladder stones.

BACKGROUND: Gallstones, a common surgical condition globally, affect around 20% of patients. The development of gallstones is linked to abnormal cholesterol and bilirubin metabolism, reduced gallbladder function, insulin resistance, biliary infections, and genetic factors. In addition to these factors, research has shown that mucins play a role in gallstone formation. This study aims to explore the connection between different types of gallstones and mucin gene polymorphisms. METHODS: For this purpose, a total of 121 patients with gallbladder stones PNS and 107 patients with healthy controls PNS were enrolled in this case-control study. One SNPs (rs4072037) of MUC1 gene&#x3001; three SNPs (rs2856111&#x3001;rs41532344&#x3001;rs41349846) of MUC2 gene&#x3001;four SNPs (rs712005&#x3001;rs2246980&#x3001;rs2258447&#x3001;rs2259292) of MUC4 gene&#x3001;seven SNPs (rs28415193&#x3001;rs56047977&#x3001;rs2037089&#x3001;rs2075854&#x3001;rs3829224&#x3001;rs2672785&#x3001;rs2735709) of MUC5 gene&#x3001;eight SNPs (rs10902268&#x3001;rs61869016&#x3001;rs573849895&#x3001;rs59257210&#x3001;rs7396383&#x3001;rs74644072&#x3001;rs7481521&#x3001;rs9704308) of MUC6 gene&#x3001;five SNPs (rs10229731&#x3001;rs73168398&#x3001;rs4729655&#x3001;rs55903219&#x3001;rs74974199) of MUC17 gene. We amplified SNP sites by polymerase chain reaction (PCR) using specific primer sets followed by DNA sequencing. RESULTS: The frequencies of MUC2 rs2856111 C/T genotype (OR&#x2009;=&#x2009;3.81, 95%CI: 1.06-13.68) was higher than the control group. MUC17 rs10229731 A/C genotype (OR&#x2009;=&#x2009;0.33, 95%CI: 0.12-0.95), rs73168398 G/A genotype (OR&#x2009;=&#x2009;0.26, 95%CI: 0.07-0.98), MUC6 rs10902268 G/A genotype (OR&#x2009;=&#x2009;0.40, 95%CI: 0.17-0.95) at lower frequencies than controls. The frequencies of MUC2 rs41532344 T allele (OR&#x2009;=&#x2009;2.55, 95%CI: 1.06-6.13), MUC4 rs712005 G allele (OR&#x2009;=&#x2009;2.51, 95%CI: 1.20-5.22), MUC5B rs2037089 C allele (OR&#x2009;=&#x2009;3.54, 95%CI: 1.14-11.01) and MUC5AC rs28415193 G allele (OR&#x2009;=&#x2009;1.77, 95%CI: 1.02-3.07) were higher than the control group. MUC6 rs10902268 A allele (OR&#x2009;=&#x2009;0.004, 95%CI: 0.00-0.27), rs61869016 C allele (OR&#x2009;=&#x2009;0.07, 95%CI: 0.01-0.63) at lower frequencies than controls. CONCLUSIONS: Polymorphisms in the mucin gene were linked to the formation of gallbladder stones. The MUC4 rs712005 G allele, MUC5B rs2037089 C allele, MUC2 rs41532344 T allele and MUC5AC rs28415193 G allele were found to predispose individuals to the development of the disease. MUC6 rs10902268 A allele and rs61869016 C allele were identified as protective factors. Meanwhile, MUC2 rs2856111 CT genotype was found to predispose individuals to the development of the disease. MUC17 rs10229731 AC genotype, rs73168398 GA genotype and MUC6 rs10902268 GA genotype were identified as protective factors.

Humans

Assessing Hardy-Weinberg equilibrium in T2T-aligned 1000 genomes project.

Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. These motivate a re-examination of HWE in sex-aware analyses. Using the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequencing data from 2,490 individuals in the 1000 Genomes Project, we examined genome-wide sex-specific deviations from HWE across five super-populations. Our analyses were restricted to bi-allelic SNPs with non-missing genotypes and minor allele frequency (MAF) &#x2265;5% in both sexes of the five super-populations. We applied an allele-based framework to quantify both the magnitude and direction of Hardy-Weinberg disequilibrium (HWD), followed by a second-order omnibus meta-analysis that combined HWD results across populations and sexes. At a genome-wide significance threshold of p&#x2009;<&#x2009;5e-8, 0.9% of autosomal SNPs exhibited significant deviations from HWE. The majority of these deviations were associated with genomic features indicative of poor sequence quality. Restricting the analysis to reliable genomic regions substantially reduced the number of signals, yielding 255 autosomal SNPs and one non-pseudoautosomal chromosome X SNP. Among these, 140 autosomal SNPs displayed significant heterogeneity across populations but not across sexes. Notably, eight SNPs within a 15-bp region on chromosome 14q31.3 showed excess heterozygosity in both sexes of the African super-population (AFR). Finally, we developed a multivariate predictor of HWD based on sequence features, providing a practical tool that can be integrated into existing quality control pipelines for whole genome sequencing studies.

Journal Article

Single Nucleotide Polymorphisms in RUNX2 and BMP2 contributes to different vertical facial profile.

The vertical facial profile is a crucial factor for facial harmony with significant implications for both aesthetic satisfaction and orthodontic treatment planning. However, the role of single nucleotide polymorphisms (SNPs) in the development of vertical facial proportions is still poorly understood. This study aimed to investigate the potential impact of some SNPs in genes associated with craniofacial bone development on the establishment of different vertical facial profiles. Vertical facial profiles were assessed by two senior orthodontists through pre-treatment digital lateral cephalograms. The vertical facial profile type was determined by recommended measurement according to the American Board of Orthodontics. Healthy orthodontic patients were divided into the following groups: "Normodivergent" (control group), "Hyperdivergent" and "Hypodivergent". Patients with a history of orthodontic or facial surgical intervention were excluded. Genomic DNA extracted from saliva samples was used for the genotyping of 7 SNPs in RUNX2, BMP2, BMP4 and SMAD6 genes using real-time polymerase chain reactions (PCR). The genotype distribution between groups was evaluated by uni- and multivariate analysis adjusted by age (alpha = 5%). A total of 272 patients were included, 158 (58.1%) were "Normodivergent", 68 (25.0%) were "Hyperdivergent", and 46 (16.9%) were "Hypodivergent". The SNPs rs1200425 (RUNX2) and rs1005464 (BMP2) were associated with a hyperdivergent vertical profile in uni- and multivariate analysis (p-value < 0.05). Synergistic effect was observed when evaluating both SNPs rs1200425- rs1005464 simultaneously (Prevalence Ratio = 4.0; 95% Confidence Interval = 1.2-13.4; p-value = 0.022). In conclusion, this study supports a link between genetic factors and the establishment of vertical facial profiles. SNPs in RUNX2 and BMP2 genes were identified as potential contributors to hyperdivergent facial profiles.

Polymorphism, Single Nucleotide

The genetic basis of adaptation to copper pollution in Drosophila melanogaster.

Introduction: Heavy metal pollutants can have long lasting negative impacts on ecosystem health and can shape the evolution of species. The persistent and ubiquitous nature of heavy metal pollution provides an opportunity to characterize the genetic mechanisms that contribute to metal resistance in natural populations. Methods: We examined variation in resistance to copper, a common heavy metal contaminant, using wild collections of the model organism Drosophila melanogaster. Flies were collected from multiple sites that varied in copper contamination risk. We characterized phenotypic variation in copper resistance within and among populations using bulked segregant analysis to identify regions of the genome that contribute to copper resistance. Results and Discussion: Copper resistance varied among wild populations with a clear correspondence between resistance level and historical exposure to copper. We identified 288 SNPs distributed across the genome associated with copper resistance. Many SNPs had population-specific effects, but some had consistent effects on copper resistance in all populations. Significant SNPs map to several novel candidate genes involved in refolding disrupted proteins, energy production, and mitochondrial function. We also identified one SNP with consistent effects on copper resistance in all populations near CG11825, a gene involved in copper homeostasis and copper resistance. We compared the genetic signatures of copper resistance in the wild-derived populations to genetic control of copper resistance in the Drosophila Synthetic Population Resource (DSPR) and the Drosophila Genetic Reference Panel (DGRP), two copper-na&#xef;ve laboratory populations. In addition to CG11825, which was identified as a candidate gene in the wild-derived populations and previously in the DSPR, there was modest overlap of copper-associated SNPs between the wild-derived populations and laboratory populations. Thirty-one SNPs associated with copper resistance in wild-derived populations fell within regions of the genome that were associated with copper resistance in the DSPR in a prior study. Collectively, our results demonstrate that the genetic control of copper resistance is highly polygenic, and that several loci can be clearly linked to genes involved in heavy metal toxicity response. The mixture of parallel and population-specific SNPs points to a complex interplay between genetic background and the selection regime that modifies the effects of genetic variation on copper resistance.

Drosophila

Adaptation to Plant Defence in an Agricultural Insect Pest: Integrating Genome Scans and Gene Expression in the Soybean Aphid Reveals Multi-Genic Pathways.

In agroecosystems, intense selection pressures cause species to adapt and spread, often leading to the evolution and persistence of pests. Understanding how pests rapidly adapt can help develop sustainable strategies for their management and improve agroecosystem health. Pest adaptation involves stable variations in DNA sequence, as well as dynamic shifts in gene expression, often mediated by non-coding regulatory elements. We examined adaptation to plant defences in the soybean aphid, Aphis glycines, in which virulent aphids have overcome plant defences and avirulent aphids have not. Previous data with laboratory colonies suggested that virulent aphids have higher overall gene expression, including transposable elements, some of which influence gene regulation. However, we lack information on how genetic variation in natural populations impacts adaptation and potentially gene regulation. We integrated population genome scans of field-collected, soybean aphid populations with gene expression profiles of virulent and avirulent laboratory colonies to uncover connections between genetic differentiation and gene regulation for virulence. Genome scan methods found 2144 single nucleotide polymorphisms (SNPs) with significant genetic differentiation (i.e., outliers) in field-collected populations. These SNPs were near 1004 genes, representing 5.16% of the effective number of genes. Based on previous RNA-Seq data with laboratory colonies, we found 3160 genes and 147 long non-coding RNAs (lncRNAs) with differential expression among virulent and avirulent biotypes. By integrating both data sets, we identified 16 genes and 5 long non-coding RNAs with differential expression and that were associated with an outlier SNP (within 10&#x2009;kbp). We validated SNPs with additional field collected aphids and found an aphid clone with stronger virulence than our laboratory virulent colony, surviving on 2 different aphid-resistant soybean varieties. This new virulent clone had fixed allele differences at 9 SNPs compared to our avirulent and other virulent colony. Field collected soybean aphids matching the phenotype of this new virulent clone had significant genetic differentiation with 3 outlier SNPs near genes related to zinc transport and lachesin compared to field collected avirulent aphids. Our entire data reinforced the importance of a potential multi-genetic response to overcome plant defence and generates new insights into complex genetic and regulatory mechanisms involved in insect-plant interactions.

Animals

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans

Genome-wide association study of angiotensinogen levels and key single nucleotide polymorphism associations with blood pressure.

OBJECTIVE: The renin angiotensin aldosterone system plays a key role in circulatory homeostasis. We sought to identify genetic determinants of measured plasma angiotensinogen levels and subsequently evaluate the association of these single nucleotide polymorphisms (SNPs) with blood pressure (BP) and hypertension in a multiethnic population. METHODS: Genome-wide association study (GWAS) of plasma angiotensinogen levels, measured using an enzyme-linked immunoassay, was conducted in 4899 Multi-Ethnic Study of Atherosclerosis (MESA) participants (self-identified as White, n = 1865; Hispanic, n &#x200a;=&#x200a;1113; Black, n &#x200a;=&#x200a;1224; and Chinese, n &#x200a;=&#x200a;629). Linear and logistic models examined the association between SNPs with angiotensinogen and hypertension, respectively. Mediation analysis evaluated the effect of angiotensinogen on BP/hypertension through the top SNPs identified by GWAS. RESULTS: In the analysis utilizing all participants, 115 SNPs were associated with angiotensinogen ( P &#x200a;<&#x200a;5&#x200a;&#xd7;&#x200a;10 -8 ), including lead SNP rs4762(G>A) in exon 2 ( P &#x200a;=&#x200a;1.51E -100 ) and rs5050(T>G) in the promoter region ( P &#x200a;=&#x200a;2.26E -69 ) of the AGT gene. Race/ethnic-specific analyses identified rs4762(G>A) as the lead SNP for White and Hispanic participants, whereas Black and Chinese participants had rs5050(T>G) and rs16852311(G>C), respectively. Both rs4762(G>A) and rs5050(T>G) indirectly increased systolic BP, diastolic BP, and the odds of hypertension through its effect of increasing angiotensinogen. CONCLUSIONS: Our findings demonstrate racial/ethnic differences in genetic effects on angiotensinogen levels across multiple SNPs. AGT rs4762(G>A) and rs5050(T>G) impact BP and hypertension through a mediated effect via angiotensinogen, though opposing direct effects may mask the overall association.

Humans

A systematic review and network meta-analysis of single nucleotide polymorphisms associated with oral submucous fibrosis risk.

BACKGROUND: Oral submucous fibrosis (OSF) is a chronic and insidious oral disease characterized by hyalinization of the subepithelial connective tissue and progressive fibrosis of the oral submucosa. It is a precancerous condition of oral squamous cell carcinoma. Studies have demonstrated that single nucleotide polymorphisms (SNPs) are closely associated with susceptibility to OSF. This study aims to comprehensively evaluate the association between SNPs and OSF risk and to rank the strength of the association between different genetic models and OSF susceptibility. METHODS: Literature related to OSF was comprehensively searched from PubMed, Web of Science, Embase, Cochrane Library, CNKI, and Wangfang databases up to July 2025. Full-text case-control studies with patients diagnosed with OSF were included. Quality assessment was performed to evaluate the risk of bias. RevMan 5.4, GeMTC 0.14.3, and STATA 17.0 were used for the pairwise and Bayesian network meta-analysis. RESULTS: A total of 24 studies with 2545 cases and 3772 controls, covering 13 SNPs in 11 genes, were included in our meta-analysis. We found that CYP1A1 rs4646903:T>C, CYP1A1 rs1048943:A>G, GSTT1 null genotype, GSTM1 null genotype, and XRCC3 rs861539:C>T were associated with an increased risk of OSF, while MMP2 rs243865:C>T and MMP3 rs3025058: 5A>6A were associated with a decreased risk of OSF. Further Bayesian network meta-analysis indicated the top 5 genetic models with the highest association with OSF risk in network group 1 were the dominant model, homozygous model, allelic model, and recessive model of CYP1A1 rs1048943:A>G (ranked 1-4), and the heterozygous/dominant model of CYP1A1 rs4646903:T>C (both ranked 5). While the allelic models of XRCC3 rs861539:C>T and MMP3 rs3025058: 5A>6A ranked first for predicting OSF in group 2 and group 3, respectively. CONCLUSION: Some specific SNPs are significantly related to the risk of OSF. Among them, the dominant model of CYP1A1 rs1048943:A>G may be the most strongly associated genetic model with OSF risk. Future large-sample, well-designed studies with detailed genotype data are needed to validate the roles of these SNPs in OSF risk.

Humans

Phylogenetic inconsistency of pairwise SNP clustering for inferring tuberculosis transmission in a high-burden, endemic setting: a case study from Thailand.

Whole-genome sequence analysis is now widely used to delineate tuberculosis transmission clusters. A standard practice is to cluster bacterial isolates based on a fixed maximum genome-wide pairwise single nucleotide polymorphism (pwSNP) distance threshold. In this study, we evaluated the phylogenetic consistency of pwSNP-distance clustering with thresholds ranging between 1 and 25 single nucleotide polymorphisms (SNPs) using two contrasting data sets: (i) a data set from the UK (N = 390) published by T. M. Walker, C. L. C. Ip, R. H. Harrell, J. T. Evans, et al. (Lancet Infect Dis 13:137-146, 2013, https://doi.org/10.1016/S1473-3099(12)70277-3), which was foundational to the establishment of this method, and (ii) a data set from Thailand (N = 3,341), characterized by persistent transmission and sparse, non-systematic sampling. For the UK data set, the standard pwSNP-distance clustering using thresholds of &#x2265;12 SNPs yielded entirely monophyletic clusters and showed high concordance with a comparative monophyly constrained, tree-based method. In contrast, for the Thai data set, pwSNP-distance clustering often generated non-monophyletic clusters, even by the 25-SNP threshold. The pwSNP-distance and comparative tree-based clustering methods only showed large consistency at thresholds of &#x2265;22 SNPs. This suggests that SNP clusters defined by low distance thresholds (i.e., <12 SNPs for the UK data set, and <22 SNPs for the Thai data set) may lack robustness, and the problem is particularly severe for data sets characterized by persistent transmission, likely due to poorer cluster separation. Moreover, our findings indicate that large cluster sizes, high maximum intra-cluster genetic distances, and broad sample collection time spans may serve as useful indicators of potentially non-monophyletic clusters. We also demonstrate that mixed infections can produce spurious, phylogenetically long-range SNP linkages, underscoring the necessity of strict sequence quality control.IMPORTANCEFixed-threshold pairwise single nucleotide polymorphism (pwSNP)-distance clustering is commonly used to delineate tuberculosis transmission clusters. From an epidemiological perspective, a genuine transmission cluster must be monophyletic, originating from a single source. However, pwSNP-distance clustering is inherently simplistic and can therefore violate this principle, making the assessment of its phylogenetic consistency critical. Our results demonstrate that while this method effectively delineated complete transmission clusters for the data set from the UK, a low-burden and non-persistent transmission setting, it frequently generated non-monophyletic clusters when applied to the Thai data set, characterized by persistent transmission alongside sparse and non-systematic sampling. Furthermore, we found that clusters derived using low distance thresholds could notably vary between the pwSNP-distance and comparative tree-based clustering methods, suggesting limited reliability and robustness. To accurately delineate tuberculosis transmission clusters, especially for complex data from high-burden, endemic settings, we recommend transitioning from pwSNP-distance clustering toward more robust, phylogenetic clustering that respects evolutionary descent.

Mycobacterium tuberculosis

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50&#xa0;K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40&#xa0;kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho&#x2009;=&#x2009;0), the beta-binomial distribution (Rho&#x2009;=&#x2009;0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5&#xa0;K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus

Genetic variants related to successful migraine prophylaxis with verapamil.

BACKGROUND: Currently, there is no biologically based rationale for drug selection in migraine prophylactic treatment. METHODS: To investigate the genetic variation underlying treatment response to verapamil prophylaxis, we selected 225 patients from a longitudinally established, deeply phenotyped migraine database (N&#xa0;=&#xa0;5983), and collected uninterrupted quantitated verapamil treatment response data and DNA for these 225 cases. We recorded the number of headache days in the four weeks preceding treatment with verapamil and for four weeks, following completion of a treatment period with verapamil lasting at least five weeks. Whole-exome sequencing (WES) was applied to a discovery cohort consisting of 21 definitive responders and 14 definitive non-responders, and the identified single nucleotide polymorphisms (SNPs) showing significant association were genotyped in a separate confirmation cohort (185 verapamil treated patients). Statistical analysis of the WES data from the discovery cohort identified 524 SNPs associated with verapamil responsiveness (p&#xa0;<&#xa0;0.01); among them, 39 SNPs were validated in the confirmatory cohort (n&#xa0;=&#xa0;185) which included the full range of response to verapamil from highly responsive to not responsive. RESULTS: Fourteen SNPs were confirmed by both percentage and arithmetic statistical approaches. Pathway and protein network analysis implicated myo-inositol biosynthetic and phospholipase-C second messenger pathways in verapamil responsiveness, emphasizing the earlier pathogenic understanding of migraine. No association was found between genetic variation in verapamil metabolic enzymes and treatment response. CONCLUSION: Our findings demonstrate that genetic analysis in well-characterized subpopulations can yield important pharmacogenetic information pertaining to the mechanism of anti-migraine prophylactic medications.

Chemoprevention

Novel genomic regions associated with adult-plant resistance to multiple fungal pathogens in wheat (Triticum aestivum L.) revealed by DArT marker sequencing.

Wheat is among the top three most important cereal crops globally and serves as a staple food for approximately 40% of the world's population. Fungal leaf diseases such as yellow and leaf rusts (YR, LR), septoria nodorum blotch (SNB), septoria tritici blotch (STB), and powdery mildew (PM) have a major effect on yield loss in wheat, and resistance breeding is so far the most effective strategy to minimize those losses. Adult plant resistance (APR) is a crucial component of durable disease resistance; it reduces the pathogen's infection rate, keeping disease levels below the damage threshold, even in the absence of complete immunity. Therefore, this study aimed to identify sources of resistance in a collection of 411 accessions from diverse global origins. These accessions were phenotyped across 2018-2019. DArTseq technology and Genome-wide association studies (GWAS) analysis were conducted to identify single-nucleotide polymorphisms (SNPs) associated with APR for evaluated pathogens. DArT analysis showed that wheat chromosome 2B contains genomic regions associated with resistance to SNB, and that SNPs on chromosome 3B are associated with resistance to YR. On chromosome 6&#xa0;A, there is a strong potential to explore, as a shared resistance locus for YR and SNB was found. SNPs: 3,937,236, 1,056,817 were consistent in both years, meaning their association with disease resistance is reliable and repeatable. Chromosome 7D is a strong region for SNPs significantly associated with both LR and SNB resistance. While multiple disease resistance genes are present on 7D, the 610&#xa0;Mb LR locus is distinct from known LR, PM, and SNB loci, making it a strong candidate for functional validation. These findings highlight the value of historical resistance sources and uncover novel genomic regions for breeding a broad-spectrum APR-based resistance. Dual-trait loci, especially those effective against both biotrophic and necrotrophic pathogens, represent a promising material for achieving durable resistance in elite wheat cultivars.

Triticum

A cross-sectional study of oxidative stress pathway genotypes and their interactions with environmental pollutant levels identifies associations with gene expression and lung function.

BACKGROUND: Asthma is a heterogeneous disease influenced by genetic and environmental factors. Fine particulate matter (PM2.5) exacerbates asthma, likely through oxidative stress pathways, but whether genetic variation modifies this effect remains unclear. METHODS: We analysed data on 948 adults with asthma from the Severe Asthma Research Program (SARP), linking ZIP-code-level PM2.5 exposure with whole-genome sequencing data. We tested 4337 single nucleotide polymorphisms (SNPs) in 120 oxidative stress pathway genes for gene-environment (GxE) interactions with PM2.5 on lung function (forced expiratory volume in 1 s [FEV1] % predicted) using weighted linear regression. Gene expression data from bronchial epithelial cells (n = 170) were used to assess cis-expression quantitative trait loci (eQTLs). FINDINGS: Higher PM2.5 exposure was associated with lower FEV1% predicted (&#x3b2; per &#x3bc;g/m3 = -0.7, p = 0.01). We identified 20 SNPs across seven genes (OXSR1, PXDN, TPO, LRRK2, APP, MSRA, MSRB2) with significant GxE interactions after multiple-testing correction. Five SNPs were also eQTLs, linking PM2.5-modified gene expression to lung function. Minor alleles in OXSR1 and PXDN were associated with reduced gene expression and worsened FEV1% under high PM2.5 exposure. Conversely, TPO variants were associated with higher baseline expression and lower lung function, but under increasing PM2.5 exposure, minor allele carriers showed suppressed TPO expression and improved FEV1%. INTERPRETATION: This study identified 20 SNPs in oxidative stress pathway genes that modify the effect of PM2.5 on lung function in asthma. These findings highlight the importance of integrating environmental context in genetic studies and suggest potential therapeutic targets for pollution-sensitive asthma phenotypes. FUNDING: Supported by NIH grants.

Cross-Sectional Studies

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28&#xd7;, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS&#xa0;&#x2248;&#xa0;1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (H&#x2092;&#xa0;&#x2248;&#xa0;0.71) and nucleotide diversity (&#x3c0;&#xa0;&#x2248;&#xa0;7&#xa0;&#xd7;&#xa0;10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation