PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genetic variants”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models

Heterogeneous effects of genetic variants and traits associated with fasting insulin on cardiometabolic outcomes.

Elevated fasting insulin levels (FI), indicative of altered insulin secretion and sensitivity, may precede type 2 diabetes (T2D) and cardiovascular disease onset. In this study, we group FI-associated genetic variants based on their genetic and phenotypic similarities and identify seven clusters with distinct mechanisms contributing to elevated FI levels. Clusters fall into two types: "non-diabetogenic hyperinsulinemia," where clusters are not associated with increased T2D risk, and "diabetogenic hyperinsulinemia," where T2D associations are driven by body fat distribution, liver function, circulating lipids, or inflammation. In over 1.1 million multi-ancestry individuals, we demonstrated that diabetogenic hyperinsulinemia cluster-specific polygenic scores exhibit varying risks for cardiovascular conditions, including coronary artery disease, myocardial infarction (MI), and stroke. Notably, the visceral adiposity cluster shows sex-specific effects for MI risk in males without T2D. This study underscores processes that decouple elevated FI levels from T2D and cardiovascular risk, offering new avenues for investigating process-specific pathways of disease.

Humans

Mapping Focal and Generalized Effects of Common Genetic Variants on Human Brain Structure.

Genome-wide association studies (GWAS) have advanced the quest to understand how specific genetic variants influence human brain structure and function. Recent work has identified hundreds of common variants associated with subcortical brain volumes, sparking interest in how these genetic markers overlap across brain networks. While this can be estimated by hierarchical clustering of the genetic correlation matrix to identify modular patterns of shared architecture, no brain-wide maps of these effects are available. To address this, we computed polygenic scores (PGS) from loci associated with ten brain volume regions of interest (ROIs): nine major subcortical structures and intracranial volume, with each locus weighted by its association with regional volume. In an independent sample from the discovery GWAS, we performed large-scale segmentation of 3D volumetric T1-weighted MRI scans using voxel-based morphometry (VBM) to map 3D profile of regions where gray matter volume (GMV) was associated with each PGS. We found statistically significant, localized effects for PGS defined for the amygdala, thalamus, and basal ganglia, but PGS for brainstem volume was associated with widespread differences throughout the brain. These brain-wide maps reveal patterns consistent with both localized and distributed genetic influences, offering a novel approach to interpret the genomic architecture of brain structure.

GWAS

SNPannotator: automated functional annotation of genetic variants and linked proxies.

SUMMARY: Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation. AVAILABILITY AND IMPLEMENTATION: The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.

Software

High-impact rare genetic variants in severe schizophrenia.

Extreme phenotype sequencing has led to the identification of high-impact rare genetic variants for many complex disorders but has not been applied to studies of severe schizophrenia. We sequenced 112 individuals with severe, extremely treatment-resistant schizophrenia, 218 individuals with typical schizophrenia, and 4,929 controls. We compared the burden of rare, damaging missense and loss-of-function variants between severe, extremely treatment-resistant schizophrenia, typical schizophrenia, and controls across mutation intolerant genes. Individuals with severe, extremely treatment-resistant schizophrenia had a high burden of rare loss-of-function (odds ratio, 1.91; 95% CI, 1.39 to 2.63; P = 7.8 × 10-5) and damaging missense variants in intolerant genes (odds ratio, 2.90; 95% CI, 2.02 to 4.15; P = 3.2 × 10-9). A total of 48.2% of individuals with severe, extremely treatment-resistant schizophrenia carried at least one rare, damaging missense or loss-of-function variant in intolerant genes compared to 29.8% of typical schizophrenia individuals (odds ratio, 2.18; 95% CI, 1.33 to 3.60; P = 1.6 × 10-3) and 25.4% of controls (odds ratio, 2.74; 95% CI, 1.85 to 4.06; P = 2.9 × 10-7). Restricting to genes previously associated with schizophrenia risk strengthened the enrichment with 8.9% of individuals with severe, extremely treatment-resistant schizophrenia carrying a damaging missense or loss-of-function variant compared to 2.3% of typical schizophrenia (odds ratio, 5.48; 95% CI, 1.52 to 19.74; P = 0.02) and 1.6% of controls (odds ratio, 5.82; 95% CI, 3.00 to 11.28; P = 2.6 × 10-8). These results demonstrate the power of extreme phenotype case selection in psychiatric genetics and an approach to augment schizophrenia gene discovery efforts.

Aged

Identifying causal genetic variants for high-altitude adaptation through blood eQTL analysis in plateau populations.

A substantial number of genetic variants have been associated with high-altitude adaptation (HAA), yet most of them are located in non-coding genomic regions, leaving their specific functions and underlying mechanisms largely unknown. In this study, we analyze whole-genome and transcriptome sequencing data from a self-established cohort comprising 61 native highlanders (NHs) and 164 acclimatized newcomers (ANs), identifying 6,586 cis- and 34,203 trans-expression quantitative trait loci (eQTLs), along with 130 cell type-specific eQTLs. By further combining these data with a large East Asia (~30% Tibetan) genome-wide association study (GWAS) cohort, we employ colocalization and causal inference analyses to prioritize 85 cis-eQTLs associated with HAA and identify several novel candidate causal genes, including EXOC8, which is experimentally confirmed to regulate erythroid differentiation. Additionally, network analysis of these causal genes uncovers multiple regulatory pathways, mainly involving energy metabolism, autophagy, ubiquitination and inflammation. Our study offers a comprehensive eQTL map and reveals causal chains of "variant-gene-phenotype" for HAA-related traits, which provides new insights into potential regulatory mechanisms and targets for prevention and treatment of altitude sickness.

Quantitative Trait Loci

A comparison between the common type and a rare genetic variant of human cupro-zinc superoxide dismutase.

Human cupro-zinc superoxide dismutase is polymorphic in northern Sweden. The genetic variant type has a lower mobility at electrophoresis in alkaline buffer. The enzyme was isolated from erythrocytes from one of the rare homozygotes and its properties compared to those of the common type. The isoelectric point of the variant was higher (4.85) than that of the common type (4.7). Small differences in amino acid composition were found but no definite amino acid substitutions could be pointed out. The molecular weights were equal as judged from electrophoreses in polyacrylamide gels in the presence of dodecylsulphate. The ultraviolet spectra were similar. Parameters related to the active site of the enzyme were very similar; i.e. specific activity and sensitivity to inhibition by cyanide and by H2O2. These parameters, especially the latter two, differ widely between species. Both enzymes were stable for weeks at neutral pH at 37 degrees C, whereas the common type was significantly more stable at pH 4 and pH 11 and also at incubation in neutral buffer at 70 degrees C. It appears that the active site of the variant is conserved whereas the stability of the enzyme is affected.

Amino Acids

CAGI, the Critical Assessment of Genome Interpretation, establishes progress and prospects for computational genetic variant interpretation methods.

BACKGROUND: The Critical Assessment of Genome Interpretation (CAGI) aims to advance the state-of-the-art for computational prediction of genetic variant impact, particularly where relevant to disease. The five complete editions of the CAGI community experiment comprised 50 challenges, in which participants made blind predictions of phenotypes from genetic data, and these were evaluated by independent assessors. RESULTS: Performance was particularly strong for clinical pathogenic variants, including some difficult-to-diagnose cases, and extends to interpretation of cancer-related variants. Missense variant interpretation methods were able to estimate biochemical effects with increasing accuracy. Assessment of methods for regulatory variants and complex trait disease risk was less definitive and indicates performance potentially suitable for auxiliary use in the clinic. CONCLUSIONS: Results show that while current methods are imperfect, they have major utility for research and clinical applications. Emerging methods and increasingly large, robust datasets for training and assessment promise further progress ahead.

Humans

Genetic Variants Associated With the Biochemical Response to Vitamin D3 in the Multi-Ethnic Study of Atherosclerosis.

CONTEXT: The response to treatment with vitamin D varies between patients. OBJECTIVE: To identify genetic variants associated with the biochemical response to vitamin D3 supplementation. DESIGN: Randomized placebo-controlled trial conducted between 2017 and 2019. SETTING: The trial was nested in an ongoing community-based cohort study, the Multi-Ethnic Study of Atherosclerosis. INTERVENTION: 2000 International Units of vitamin D3 or placebo daily for 16 weeks. PARTICIPANTS: The analytic sample included 427 participants assigned to vitamin D3 (mean age, 73 years; 54% females) and was 36% White, 33% Black, 18% Hispanic, and 14% Chinese. MAIN OUTCOME MEASURES: The biochemical response to vitamin D3 included changes in serum concentrations of 1,25-dihydroxyvitamin D3 [1,25(OH)2D3], PTH, and 25-hydroxyvitamin D3 [25(OH)D3]. RESULTS: In genome-wide analyses, single nucleotide polymorphisms in 8 regions of the genome had significant association (P < 5E-08) with 1 of the traits (2 with change in 1,25(OH)2D3, 1 with change in PTH, and 5 with change in 25(OH)D3). rs16867276 within an intergenic region on 2q31 was associated with change in serum 1,25(OH)2D3 (+8.37&#x2005;pg/mL difference per effect allele; P = 4.93E-08) and was the only locus that achieved genome-wide significance in transethnic meta-analysis. rs114044709 adjacent to FAM20A, which encodes a protein required for biomineralization, was associated with change in PTH among Black participants (+20.32&#x2005;pg/mL difference per effect allele; P = 1.34E-08). In candidate analyses, single nucleotide polymorphisms within SULT2A1 and CYP24A1 had significant association (P < .05&#xf7;36 = .0014) with the changes in 1,25(OH)2D3 and PTH, respectively. CONCLUSION: Our results reveal potential new pathways of vitamin D regulation that require replication in other vitamin D trials.

Humans

Rare genetic variant risks in patients with sepsis-associated acute respiratory distress syndrome.

BACKGROUND: Acute respiratory distress syndrome (ARDS) is a complex, heterogeneous, and deadly condition often resulting from pulmonary lesions due to sepsis, among other causes. There is a lack of targeted therapies to specifically treat the patients. Common genetic factors in the population (frequency&#x2009;>&#x2009;1%) have been associated with ARDS susceptibility, but systematic genetic screens of the role of rare genetic variants are lacking. We used the network of known molecular interactions to identify ARDS risks from clusters of biologically related genes containing qualifying variants (QVs) with frequency&#x2009;<&#x2009;1% likely affecting function. METHODS: We conducted whole-exome sequencing in sepsis patients from the GEN-SEP cohort (n&#x2009;=&#x2009;822, of which 272 developed ARDS). A network-based heterogeneity clustering algorithm was used to discover significant gene clusters (p&#x2009;<&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5). Gene-set enrichment analysis and logistic regression models aggregating QVs were used for cross-verification to confirm consistency and deepen understanding of the effect sizes of gene clusters. RESULTS: We identified 19 significant clusters (plowest&#x2009;=&#x2009;3.29&#x2009;&#xd7;&#x2009;10&#x2013;10), each containing an average of 102 genes (11.6% mean similarity). QVs in nine gene clusters were associated with sepsis-associated ARDS (plowest&#x2009;=&#x2009;1&#x2009;&#xd7;&#x2009;10&#x2013;5) but were not associated with 28-day survival. Clusters were enriched in several biological pathways, notably the Toll-like receptor cascades. CONCLUSIONS: These results support a marked genetic heterogeneity underlying ARDS susceptibility and the presence of rare risk variants involving multiple biological processes that are associated with sepsis outcomes. Particularly, they underscore the importance of rare variants in genes of the Toll-like receptor cascades in the risk for sepsis-associated ARDS.

Humans

Genetic variants reduced POPs-related colorectal cancer risk via altering miRNA binding affinity and m6A modification.

Exposure to persistent organic pollutants (POPs) may contribute to colorectal cancer risk, but the underlying mechanisms of crucial POPs exposure remain unclear. Hence, we systematically investigated the associations among POPs exposure, genetics and epigenetics and their effects on colorectal cancer. A case-control study was conducted in the Chinese population for detecting POPs levels. We measured the concentrations of 24 POPs in the plasma using gas chromatography-tandem mass spectrometry (GC-MS/MS) and evaluated the clinical significance of POPs by calculating the area under the receiver operating characteristic curve (AUC). To assess the associations between candidate genetic variants and colorectal cancer risk, unconditional logistic regression was used. Compared with healthy control individuals, individuals with colorectal cancer exhibited higher concentrations of the majority of POPs. Exposure to PCB153 was positively associated with colorectal cancer risk, and PCB153 demonstrated superior accuracy (AUC=0.72) for predicting colorectal cancer compared to other analytes. On PCB153-related genes, the rs67734009 C allele was significantly associated with reduced colorectal cancer risk and lower plasma levels of PCB153. Moreover, rs67734009 exhibited an expression quantitative trait locus (eQTL) effect on ESR1, of which the expression level was negatively related to PCB153 concentration. Mechanistically, the risk allele of rs67734009 increased ESR1 expression via miR-3492 binding and m6A modification. Collectively, this study sheds light on potential genetic and epigenetic mechanisms linking PCB153 exposure and colorectal cancer risk, thereby providing insight into the accurate protection against POPs exposure.

Humans

Contributions of Common, Rare, and Somatic Genetic Variants to Incidence of Atrial Fibrillation.

IMPORTANCE: Atrial fibrillation (AF) has a complex genetic architecture involving common, rare, and somatic variants. The association between these components requires further investigation. OBJECTIVE: To examine the individual and combined contributions of polygenic, monogenic, and somatic genetic variants to AF incidence, and develop an integrated genomic model (IGM-AF) for improved risk prediction. DESIGN, SETTING, AND PARTICIPANTS: This cohort study used whole-genome sequence data from participants of the UK Biobank, with follow-up for AF events through hospital records, death registries, and self-report. The UK Biobank recruited participants aged 40 to 69 years in the UK between 2006 and 2010. Study data were analyzed from August 2022 to November 2024. EXPOSURES: IGM-AF comprising an AF polygenic risk score (PRS), a composite rare variant gene set (AFgeneset), and somatic variants associated with clonal hematopoiesis of indeterminate potential (CHIP). Clinical AF risk was estimated using the Cohorts for Heart and Aging Research in Genomic Epidemiology AF (CHARGE-AF) score. MAIN OUTCOMES AND MEASURES: The primary outcome was hazard ratios (HRs) for 5-year incident AF attributable to PRS, AFgeneset, CHIP, and their interactions. The predictive performance of IGM-AF and its components was quantified using HRs, C statistics, and reclassification indices. RESULTS: A total of 416&#x202f;085 individuals (mean [SD] age, 56.6 [8.0] years; 224&#x202f;642 female [54.0%]) with 30&#x202f;797 AF cases were included. The PRS (HR per 1 SD, 1.65; 95% CI, 1.63-1.67; P&#x2009;<&#x2009;1&#x2009;&#xd7;&#x2009;10-8), AFgeneset (HR, 1.63; 95% CI, 1.52-1.75; P&#x2009;=&#x2009;1.46&#x2009;&#xd7;&#x2009;10-42), and CHIP (HR, 1.26; 95% CI, 1.15-1.38; P&#x2009;=&#x2009;1.41&#x2009;&#xd7;&#x2009;10-6) were associated with incident AF. The 5-year cumulative incidence of AF was at least 2-fold among individuals having all 3 genetic drivers (common, rare, and somatic drivers) compared with those with only 1 driver. Integration of IGM-AF with a clinical risk model (CHARGE-AF) showed higher predictive performance (C statistic, 0.80; 95% CI, 0.80-0.80) compared with IGM-AF and CHARGE-AF alone. The classification of the at-risk population for AF was improved when IGM-AF was added to CHARGE-AF (net reclassification index, 0.08; 95% CI, 0.07-0.09). CONCLUSIONS AND RELEVANCE: Results of this cohort study demonstrated the complementary value of common, rare, and somatic variants in shaping genomic AF risk. Leveraging comprehensive genetic information may enhance screening and preventive interventions for AF.

Humans

Structural comparison of glycophorins and immunochemical analysis of genetic variants.

Differences in amino acid sequence of erythrocyte membrane glycophorin A are correlated with M or N blood group activity. A second sialoglycoprotein, glycophorin B, has an amino acid sequence identical to that of glycophorin AN in the first 23 positions and carries N activity only, suggesting that different structural genes code for the glycoproteins carrying these antigens. Certain genetically variant cells lack glycophorin A, as determined by immunochemical methods, and serological MN activity. Other variants lack MN activity, but contain normal amounts of glycophorin A in the membrane.

Amino Acid Sequence

The effect of neuraminidase on genetic variants of alpha anitrypsin.

Sera of Pi types M, F, S, Z, IM, FM, MS, and MZ were incubated with neuraminidase and the reaction products followed by electrophoresis. The alpha1 antitrypsin components showed a series of changes in mobility as sialic residues were removed. Removal of sialic acid was confirmed by chemical assay. Results of studies with two different electrophoretic systems suggested that the Z type alpha1 antitrypsin has less sialic acid than the M, F, and S types. There was no evidence that other genetic variants have a reduced sialic acid content. The two major bands of alpha1 antitrypsin seen in certain electrophoretic systems may reflect a difference of one sialic acid residue. It is proposed that the Z protein lacks a carbohydrate chain with two terminal sialic acid residues. This carbohydrate deficiency results in lack of secretion of type Z alpha1 antitrypsin from the endoplasmic reticulum, perhaps because of binding to sites specific for the incomplete glycoprotein or because of aggregation of the Z asialo protein. A carbohydrate chain could be prevented from attaching to the Z type either because of a conformational change or because of the replacement of a carbohydrate-binding asparagine residue in the Z protein.

Alleles

Purification and characterization of rat alpha-lactalbumins: apparent genetic variants.

Rat alpha-lactalbumin, from the milk of Fischer 344 (CDF) rats, was isolated and purified by a combination of gel filtration and diethylaminoethyl-cellulose ion exchange chromatography. Three electrophoretically distinct proteins had alpha-lactalbumin activity. Staining for carbohydrate indicated that at least two of the three forms were glycoproteins. The low molecular weight protein fraction from the wheys of two additional strains of laboratory rat were compared to ascertain whether the composition of this fraction was common in the divergent strains. Outbred Wistar and Long-Evans dams yielded wheys containing up to six forms of alpha-lactalbumin. Either one or both of two groups of three alpha-lactalbumins were in a given milk sample. The two groups of three alpha-lactalbumins appear to represent two genetic variants upon which is imposed a polymorphic character. All forms of alpha-lactalbumin, within and between strains, were immunologically identical.

Animals