PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Genetic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

prot4EST: translating expressed sequence tags from neglected genomes.

BACKGROUND: The genomes of an increasing number of species are being investigated through generation of expressed sequence tags (ESTs). However, ESTs are prone to sequencing errors and typically define incomplete transcripts, making downstream annotation difficult. Annotation would be greatly improved with robust polypeptide translations. Many current solutions for EST translation require a large number of full-length gene sequences for training purposes, a resource that is not available for the majority of EST projects. RESULTS: As part of our ongoing EST programs investigating these "neglected" genomes, we have developed a polypeptide prediction pipeline, prot4EST. It incorporates freely available software to produce final translations that are more accurate than those derived from any single method. We show that this integrated approach goes a long way to overcoming the deficit in training data. CONCLUSIONS: prot4EST provides a portable EST translation solution and can be usefully applied to >95% of EST projects to improve downstream annotation. It is freely available from http://www.nematodes.org/PartiGene.

Animals

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals

Bioinformatics identification and validation of pyroptosis-related gene for ischemic stroke.

BACKGROUND: Ischemic stroke (IS) is one of the common and frequent diseases with extremely high lethality and disability in the world, and there is no effective treatment at present. This study aimed to screen hub genes involved in cerebral ischemia/reperfusion injury (CIRI) and pyroptosis, and explore promising intervention targets. METHODS: CIRI-related genes (GSE202659 and GSE131193) and pyroptosis-related genes (PRGs) in mice were obtained from the Gene Expression Omnibus (GEO) and GeneCards database. We screened for LASSO regression to construct a prognostic model of GSE131193 and PRGs and examined by GSE137482. The functional enrichment analysis of Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Gene Set Enrichment Analysis (GSEA) and Gene Set Variation Analysis (GSVA) were performed on pyroptosis-related differentially expressed genes (PRDEGs) of GSE202659.The key modules for CIRI and pyroptosis were identified by Weight Gene Co-expression Network Analysis (WGCNA). Subsequently, Protein-protein Interaction (PPI) network and the Cytoscape was constructed to screen out hub genes. Used the starBase to predict miRNA interacting with hub genes and constructed mRNA-miRNA-lncRNA interaction networks. CIRI-related Molecular Subtypes were constructed for hub genes. The relationship between immune cells and hub genes was verified via CIBERSORT. Finally, we selected C57BL/6 mice to construct models to confirm hub genes by enzyme linked immunosorbent assay (ELISA), reverse transcription-polymerase chain reaction (RT-PCR), western blot, and Immunofluorescence. RESULTS: A total of 272 PRGs and 35 PRDEGs were screened. An eight-gene risk prediction models were established (AUC = 0.868). GO, KEGG, GSEA and GSVA analyses revealed that PRDEGs were mainly involved in positive regulation of cytokine production, and NOD-like receptor signaling pathway. And then, seven hub genes (Irf1, Icam1, Tlr2, Tnf, Cebpb, Il1rn, and Casp8) were identified by PPI. Icam1, Tnf, Cebpb, Il1rn, and Casp8 had high expression profiles in Cluster2 by hierarchical clustering. The immune infiltration analysis results showed that among the hub genes, Cebpb, Il1rn, and Casp8, showed a significant positive correlation with the degree of NK.Actived, and Icam1 showed a significant negative correlation with B.Cells.Memory. The results of animal experiments significantly demonstrated an upregulation of Irf1, Icam1, Tlr2, Cebpb, and Il1rn. CONCLUSION: Our finding indicated that Irf1, Icam1, Tlr2, Cebpb, and Il1rn are hub genes associated with pyroptosis, and these genes are all associated with different immune cells, so as to provide new targets for the prevention and treatment of IS from the perspective of pyroptosis.

Pyroptosis

The Aggregated Gut Viral Catalogue (AVrC): A unified resource for exploring the viral diversity of the human gut.

The growing interest in the role of the gut virome in human health and disease, has led to several recent large-scale viral catalogue projects mining human gut metagenomes each using varied computational tools and quality control criteria. Importantly, there has been to date no consistent comparison of these catalogues' quality, diversity, and overlap. In this project, we therefore systematically surveyed nine previously published human gut viral catalogues. While these catalogues collectively screened >40,000 human fecal metagenomes, 82% of the recovered 345,613 viral sequences were unique to one catalogue, highlighting limited redundancy between the ressources and suggesting the need for an aggregated resource bringing these viral sequences together. We further expanded these viral catalogues by mining 7,867 infant gut metagenomes from 12 large-scale infant studies collected in 9 different countries. From these datasets, we constructed the Aggregated Gut Viral Catalogue (AVrC), a unified modular resource containing 1,018,941 dereplicated viral sequences (449,859 species-level vOTUs). Using computational inference tools, annotations were obtained for each vOTU representative sequence quality, viral taxonomy, predicted viral lifestyle, and putative host. This project aims to facilitate the reuse of previously published viral catalogues by the research community and follows a modular framework to enable future expansions as novel data becomes available.

Humans

GeneCOCOA: Detecting context-specific functions of individual genes using co-expression data.

Extraction of meaningful biological insight from gene expression profiling often focuses on the identification of statistically enriched terms or pathways. These methods typically use gene sets as input data, and subsequently return overrepresented terms along with associated statistics describing their enrichment. This approach does not cater to analyses focused on a single gene-of-interest, particularly when the gene lacks prior functional characterization. To address this, we formulated GeneCOCOA, a method which utilizes context-specific gene co-expression and curated functional gene sets, but focuses on a user-supplied gene-of-interest (GOI). The co-expression between the GOI and subsets of genes from functional groups (e.g. pathways, GO terms) is derived using linear regression, and resulting root-mean-square error values are compared against background values obtained from randomly selected genes. The resulting p values provide a statistical ranking of functional gene sets from any collection, along with their associated terms, based on their co-expression with the gene of interest in a manner specific to the context and experiment. GeneCOCOA thereby provides biological insight into both gene function, and putative regulatory mechanisms by which the expression of the GOI is controlled. Despite its relative simplicity, GeneCOCOA outperforms similar methods in the accurate recall of known gene-disease associations. We furthermore include a differential GeneCOCOA mode, thus presenting the first implementation of a gene-focused approach to experiment-specific gene set enrichment analysis. GeneCOCOA is formulated as an R package for ease-of-use, available at https://github.com/si-ze/geneCOCOA.

Gene Expression Profiling

TIPP3 and TIPP3-fast: Improved abundance profiling in metagenomics.

We present TIPP3 and TIPP3-fast, new tools for abundance profiling in metagenomic datasets. Like its predecessor, TIPP2, the TIPP3 pipeline uses a maximum likelihood approach to place reads into labeled taxonomies using marker genes, but it achieves superior accuracy to TIPP2 by enabling the use of much larger taxonomies through improved algorithmic techniques. We show that TIPP3 is generally more accurate than leading methods for abundance profiling in two important contexts: when reads come from genomes not already in a public database (i.e., novel genomes) and when reads contain sequencing errors. We also show that TIPP3-fast has slightly lower accuracy than TIPP3, but is also generally more accurate than other leading methods and uses a small fraction of TIPP3's runtime. Additionally, we highlight the potential benefits of restricting abundance profiling methods to those reads that map to marker genes (i.e., using a filtered marker-gene based analysis), which we show typically improves accuracy. TIPP3 is freely available at https://github.com/c5shen/TIPP3.

Metagenomics

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models

PoweREST: Statistical power estimation for spatial transcriptomics experiments to detect differentially expressed genes between two conditions.

Recent advancements in spatial transcriptomics (ST) have significantly enhanced biological research in various domains. However, the high cost for current ST data generation techniques restricts the large-scale application of ST. Consequently, maximization of the use of available resources to achieve robust statistical power for ST data is a pressing need. One fundamental question in ST analysis is detection of differentially expressed genes (DEGs) under different conditions using ST data. Such DEG analyses are performed frequently, but their power calculations are rarely discussed in the literature. To address this gap, we developed PoweREST, a power estimation tool designed to support the power calculation for DEG detection with 10X Genomics Visium data. PoweREST enables power estimation both before any ST experiments and after preliminary data are collected, making it suitable for a wide variety of power analyses in ST studies. We also provide a user-friendly, program-free web application that allows users to interactively calculate and visualize study power along with relevant parameters.

Gene Expression Profiling

Assessment of Gene Set Enrichment Analysis using curated RNA-seq-based benchmarks.

Pathway enrichment analysis is a ubiquitous computational biology method to interpret a list of genes (typically derived from the association of large-scale omics data with phenotypes of interest) in terms of higher-level, predefined gene sets that share biological function, chromosomal location, or other common features. Among many tools developed so far, Gene Set Enrichment Analysis (GSEA) stands out as one of the pioneering and most widely used methods. Although originally developed for microarray data, GSEA is nowadays extensively utilized for RNA-seq data analysis. Here, we quantitatively assessed the performance of a variety of GSEA modalities and provide guidance in the practical use of GSEA in RNA-seq experiments. We leveraged harmonized RNA-seq datasets available from The Cancer Genome Atlas (TCGA) in combination with large, curated pathway collections from the Molecular Signatures Database to obtain cancer-type-specific target pathway lists across multiple cancer types. We carried out a detailed analysis of GSEA performance using both gene-set and phenotype permutations combined with four different choices for the Kolmogorov-Smirnov enrichment statistic. Based on our benchmarks, we conclude that the classic/unweighted gene-set permutation approach offered comparable or better sensitivity-vs-specificity tradeoffs across cancer types compared with other, more complex and computationally intensive permutation methods. Finally, we analyzed other large cohorts for thyroid cancer and hepatocellular carcinoma. We utilized a new consensus metric, the Enrichment Evidence Score (EES), which showed a remarkable agreement between pathways identified in TCGA and those from other sources, despite differences in cancer etiology. This finding suggests an EES-based strategy to identify a core set of pathways that may be complemented by an expanded set of pathways for downstream exploratory analysis. This work fills the existing gap in current guidelines and benchmarks for the use of GSEA with RNA-seq data and provides a framework to enable detailed benchmarking of other RNA-seq-based pathway analysis tools.

Humans

Identification of Critical Genes Related to Breast Cancer with Brain Metastasis Through Bioinformatics Analysis.

INTRODUCTION: Distant metastasis accounts for the majority of Breast Cancer (BC)-related mortality. The brain is one of the most common regions of metastasis. However, the underlying molecular mechanisms remain uncertain. METHODS: In this study, gene expression profiles were downloaded from the Gene Expression Omnibus (GEO) database. Datasets GSE100534 and GSE52604, containing 16 primary brain tumor samples and 38 breast cancer brain metastasis samples, were used to identify the Differentially Expressed Genes (DEGs). The Metascape database was used to analyze enriched Gene Ontology (GO) entries and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway entries in DEGs. The STRING database was then used to construct a Protein-Protein Interaction (PPI) network, and the Cytoscape platform was employed to visualize the network. Furthermore, the Kaplan-Meier curve was used to analyze the Relapse-Free Survival (RFS) among the hub genes. Finally, the iRegulon plugin was used to construct a regulatory network to find the transcription factors (TFs) that regulate the expression of the hub genes. RESULTS: A total of 344 DEGs, including 182 up-regulated and 162 down-regulated genes, were identified by using the limma package in R. A module with 18 nodes and 9 hub genes was selected from the PPI network by using the plugins MCODE and Cyto- Hubba, respectively. KEGG pathway analysis demonstrated that brain metastasis in BC was closely related to the oocyte cell cycle. The Kaplan-Meier curve showed that high expression of these 9 hub genes was associated with poor RFS in BC patients. TFs' analysis showed that E2F4, SIN3A, FOXM1, and TFDP1 interacted with these hub genes. DISCUSSION: This study revealed that Breast Cancer Brain Metastasis (BCBM) may have a promoting effect on the cell cycle of oocytes and affect the maturation and division of oocytes through the KEGG and GO analyses of 344 DEGs. The selected 9 hub genes (ASPM, BUB1, BUB1B, CCNA2, CCNB1, CDK1, NDC80, NCAPG, and TOP2A) and 4 transcription factors (E2F4, SIN3A, FOXM1, TFDP1) may play a critical role in brain metastasis of BC. CONCLUSION: The results of this study may aid in the early diagnosis and suggest potential targets for the treatment of BCBM.

Brain Neoplasms

Identification of NR4A2 as a Potential Predictive Biomarker for Atherosclerosis.

INTRODUCTION/OBJECTIVE: Atherosclerosis, a leading cause of death globally, is characterized by the buildup of immune cells and lipids in medium to large-sized arteries. However, its precise mechanism remains unclear. The purpose of this study is to explore innovative and reliable biomarkers as a viable approach for the identification and management of atherosclerosis. METHODS: The atherosclerosis-related datasets GSE100927 and GSE66360 were retrieved from the Gene Expression Omnibus (GEO) database. The Limma package in the R programming language was utilized, applying the criteria of |logFC| > 1 and P < 0.05. Subsequently, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the 127 identified DEGs using R. Machine learning techniques were then applied to these data to explore and pinpoint potential biomarkers. The diagnostic potential of these markers was assessed via Receiver Operating Characteristic (ROC) curve analysis. Finally, western blot, real-time quantitative PCR (qRT-PCR), and immunohistochemistry (IHC) were employed to confirm the key biomarkers. RESULTS: Our research indicated that a total of 127 DEGs linked to atherosclerosis were successfully identified. Through the application of machine learning methods, eight critical genes were highlighted. Among these, Nuclear Receptor Subfamily 4 Group A Member-2 (NR4A2) emerged as the most promising marker for further investigation. CIBERSORT analysis revealed that NR4A2 expression levels were significantly correlated with multiple immune cell types, including B cells, plasma cells, and macrophages. Additional validation experiments confirmed that NR4A2 expression was indeed elevated in atherosclerotic plaques, supporting its potential as a biomarker for atherosclerosis. CONCLUSION: Our study identified NR4A2 as a potential immune-related biomarker for the diagnosis and treatment of atherosclerosis.

Atherosclerosis

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, &#x201c;DESeq&#x201d; and &#x201c;Limma&#x201d; R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology

Polycystic ovarian syndrome (PCOS) and recurrent spontaneous abortion (RSA) are associated with the PI3K-AKT pathway activation.

AIMS: We aimed to elucidate the mechanism leading to polycystic ovarian syndrome (PCOS) and recurrent spontaneous abortion (RSA). BACKGROUND: PCOS is an endocrine disorder. Patients with RSA also have a high incidence rate of PCOS, implying that PCOS and RSA may share the same pathological mechanism. OBJECTIVE: The single-cell RNA-seq datasets of PCOS (GSE168404 and GSE193123) and RSA GSE113790 and GSE178535) were downloaded from the Gene Expression Omnibus (GEO) database. METHODS: Datasets of PSCO and RSA patients were retrieved from the Gene Expression Omnibus (GEO) database. The "WGCNA" package was used to determine the module eigengenes associated with the PCOS and RSA phenotypes and the gene functions were analyzed using the "DAVID" database. The GSEA analysis was performed in "clusterProfiler" package, and key genes in the activated pathways were identified using the Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis. Real-time quantitative PCR (RT-qPCR) was conducted to determine the mRNA level. Cell viability and apoptosis were measured by cell counting kit-8 (CCK-8) and flow cytometry, respectively. RESULTS: The modules related to PCOS and RSA were sectioned by weighted gene co-expression network analysis (WGCNA) and positive correlation modules of PCOS and RSA were all enriched in angiogenesis and Wnt pathways. The GSEA further revealed that these biological processes of angiogenesis, Wnt and regulation of cell cycle were significantly positively correlated with the PCOS and RSA phenotypes. The intersection of the positive correlation modules of PCOS and RSA contained 80 key genes, which were mainly enriched in kinase-related signal pathways and were significant high-expressed in the disease samples. Subsequently, visualization of these genes including PDGFC, GHR, PRLR and ITGA3 showed that these genes were associated with the PI3K-AKT signal pathway. Moreover, the experimental results showed that PRLR had a higher expression in KGN cells, and that knocking PRLR down suppressed cell viability and promoted apoptosis of KGN cells. CONCLUSION: This study revealed the common pathological mechanisms between PCOS and RSA and explored the role of the PI3K-AKT signaling pathway in the two diseases, providing a new direction for the clinical treatment of PCOS and RSA.

Humans

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A&#xa0;key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

Imprecision medicine: Systematic gaps in reporting variants of uncertain significance (VUS) and their reclassifications.

PURPOSE: Variants of uncertain significance (VUS) are frequently encountered during clinical genetic testing. To explore the clinical burden of VUS, we developed the Brotman Baty Institute Clinical Variant Database, which is an electronic health record (EHR)-linked database of clinical germline genetic variant information from patients with rare genetic disorders seen at 2 tertiary academic medical centers. METHODS: We retrospectively reviewed EHRs and genetic testing reports from 5158 patients seen across diverse adult genetics practices at these institutions from 2015 to 2024. We also compared these EHR-based variant classifications with those in ClinVar. RESULTS: The number of reported VUS relative to pathogenic or likely pathogenic variants can vary by over 14-fold depending on the primary indication for genetic testing and 3-fold depending on self-reported race. Furthermore, at least 1.6% of variant classifications used in the EHR for clinical care are outdated based on ClinVar variant classifications, including 26 instances in which the testing lab updated ClinVar, but the reclassification was never communicated to the patient. CONCLUSION: Our findings reveal that the clinical burden of VUS in adult medical genetics is unequally distributed across patients. We also highlight a deficiency in existing systems for communicating variant reclassifications to ClinVar, patients, and providers.

Humans