PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Genetic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, “DESeq” and “Limma” R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology

Polycystic ovarian syndrome (PCOS) and recurrent spontaneous abortion (RSA) are associated with the PI3K-AKT pathway activation.

AIMS: We aimed to elucidate the mechanism leading to polycystic ovarian syndrome (PCOS) and recurrent spontaneous abortion (RSA). BACKGROUND: PCOS is an endocrine disorder. Patients with RSA also have a high incidence rate of PCOS, implying that PCOS and RSA may share the same pathological mechanism. OBJECTIVE: The single-cell RNA-seq datasets of PCOS (GSE168404 and GSE193123) and RSA GSE113790 and GSE178535) were downloaded from the Gene Expression Omnibus (GEO) database. METHODS: Datasets of PSCO and RSA patients were retrieved from the Gene Expression Omnibus (GEO) database. The "WGCNA" package was used to determine the module eigengenes associated with the PCOS and RSA phenotypes and the gene functions were analyzed using the "DAVID" database. The GSEA analysis was performed in "clusterProfiler" package, and key genes in the activated pathways were identified using the Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis. Real-time quantitative PCR (RT-qPCR) was conducted to determine the mRNA level. Cell viability and apoptosis were measured by cell counting kit-8 (CCK-8) and flow cytometry, respectively. RESULTS: The modules related to PCOS and RSA were sectioned by weighted gene co-expression network analysis (WGCNA) and positive correlation modules of PCOS and RSA were all enriched in angiogenesis and Wnt pathways. The GSEA further revealed that these biological processes of angiogenesis, Wnt and regulation of cell cycle were significantly positively correlated with the PCOS and RSA phenotypes. The intersection of the positive correlation modules of PCOS and RSA contained 80 key genes, which were mainly enriched in kinase-related signal pathways and were significant high-expressed in the disease samples. Subsequently, visualization of these genes including PDGFC, GHR, PRLR and ITGA3 showed that these genes were associated with the PI3K-AKT signal pathway. Moreover, the experimental results showed that PRLR had a higher expression in KGN cells, and that knocking PRLR down suppressed cell viability and promoted apoptosis of KGN cells. CONCLUSION: This study revealed the common pathological mechanisms between PCOS and RSA and explored the role of the PI3K-AKT signaling pathway in the two diseases, providing a new direction for the clinical treatment of PCOS and RSA.

Humans

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

Imprecision medicine: Systematic gaps in reporting variants of uncertain significance (VUS) and their reclassifications.

PURPOSE: Variants of uncertain significance (VUS) are frequently encountered during clinical genetic testing. To explore the clinical burden of VUS, we developed the Brotman Baty Institute Clinical Variant Database, which is an electronic health record (EHR)-linked database of clinical germline genetic variant information from patients with rare genetic disorders seen at 2 tertiary academic medical centers. METHODS: We retrospectively reviewed EHRs and genetic testing reports from 5158 patients seen across diverse adult genetics practices at these institutions from 2015 to 2024. We also compared these EHR-based variant classifications with those in ClinVar. RESULTS: The number of reported VUS relative to pathogenic or likely pathogenic variants can vary by over 14-fold depending on the primary indication for genetic testing and 3-fold depending on self-reported race. Furthermore, at least 1.6% of variant classifications used in the EHR for clinical care are outdated based on ClinVar variant classifications, including 26 instances in which the testing lab updated ClinVar, but the reclassification was never communicated to the patient. CONCLUSION: Our findings reveal that the clinical burden of VUS in adult medical genetics is unequally distributed across patients. We also highlight a deficiency in existing systems for communicating variant reclassifications to ClinVar, patients, and providers.

Humans

A microcomputer-based relational database for an academic department of human genetics.

Recent advances in microcomputer technology have made it possible for academic departments to establish their own discrete databases. This is in keeping with the modern tendency towards distributed processing on smaller systems as opposed to depending on large shared remote centralized mainframes. A database that had been implemented on a mainframe for the Department of Human Genetics of the University of Cape Town has been successfully transferred to a microcomputer system, resulting in a redesigned relational system with several significant advantages. These include faster data capture, enhanced consistency, greater computer awareness, improved economy and increased confidentiality.

Computers

Biomedical database inter-connectivity: an experiment linking MIM, GENBANK, and META-1 via MEDLINE.

The linkage of disparate biomedical databases is an important goal of the Unified Medical Language (UMLS) Project. We conducted an experiment to investigate the feasibility of using UMLS resources to link databases in clinical genetics and molecular biology. References from MIM ("Mendelian Inheritance in Man") were lexically mapped to the equivalent citations in MEDLINE. The MeSH major subject headings by which the citations in a particular MIM entry had been indexed were used to develop a "genetic-disorder-centered view of the world" in Meta-1 (the first official version of the UMLS Metathesaurus). Our hypothesis was that these MeSH subject headings could provide access to a "semantic neighborhood" in Meta-1 that would be relevant to a particular genetic disorder. By browsing in this "semantic neighborhood," a user could select various combinations of terms with which to search MEDLINE through an interface between Meta-1 and Grateful Med. Such searches might retrieve citations that were more recent than those in MIM or that provided useful supplementary information. Since some MEDLINE records contain pointers to entries in GENBANK, information about genetic sequences related to a particular clinical genetic disorder could also be retrieved. This scenario was implemented for a small number of MIM entries, providing a concrete demonstration that linking disparate electronic databases in an important subdomain of biomedicine is relatively straightforward.

Databases, Factual

Use of a microcomputer database system in a statewide effort for data collection in medical genetics.

The Genetics Office Automation System (GOAS) is a database management system for the collection and reporting of medical genetics data. We have previously reported on its implementation in a single university center [1,2]. We report here on its implementation in a coordinated data collection effort for the State of Missouri. We discuss the current status of the data collection activities and procedures to share data collected at an individual center with state, regional, and national data collection efforts.

Data Collection

A novel, highly conserved structural motif is present in all members of the steroid receptor superfamily.

Steroid and thyroid hormone receptor superfamily members are ligand potentiated transcription factors. Recent evidence indicates that one aspect of steroid receptor action is an interaction with other trans-acting factors, such as the glucocorticoid receptor with the AP1 transcription factor, for example. Using a structural approach to identify domains of the glucocorticoid receptor responsible for interactions with affiliated transacting factors and DNA, we have identified a putative helix-turn-zipper motif that is conserved in all steroid, thyroid hormone, retinoic acid, and vitamin-D3 receptors. This structural motif is also conserved among new members of the family, the peroxisome proliferator-activated receptors and the retinoid-X receptors. This structural domain is characterized by a pair of amino acids (I,L,V)P that is conserved in all superfamily members. Additional characteristics include six heptad repeats of hydrophobic amino acids, four of which form a canonical leucine zipper in the rat glucocorticoid receptor. Although this leucine repeat is not absolutely conserved among superfamily members, the periodicity of hydrophobic residues is conserved throughout. Based on sequence analyses from the GenEMBL and SwissProt databases using the Genetics Computer Group and MacVector sequence analysis software packages, and the Brookhaven structural database, we present evidence for a novel structural domain, a helix-turn-zipper that is conserved in all superfamily members, and may function in transactivation of cognate genes.

Amino Acid Sequence

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements

Current status of the Gene-Tox Program.

The U.S. Environmental Protection Agency's Gene-Tox Program is a multiphased effort to review and evaluate the existing literature in assay systems available in the field of genetic toxicology. The first phase of the Gene-Tox Program selected assay systems for evaluation, generated expert panel reviews of the data from the scientific literature, and recommended testing protocols for the systems. Phase II established and evaluated the database of chemical genetic toxicity data for its relevance to identifying human health hazards. The ongoing phase III continues reviewing and updating chemical data in selected assay systems. Currently, data exist on over 4000 chemicals in 27 assay systems; two additional assay systems will be included in phase III. The review data are published in the scientific literature and are also publicly available through the National Library of Medicine TOXNET system. The review and analysis components of Gene-Tox comprise 45 published papers, and several others are in preparation. Differences that have been observed between Gene-Tox and National Toxicology Program databases relative to the sensitivity, specificity, accuracy, and predictivity of genetic toxicity data compared to carcinogenesis data are ascribable to differences between the two databases in chemical selection criteria, testing protocols, and chemical class distributions.

Animals

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics

Palaeoproteomic Deconvolution of Physical and Genetic Collagen Mixtures.

Species identification in palaeoproteomics relies on genome-derived protein sequences which are often poor-quality, and lacks tools to cope with multi-species samples. Here, we address both challenges through the analysis of "physical and genetic mixtures". Species that are absent from our database are considered a "genetic mixture", i.e. a patchwork of peptides from closely related species. Inversely, various overlapping peptide stretches allow us to resolve complex "physical mixtures". This is benchmarked by analysing physical mixtures of modern bone fragments, including genetic mixtures. We illustrate the impact of our approach via a rapid and high-throughput analysis of >2500 bone fragments, revealing the Eemian-era faunal environment around Scladina Cave, including the first Palaeoloxodon antiquus identified at this site.

bioarchaeology

A current genotoxicity database for heterocyclic thermic food mutagens. I. Genetically relevant endpoints.

Cooking, heat processing, or pyrolysis of protein-rich foods induce the formation of a series of structurally related heterocyclic aromatic bases that have been found to be mutagens. The primary genetic assay utilized to detect and isolate these mutagens has been the his reversion assay in Salmonella typhimurium. The classification and nomenclature of these chemicals is revised to reflect recent advances. The findings of short-term tests for genetic injury that have been applied to these agents are presented in a systematic way. Cell-free, bacterial, mammalian cell culture, and in vivo systems are included. Major results, the mutagens tested, and key references are presented in tabular form, with text commentary. Integrated conclusions on the state of current knowledge of the genetic toxicity of thermic food mutagens are presented. Areas in need of further research are defined. Finally, an outline is presented of a suggested path leading to the determination whether normal methods of food preparation and processing constitute a human health hazard.

Animals

Unique signatures of highly constrained genes across publicly available genomic databases.

PURPOSE: Publicly available genomic databases are critical in understanding human genetic variation. They also provide unique insights into patterns of genetic constraints and their relationship with human disease. METHODS: We utilized one of the largest publicly available databases, Genome Aggregate Database, to determine genes that are highly constrained for only loss-of-function, only missense, and both loss-of-function/missense variants. We identified their unique signatures and explored their causal relationship with human diseases. Those genes were also evaluated for chromosomal location, tissue-level expression, Gene Ontology analysis, and gene family categorization using multiple publicly available databases. RESULTS: We identified unique patterns of inheritance, protein size, and enrichment in distinct molecular pathways for those constrained genes associated with human disease. In addition, we identified genes that are currently not known to cause human disease, which may be excellent gene discovery candidates. CONCLUSION: We elucidate biological pathways of highly constrained genes that expand our understanding of critical cellular proteins. The findings can also advance research in rare diseases.

Humans

Efficacy of alternative multivariate best linear unbiased prediction models for genetic evaluation of swine.

A comparison of the accuracy of alternative BLUP evaluations of swine performance data is reported. A simulated data set of performance for days to market, backfat, number born alive, number weaned, and litter weight in six herds was used for the evaluation. The data structure was derived by using the performance records of six herds sampled from the American Yorkshire Club's STAGES (Swine Testing and Genetic Evaluation System) database. For growth traits, 10,360 pig records from 129 sires and 897 dams were recorded. For maternal traits, records on 2,598 litters of 1,209 sows from 147 sires and 585 dams were included. The actual observed performance of each record was removed and replaced with simulated performance. These simulated data were analyzed by within- and across-herd BLUP models using STAGES and Animal Model (AM) procedures. Results indicate that the alternative BLUP procedures produced similar estimates. Correlations between STAGES and AM EPD ranged from .84 to .95. Correlations between STAGES EPD and true genetic value (G) ranged from .41 to .74, and correlations between AM EPD and G ranged from .40 to .74. On average, AM EPD had a .04 larger correlation with G than did STAGES EPD, although the difference in the correlations was not significant (P greater than .05). Trends in EPD for sires, dams, and pigs or sows were the same. Likewise, standard errors of prediction for AM EPD averaged 4% smaller than those for STAGES EPD. Computationally, the AM procedures used 15 to 20 times as much processing time as did STAGES procedures.

Animals