PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics↗

A mathematical and computational framework for quantitative comparison and integration of large-scale gene expression data.

Analysis of large-scale gene expression studies usually begins with gene clustering. A ubiquitous problem is that different algorithms applied to the same data inevitably give different results, and the differences are often substantial, involving a quarter or more of the genes analyzed. This raises a series of important but nettlesome questions: How are different clustering results related to each other and to the underlying data structure? Is one clustering objectively superior to another? Which differences, if any, are likely candidates to be biologically important? A systematic and quantitative way to address these questions is needed, together with an effective way to integrate and leverage expression results with other kinds of large-scale data and annotations. We developed a mathematical and computational framework to help quantify, compare, visualize and interactively mine clusterings. We show that by coupling confusion matrices with appropriate metrics (linear assignment and normalized mutual information scores), one can quantify and map differences between clusterings. A version of receiver operator characteristic analysis proved effective for quantifying and visualizing cluster quality and overlap. These methods, plus a flexible library of clustering algorithms, can be called from a new expandable set of software tools called CompClust 1.0 (http://woldlab.caltech.edu/compClust/). CompClust also makes it possible to relate expression clustering patterns to DNA sequence motif occurrences, protein-DNA interaction measurements and various kinds of functional annotations. Test analyses used yeast cell cycle data and revealed data structure not obvious under all algorithms. These results were then integrated with transcription motif and global protein-DNA interaction data to identify G1 regulatory modules.

Algorithms↗

Whole genome sequence-based association analysis of African American individuals with bipolar disorder and schizophrenia.

In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole genome sequencing of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls. To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of BD association with single-variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD GWAS loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.

Journal Article↗

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals↗

An evaluation of GO annotation retrieval for BioCreAtIvE and GOA.

BACKGROUND: The Gene Ontology Annotation (GOA) database http://www.ebi.ac.uk/GOA aims to provide high-quality supplementary GO annotation to proteins in the UniProt Knowledgebase. Like many other biological databases, GOA gathers much of its content from the careful manual curation of literature. However, as both the volume of literature and of proteins requiring characterization increases, the manual processing capability can become overloaded. Consequently, semi-automated aids are often employed to expedite the curation process. Traditionally, electronic techniques in GOA depend largely on exploiting the knowledge in existing resources such as InterPro. However, in recent years, text mining has been hailed as a potentially useful tool to aid the curation process. To encourage the development of such tools, the GOA team at EBI agreed to take part in the functional annotation task of the BioCreAtIvE (Critical Assessment of Information Extraction systems in Biology) challenge. BioCreAtIvE task 2 was an experiment to test if automatically derived classification using information retrieval and extraction could assist expert biologists in the annotation of the GO vocabulary to the proteins in the UniProt Knowledgebase. GOA provided the training corpus of over 9000 manual GO annotations extracted from the literature. For the test set, we provided a corpus of 200 new Journal of Biological Chemistry articles used to annotate 286 human proteins with GO terms. A team of experts manually evaluated the results of 9 participating groups, each of which provided highlighted sentences to support their GO and protein annotation predictions. Here, we give a biological perspective on the evaluation, explain how we annotate GO using literature and offer some suggestions to improve the precision of future text-retrieval and extraction techniques. Finally, we provide the results of the first inter-annotator agreement study for manual GO curation, as well as an assessment of our current electronic GO annotation strategies. RESULTS: The GOA database currently extracts GO annotation from the literature with 91 to 100% precision, and at least 72% recall. This creates a particularly high threshold for text mining systems which in BioCreAtIvE task 2 (GO annotation extraction and retrieval) initial results precisely predicted GO terms only 10 to 20% of the time. CONCLUSION: Improvements in the performance and accuracy of text mining for GO terms should be expected in the next BioCreAtIvE challenge. In the meantime the manual and electronic GO annotation strategies already employed by GOA will provide high quality annotations.

Animals↗

A Protein Classification Benchmark collection for machine learning.

Protein classification by machine learning algorithms is now widely used in structural and functional annotation of proteins. The Protein Classification Benchmark collection (http://hydra.icgeb.trieste.it/benchmark) was created in order to provide standard datasets on which the performance of machine learning methods can be compared. It is primarily meant for method developers and users interested in comparing methods under standardized conditions. The collection contains datasets of sequences and structures, and each set is subdivided into positive/negative, training/test sets in several ways. There is a total of 6405 classification tasks, 3297 on protein sequences, 3095 on protein structures and 10 on protein coding regions in DNA. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences or structures, as well as various functional and taxonomic classification problems. In the case of hierarchical classification schemes, the classification tasks can be defined at various levels of the hierarchy (such as classes, folds, superfamilies, etc.). For each dataset there are distance matrices available that contain all vs. all comparison of the data, based on various sequence or structure comparison methods, as well as a set of classification performance measures computed with various classifier algorithms.

Algorithms↗

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization↗

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans↗

Assessing the precision of high-throughput computational and laboratory approaches for the genome-wide identification of protein subcellular localization in bacteria.

BACKGROUND: Identification of a bacterial protein's subcellular localization (SCL) is important for genome annotation, function prediction and drug or vaccine target identification. Subcellular fractionation techniques combined with recent proteomics technology permits the identification of large numbers of proteins from distinct bacterial compartments. However, the fractionation of a complex structure like the cell into several subcellular compartments is not a trivial task. Contamination from other compartments may occur, and some proteins may reside in multiple localizations. New computational methods have been reported over the past few years that now permit much more accurate, genome-wide analysis of the SCL of protein sequences deduced from genomes. There is a need to compare such computational methods with laboratory proteomics approaches to identify the most effective current approach for genome-wide localization characterization and annotation. RESULTS: In this study, ten subcellular proteome analyses of bacterial compartments were reviewed. PSORTb version 2.0 was used to computationally predict the localization of proteins reported in these publications, and these computational predictions were then compared to the localizations determined by the proteomics study. By using a combined approach, we were able to identify a number of contaminants and proteins with dual localizations, and were able to more accurately identify membrane subproteomes. Our results allowed us to estimate the precision level of laboratory subproteome studies and we show here that, on average, recent high-precision computational methods such as PSORTb now have a lower error rate than laboratory methods. CONCLUSION: We have performed the first focused comparison of genome-wide proteomic and computational methods for subcellular localization identification, and show that computational methods have now attained a level of precision that is exceeding that of high-throughput laboratory approaches. We note that analysis of all cellular fractions collectively is required to effectively provide localization information from laboratory studies, and we propose an overall approach to genome-wide subcellular localization characterization that capitalizes on the complementary nature of current laboratory and computational methods.

Bacterial Proteins↗

Dynamic transcriptome of mice.

Life science in the 21st century is developing rapidly through the structural analysis of biomolecules, the completion of the human genome sequence and the analysis of transcriptomes. The mouse transcriptome has been comprehensively analyzed using a gene discovery approach to collect full-length cDNA (FL-cDNA) clones. The framework of the transcriptome was then mapped out by an international Functional ANnoTation Of Mouse cDNA (FANTOM) effort, and a significant new population of noncoding transcripts was discovered. The geographical analogy of a second "RNA continent," separate from the "continent" of expressed proteins, aids the visualization of this concept. An unexpected number of variations was discovered in the mouse transcriptome. The animal transcriptome has evolved to produce several transcripts and proteins from a single "transcriptional unit". Transcriptome analysis has given rise to the FL-cDNA database and to the 60 770 FANTOM FL-cDNA clone set, and the DNABook was developed as an easier way to distribute these clones. In conjunction with genome sequence databases, transcriptome databases and clone banks will be platforms for developing advanced databases of gene function (e.g. the Genome Function Database). This will enable life science to make rapid progress towards understanding life as a system of molecules.

Animals↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗

MetaCyc and AraCyc. Metabolic pathway databases for plant research.

MetaCyc (http://metacyc.org) contains experimentally determined biochemical pathways to be used as a reference database for metabolism. In conjunction with the Pathway Tools software, MetaCyc can be used to computationally predict the metabolic pathway complement of an annotated genome. To increase the breadth of pathways and enzymes, more than 60 plant-specific pathways have been added or updated in MetaCyc recently. In contrast to MetaCyc, which contains metabolic data for a wide range of organisms, AraCyc is a species-specific database containing only enzymes and pathways found in the model plant Arabidopsis (Arabidopsis thaliana). AraCyc (http://arabidopsis.org/tools/aracyc/) was the first computationally predicted plant metabolism database derived from MetaCyc. Since its initial computational build, AraCyc has been under continued curation to enhance data quality and to increase breadth of pathway coverage. Twenty-eight pathways have been manually curated from the literature recently. Pathway predictions in AraCyc have also been recently updated with the latest functional annotations of Arabidopsis genes that use controlled vocabulary and literature evidence. AraCyc currently features 1,418 unique genes mapped onto 204 pathways with 1,156 literature citations. The Omics Viewer, a user data visualization and analysis tool, allows a list of genes, enzymes, or metabolites with experimental values to be painted on a diagram of the full pathway map of AraCyc. Other recent enhancements to both MetaCyc and AraCyc include implementation of an evidence ontology, which has been used to provide information on data quality, expansion of the secondary metabolism node of the pathway ontology to accommodate curation of secondary metabolic pathways, and enhancement of the cellular component ontology for storing and displaying enzyme and pathway locations within subcellular compartments.

4-Hydroxyphenylpyruvate Dioxygenase↗

Temporal and spatial transcriptional programs in murine kidney development.

We have performed a systematic temporal and spatial expression profiling of the developing mouse kidney using Compugen long-oligonucleotide microarrays. The activity of 18,000 genes was monitored at 24-h intervals from 10.5-day-postcoitum (dpc) metanephric mesenchyme (MM) through to neonatal kidney, and a cohort of 3,600 dynamically expressed genes was identified. Early metanephric development was further surveyed by directly comparing RNA from 10.5 vs. 11.5 vs. 13.5dpc kidneys. These data showed high concordance with the previously published dynamic profile of rat kidney development (Stuart RO, Bush KT, and Nigam SK. Proc Natl Acad Sci USA 98: 5649-5654, 2001) and our own temporal data. Cluster analyses were used to identify gene ontological terms, functional annotations, and pathways associated with temporal expression profiles. Genetic network analysis was also used to identify biological networks that have maximal transcriptional activity during early metanephric development, highlighting the involvement of proliferation and differentiation. Differential gene expression was validated using whole mount and section in situ hybridization of staged embryonic kidneys. Two spatial profiling experiments were also undertaken. MM (10.5dpc) was compared with adjacent intermediate mesenchyme to further define metanephric commitment. To define the genes involved in branching and in the induction of nephrogenesis, expression profiling was performed on ureteric bud (GFP+) FACS sorted from HoxB7-GFP transgenic mice at 15.5dpc vs. the GFP- mesenchymal derivatives. Comparisons between temporal and spatial data enhanced the ability to predict function for genes and networks. This study provides the most comprehensive temporal and spatial survey of kidney development to date, and the compilation of these transcriptional surveys provides important insights into metanephric development that can now be functionally tested.

Animals↗

ProbeLynx: a tool for updating the association of microarray probes to genes.

As genome sequence data and gene prediction improve, probes developed for a given microarray experiment should be continuously re-evaluated for their specificity for given genes. ProbeLynx(www.pathogenomics.ca/probelynx) is a new web service which uses current genomic sequence information to re-examine microarray probe specificity and provide annotation updates relevant to determining which gene(s) and transcript(s) are associated with a given probe. Probe sequences (either oligonucleotide- or cDNA-based) are uploaded in FASTA format and the results returned as a tab-delimited flat file for insertion into a spreadsheet application or database management system for further analysis. ProbeLynx has been initially developed to focus on arrays derived from human, mouse, chicken and bovine genomes, but may be expanded to handle other genomic datasets. ProbeLynx offers microarray users the important ability to continuously assess the potential of a probe to cross-hybridize to paralogous genes and the suitability of a given probe to investigate a transcript of interest. By also including the latest gene function annotation information in the output, ProbeLynx provides the critical first step in updating microarray data annotation.

Animals↗

Mouse models of Down syndrome: how useful can they be? Comparison of the gene content of human chromosome 21 with orthologous mouse genomic regions.

With an incidence of approximately 1 in 700 live births, Down syndrome (DS) remains the most common genetic cause of mental retardation. The phenotype is assumed to be due to overexpression of some number of the >300 genes encoded by human chromosome 21. Mouse models, in particular the chromosome 16 segmental trisomies, Ts65Dn and Ts1Cje, are indispensable for DS-related studies of gene-phenotype correlations. Here we compare the updated gene content of the finished sequence of human chromosome 21 (364 genes and putative genes) with the gene content of the homologous mouse genomic regions (291 genes and putative genes) obtained from annotation of the public sector C57Bl/6 draft sequence. Annotated genes fall into one of three classes. First, there are 170 highly conserved, human/mouse orthologues. Second, there are 83 minimally conserved, possible orthologues. Included among the conserved and minimally conserved genes are 31 antisense transcripts. Third, there are species-specific genes: 111 spliced human transcripts show no orthologues in the syntenic mouse regions although 13 have homologous sequences elsewhere in the mouse genomic sequence, and 38 spliced mouse transcripts show no identifiable human orthologues. While these species-specific genes are largely based solely on spliced EST data, a majority can be verified in RNA expression experiments. In addition, preliminary data suggest that many human-specific transcripts may represent a novel class of primate-specific genes. Lastly, updated functional annotation of orthologous genes indicates genes encoding components of several cellular pathways are dispersed throughout the orthologous mouse chromosomal regions and are not completely represented in the Down syndrome segmental mouse models. Together, these data point out the potential for existing mouse models to produce extraneous phenotypes and to fail to produce DS-relevant phenotypes.

Alternative Splicing↗

Frequency, type, distribution and annotation of simple sequence repeats in Rosaceae ESTs.

Genomic resources for peach, a model species for Rosaceae, are being developed to accelerate gene discovery in other Rosaceae species by comparative mapping. Simple sequence repeats (SSRs) are an important tool for comparative mapping because of their high polymorphism and transportability. To accelerate the development of SSR markers, we analyzed publicly available Rosaceae expressed sequence tags (ESTs) for SSRs. A total of 17,284 ESTs from almond, peach and rose were assembled into putatively non-redundant EST sets. For comparison, 179,099 ESTs from Arabidopsis were also used in the analysis. About 4% of the assembled ESTs contained SSRs in Rosaceae, which was higher than the 2.4% found in Arabidopsis. About half of the SSRs were found in the putative UTR, and the estimated average distance between SSRs in the UTR was 5.5 kb in rose, 5.1 kb in almond, 7 kb in peach and 13 kb in Arabidopsis. In the putative coding region, the estimated average distance was two to four times longer than in the UTR. Rosaceae ESTs containing SSRs were functionally annotated using the GenBank nr database and further classified using the gene ontology terms associated with the matching sequences in the SwissProt database. The detailed data including the sequences and annotation results are available from http://www.genome.clemson.edu/gdr/rosaceaessr/.

Arabidopsis↗

Gene expression patterns in the developing murine placenta.

OBJECTIVE: Successful placental development is crucial for optimal growth, maturation, and survival of the embryo/fetus. To examine genetic aspects of placental development, we investigated gene expression patterns in the murine placenta at embryonic day 10.5 (E10.5), E12.5, E15.5, and E17.5. METHODS: By use of the Affymetrix MU74A array (Affymetrix, Santa Clara, CA), we measured expression levels for 12,473 probe sets. Using pairwise analysis we selected 622 probe sets, corresponding to 599 genes, that were up- or down-regulated by more than fourfold between time points E10.5 and E12.5, E12.5 and E15.5, E15.5 and E17.5. We analyzed and functionally annotated those genes regulated during development. RESULTS: In comparing E10.5 to E12.5 we found that angiogenesis and fatty acid metabolism and transport related genes were up-regulated at E10.5, while genes involved in hormonal control and ribosomal proteins were up-regulated at E12.5. When comparing E12.5 to E15.5 we noted that genes involved in the cell cycle and RNA metabolism were strongly up-regulated at E12.5, while genes involved in cellular transport were up-regulated at E15.5. Finally, when comparing E15.5 to E17.5, we found genes related to cell cycle control, genes expressed in the nucleus and involved in RNA metabolism were up-regulated at E17.5. CONCLUSION: Microarray analysis has allowed us to describe gene expression patterns and profiles in the developing mouse placenta. Further analysis has demonstrated that several functional classes are up- and down-regulated at specific time points in placental development. These changes may have significant implications for placental development in the human.

Animals↗

Genetic-epigenetic interactions (meQTLs) in orofacial clefts etiology.

OBJECTIVES: Nonsyndromic orofacial clefts (OFCs) involve complex genetic and environmental factors, with over 60 risk loci accounting for only a minority of estimated heritability and residing in non-coding regions with unclear functional relevance. We hypothesize that some genetic variants alter orofacial cleft risk by modifying DNA methylation (DNAm) at regulatory sequences essential for craniofacial development, acting as methylation quantitative trait loci (meQTLs). METHODS: We analyzed 10 well-established OFC-associated SNPs against genome-wide DNAm profiles in 409 cases and 456 controls, identifying 23 potential meQTLs. We validated findings using 358 cleft-discordant sibling pairs analyzed with quantitative MethyLight assays. Cross-referencing with the mQTL Database assessed temporal patterns across human development. Functional annotation used GeneHancer and craniofacial enhancer databases. RESULTS: Nine meQTLs were successfully replicated, including the highly significant rs987525 (8q24) - cg16561172 (MYC) association (P = 9.610E-6). This association mapped to a mesendoderm-active enhancer upstream of MYC, providing mechanistic explanation for the longstanding 8q24 cleft locus. Additional validated associations involved MAFB-PLCG1, NOG-PPM1E, FOXE1-FRZB, and SPRY2-LGR4 interactions. Independent differential methylation analysis revealed significant differences between discordant siblings at three CpG sites. Cross-referencing confirmed concordance with population-level methylation effects, with childhood representing the critical developmental window for most associations. CONCLUSIONS: This systematic meQTL characterization in OFCs demonstrates that genetic variants influence disease risk through epigenetic mechanisms. The 8q24-MYC regulatory pathway evidence provides crucial mechanistic insight into a major OFC risk locus. These findings bridge genetic associations with functional consequences, address missing heritability challenges, and suggest potential biomarkers and therapeutic targets for OFC prevention and treatment.

Journal Article↗