PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein↗

How well is enzyme function conserved as a function of pairwise sequence identity?

Enzyme function conservation has been used to derive the threshold of sequence identity necessary to transfer function from a protein of known function to an unknown protein. Using pairwise sequence comparison, several studies suggested that when the sequence identity is above 40%, enzyme function is well conserved. In contrast, Rost argued that because of database bias, the results from such simple pairwise comparisons might be misleading. Thus, by grouping enzyme sequences into families based on sequence similarity and selecting representative sequences for comparison, he showed that enzyme function starts to diverge quickly when the sequence identity is below 70%. Here, we employ a strategy similar to Rost's to reduce the database bias; however, we classify enzyme families based not only on sequence similarity, but also on functional similarity, i.e. sequences in each family must have the same four digits or the same first three digits of the enzyme commission (EC) number. Furthermore, instead of selecting representative sequences for comparison, we calculate the function conservation of each enzyme family and then average the degree of enzyme function conservation across all enzyme families. Our analysis suggests that for functional transferability, 40% sequence identity can still be used as a confident threshold to transfer the first three digits of an EC number; however, to transfer all four digits of an EC number, above 60% sequence identity is needed to have at least 90% accuracy. Moreover, when PSI-BLAST is used, the magnitude of the E-value is found to be weakly correlated with the extent of enzyme function conservation in the third iteration of PSI-BLAST. As a result, functional annotation based on the E-values from PSI-BLAST should be used with caution. We also show that by employing an enzyme family-specific sequence identity threshold above which 100% functional conservation is required, functional inference of unknown sequences can be accurately accomplished. However, this comes at a cost: those true positive sequences below this threshold cannot be uniquely identified.

Animals↗

Preparation of a set of expression-ready clones of mammalian long cDNAs encoding large proteins by the ORF trap cloning method.

Although we have so far identified and sequenced >2000 human long cDNAs, known as KIAA cDNAs, half of them have yet to be functionally annotated. Expression-ready cDNA clones derived from these genes, where the open reading frame (ORF) of the gene of interest is placed under the control of an appropriate promoter, are critical for functional characterization of these gene products. In this study, we attempted to systematically convert original cDNA clones to expression-ready forms for native and fusion proteins. For this purpose, we developed a new method for ORF cloning based on a homologous recombination in Escherichia coli to avoid laborious manipulations and artificial introduction of mutations in ORF. Using 1589 putative full-length ORFs (from 1002 KIAA genes, 119 human known genes and 468 mouse genes) with an average size of 2.8 kb, we successfully prepared expression plasmids for 1463 native proteins and for 1343 fusion proteins by this method. The resultant expression-ready clones were examined using an in vitro transcription/translation system followed by SDS-polyacrylamide gel electrophoresis and by transient expression of GFP-fusion proteins in human embryonic kidney (HEK) 293 cells. This set of expression-ready clones of long cDNAs encoding large proteins would open a new route to experimentally analyze their functions on a proteomic scale, since unavailability of expression-ready clones for mammalian large proteins has been a major obstacle to the functional analysis of these cDNAs.

Animals↗

A protein-protein interaction map of the Caenorhabditis elegans 26S proteasome.

The ubiquitin-proteasome proteolytic pathway is pivotal in most biological processes. Despite a great level of information available for the eukaryotic 26S proteasome-the protease responsible for the degradation of ubiquitylated proteins-several structural and functional questions remain unanswered. To gain more insight into the assembly and function of the metazoan 26S proteasome, a two-hybrid-based protein interaction map was generated using 30 Caenorhabditis elegans proteasome subunits. The results recapitulate interactions reported for other organisms and reveal new potential interactions both within the 19S regulatory complex and between the 19S and 20S subcomplexes. Moreover, novel potential proteasome interactors were identified, including an E3 ubiquitin ligase, transcription factors, chaperone proteins and other proteins not yet functionally annotated. By providing a wealth of novel biological hypotheses, this interaction map constitutes a framework for further analysis of the ubiquitin-proteasome pathway in a multicellular organism amenable to both classical genetics and functional genomics.

Animals↗

Iron-related transcriptomic variations in CaCo-2 cells, an in vitro model of intestinal absorptive cells.

Regulation of iron absorption by duodenal enterocytes is essential for the maintenance of homeostasis by preventing iron deficiency or overload. Despite the identification of a number of genes implicated in iron absorption and its regulation, it is likely that further factors remain to be identified. For that purpose, we used a global transcriptomic approach, using the CaCo-2 cell line as an in vitro model of intestinal absorptive cells. Pangenomic screening for variations in gene expression correlating with intracellular iron content allowed us to identify 171 genes. One hundred nine of these genes are clustered into five types of expression profile. This is the first time that most of these genes have been associated with iron metabolism. Functional annotation of these five clusters indicates potential links between the immune response, proteolysis processes, and iron depletion. In contrast, iron overload is associated with cellular metabolism, especially that of lipids and glutathione involving redox function and electron transfer.

Caco-2 Cells↗

Genes involved in complex adaptive processes tend to have highly conserved upstream regions in mammalian genomes.

BACKGROUND: Recent advances in genome sequencing suggest a remarkable conservation in gene content of mammalian organisms. The similarity in gene repertoire present in different organisms has increased interest in studying regulatory mechanisms of gene expression aimed at elucidating the differences in phenotypes. In particular, a proximal promoter region contains a large number of regulatory elements that control the expression of its downstream gene. Although many studies have focused on identification of these elements, a broader picture on the complexity of transcriptional regulation of different biological processes has not been addressed in mammals. The regulatory complexity may strongly correlate with gene function, as different evolutionary forces must act on the regulatory systems under different biological conditions. We investigate this hypothesis by comparing the conservation of promoters upstream of genes classified in different functional categories. RESULTS: By conducting a rank correlation analysis between functional annotation and upstream sequence alignment scores obtained by human-mouse and human-dog comparison, we found a significantly greater conservation of the upstream sequence of genes involved in development, cell communication, neural functions and signaling processes than those involved in more basic processes shared with unicellular organisms such as metabolism and ribosomal function. This observation persists after controlling for G+C content. Considering conservation as a functional signature, we hypothesize a higher density of cis-regulatory elements upstream of genes participating in complex and adaptive processes. CONCLUSION: We identified a class of functions that are associated with either high or low promoter conservation in mammals. We detected a significant tendency that points to complex and adaptive processes were associated with higher promoter conservation, despite the fact that they have emerged relatively recently during evolution. We described and contrasted several hypotheses that provide a deeper insight into how transcriptional complexity might have been emerged during evolution.

Animals↗

Automatic prediction of protein function.

Most methods annotating protein function utilise sequence homology to proteins of experimentally known function. Such a homology-based annotation transfer is problematic and limited in scope. Therefore, computational biologists have begun to develop ab initio methods that predict aspects of function, including subcellular localization, post-translational modifications, functional type and protein-protein interactions. For the first two cases, the most accurate approaches rely on identifying short signalling motifs, while the most general methods utilise tools of artificial intelligence. An outstanding new method predicts classes of cellular function directly from sequence. Similarly, promising methods have been developed predicting protein-protein interaction partners at acceptable levels of accuracy for some pairs in entire proteomes. No matter how difficult the task, successes over the last few years have clearly paved the way for ab initio prediction of protein function.

Computational Biology↗

The European dimension for the mouse genome mutagenesis program.

The European Mouse Mutagenesis Consortium is the European initiative contributing to the international effort on functional annotation of the mouse genome. Its objectives are to establish and integrate mutagenesis platforms, gene expression resources, phenotyping units, storage and distribution centers and bioinformatics resources. The combined efforts will accelerate our understanding of gene function and of human health and disease.

Animals↗

RAREsim2: flexible simulation of rare variant genetic data using real haplotypes.

MOTIVATION: Realistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., case-control status, technological or batch effects) and variant-level differences to represent a variety of causal models. RESULTS: We demonstrate RAREsim2's utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios, including various genetic ancestries, gene sizes, strengths of association, and proportions of risk variants. Type I Error was maintained and the test with the highest power matched previously known patterns. Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. AVAILABILITY AND IMPLEMENTATION: The RAREsim2 Python package is publicly available on Github (https://github.com/Hendricks-Research-Team/RAREsim2), PyPI (https://pypi.org/project/raresim/), and Zenodo (https://doi.org/10.5281/zenodo.19442523). Code for the example demonstration can be found at https://github.com/JessMurphy/RAREsim2-demo.

Software↗

Complete reannotation of the Arabidopsis genome: methods, tools, protocols and the final release.

BACKGROUND: Since the initial publication of its complete genome sequence, Arabidopsis thaliana has become more important than ever as a model for plant research. However, the initial genome annotation was submitted by multiple centers using inconsistent methods, making the data difficult to use for many applications. RESULTS: Over the course of three years, TIGR has completed its effort to standardize the structural and functional annotation of the Arabidopsis genome. Using both manual and automated methods, Arabidopsis gene structures were refined and gene products were renamed and assigned to Gene Ontology categories. We present an overview of the methods employed, tools developed, and protocols followed, summarizing the contents of each data release with special emphasis on our final annotation release (version 5). CONCLUSION: Over the entire period, several thousand new genes and pseudogenes were added to the annotation. Approximately one third of the originally annotated gene models were significantly refined yielding improved gene structure annotations, and every protein-coding gene was manually inspected and classified using Gene Ontology terms.

Alternative Splicing↗

GPAC: benchmarking the sensitivity of genome informatics analysis to genome annotation completeness.

In view of the recent explosion in genome sequence data, and the 200 or more complete genome sequences currently available, the importance of genome-scale bioinformatics analysis is increasing rapidly. However, computational genome informatics analyses often lack a statistical assessment of their sensitivity to the completeness of the functional annotation. Therefore, a pre-analysis method to automatically validate the sensitivity of computational genome analyses with regard to genome annotation completeness is useful for this purpose. In this report we developed the Gene Prediction Accuracy Classification (GPAC) test, which provides statistical evidence of sensitivity by repeating the same analysis for five different gene groups (classified according to annotation accuracy level), and for randomly sampled gene groups, with the same number of genes as each of the five classified groups. Variability in these results is then assessed, and if the results vary significantly with different data subsets, the analysis is considered "sensitive" to annotation completeness, and careful selection of data is advised prior to the actual in silico analysis. The GPAC test has been applied to the analyses of Sakai et al., 2001, and Ohno et al., 2001, and it revealed that the analysis of Ohno et al. was more sensitive to annotation completeness. It showed that GPAC could be employed to ascertain the sensitivity of an analysis. The GPAC bendhmarking software is freely available in the latest G-language Genome Analysis Environment package, at http://www.g-language.org/.

Benchmarking↗

Gene expression in the brain and kidney of rainbow trout in response to handling stress.

BACKGROUND: Microarray technologies are rapidly becoming available for new species including teleost fishes. We constructed a rainbow trout cDNA microarray targeted at the identification of genes which are differentially expressed in response to environmental stressors. This platform included clones from normalized and subtracted libraries and genes selected through functional annotation. Present study focused on time-course comparisons of stress responses in the brain and kidney and the identification of a set of genes which are diagnostic for stress response. RESULTS: Fish were stressed with handling and samples were collected 1, 3 and 5 days after the first exposure. Gene expression profiles were analysed in terms of Gene Ontology categories. Stress affected different functional groups of genes in the tissues studied. Mitochondria, extracellular matrix and endopeptidases (especially collagenases) were the major targets in kidney. Stress response in brain was characterized with dramatic temporal alterations. Metal ion binding proteins, glycolytic enzymes and motor proteins were induced transiently, whereas expression of genes involved in stress and immune response, cell proliferation and growth, signal transduction and apoptosis, protein biosynthesis and folding changed in a reciprocal fashion. Despite dramatic difference between tissues and time-points, we were able to identify a group of 48 genes that showed strong correlation of expression profiles (Pearson r > /0.65/) in 35 microarray experiments being regulated by stress. We evaluated performance of the clone sets used for preparation of microarray. Overall, the number of differentially expressed genes was markedly higher in EST than in genes selected through Gene Ontology annotations, however 63% of stress-responsive genes were from this group. CONCLUSIONS: 1. Stress responses in fish brain and kidney are different in function and time-course. 2. Identification of stress-regulated genes provides the possibility for measuring stress responses in various conditions and further search for the functionally related genes.

Animals↗

Genome-Wide Characterization of β-Glucosidase (TaBGLU) Genes in Bread Wheat and Their Expression Under Drought, Cold, and Combined Stress.

Glycoside hydrolase 1 (GH1) β-glucosidases were known to activate hormone conjugates and defense metabolites, yet their genomic organization and stress-response dynamics in wheat remained incompletely defined. We therefore performed an integrated characterization of TaBGLUs spanning phylogeny, gene structure and conserved motifs, subcellular localization, promoter cis-elements, Gene Ontology enrichment, protein-protein interaction networks, and targeted expression profiling. Wheat TaBGLUs partitioned into well-supported clades that shared canonical GH1 catalytic residues and a largely conserved motif scaffold. Subcellular localization predictions indicated predominant nuclear and chloroplast targeting, with a smaller cohort directed to secretory or endomembrane compartments. Promoters were enriched for light-responsive, hormone-related (ABA, JA/SA, auxin, GA) and stress-associated (MYB/WRKY, heat, low temperature) cis-elements, and functional annotations were consistent with roles in carbohydrate and cell-wall metabolism, hormone homeostasis, and defense. Network analysis revealed a densely connected TaBGLU submodule embedded within broader carbohydrate and defense interaction networks, suggesting coordinated or cooperative functions. Expression profiling under cold, drought, and combined drought and cold demonstrated broad stress inducibility, with early activation detected by 6 h, cold-responsive maxima typically at 12 h, drought-responsive peaks predominating at 24 h, and combined stress eliciting both earlier and more sustained expression maxima between 12-24 h. Representative strongly responsive genes included TaBGLU20, TaBGLU44, TaBGLU6, and TaBGLU23, which showed pronounced late induction under combined stress, TaBGLU30, which exhibited an earlier combined-stress peak, and TaBGLU12, which displayed a marked late drought-specific response. Taken together, this integrated genomic, regulatory, and expression atlas refined the wheat BGLU repertoire relative to previous gene model inventories, highlighted candidate TaBGLUs with central network positions and strong stress inducibility, and provided concrete entry points for functional validation and breeding for improved stress resilience.

Triticum↗

Unexpected catalytic site variation in phosphoprotein phosphatase homologues of cofactor-dependent phosphoglycerate mutase.

The cofactor-dependent phosphoglycerate mutase (dPGM) superfamily contains, besides mutases, a variety of phosphatases, both broadly and narrowly substrate-specific. Distant dPGM homologues, conspicuously abundant in microbial genomes, represent a challenge for functional annotation based on sequence comparison alone. Here we carry out sequence analysis and molecular modelling of two families of bacterial dPGM homologues, one the SixA phosphoprotein phosphatases, the other containing various proteins of no known molecular function. The models show how SixA proteins have adapted to phosphoprotein substrate and suggest that the second family may also encode phosphoprotein phosphatases. Unexpected variation in catalytic and substrate-binding residues is observed in the models.

Amino Acid Sequence↗

Annotated expressed sequence tags and cDNA microarrays for studies of brain and behavior in the honey bee.

To accelerate the molecular analysis of behavior in the honey bee (Apis mellifera), we created expressed sequence tag (EST) and cDNA microarray resources for the bee brain. Over 20,000 cDNA clones were partially sequenced from a normalized (and subsequently subtracted) library generated from adult A. mellifera brains. These sequences were processed to identify 15,311 high-quality ESTs representing 8912 putative transcripts. Putative transcripts were functionally annotated (using the Gene Ontology classification system) based on matching gene sequences in Drosophila melanogaster. The brain ESTs represent a broad range of molecular functions and biological processes, with neurobiological classifications particularly well represented. Roughly half of Drosophila genes currently implicated in synaptic transmission and/or behavior are represented in the Apis EST set. Of Apis sequences with open reading frames of at least 450 bp, 24% are highly diverged with no matches to known protein sequences. Additionally, over 100 Apis transcript sequences conserved with other organisms appear to have been lost from the Drosophila genome. DNA microarrays were fabricated with over 7000 EST cDNA clones putatively representing different transcripts. Using probe derived from single bee brain mRNA, microarrays detected gene expression for 90% of Apis cDNAs two standard deviations greater than exogenous control cDNAs. [The sequence data described in this paper have been submitted to Genbank data library under accession nos. BI502708-BI517278. The sequences are also available at http://titan.biotec.uiuc.edu/bee/honeybee_project.htm.]

Animals↗

Zinc coordination environments in proteins determine zinc functions.

Estimates of the number of zinc proteins in humans are now possible and a functional annotation of the zinc proteome can begin. The catalytic and structural roles of zinc in hundreds of enzymes and thousands of so-called "zinc finger" protein domains have provided a molecular basis for the numerous biological functions of this essential element. Additional, regulatory functions of zinc/protein interactions are being recognized. They include roles of the zinc ion in signal transduction, in controlling the architecture of protein complexes, and in redox-active zinc sites, where the binding and release of zinc is under redox control. Moreover, a considerable number of proteins participate in cellular zinc homeostasis, e.g. membrane transporters, and cellular storage, sensor, and trafficking proteins. These proteins have evolved with mechanisms to handle zinc ions rather specifically and selectively. They perform their functions with a remarkably modest set: One redox state of the zinc ion and nitrogen, oxygen, and sulfur ligands from the side chains of histidine, glutamate/aspartate, and cysteine, respectively. By permutation of the ligands in this set, the functional potential of the zinc ion has been fully explored. Different coordination environments modulate the chemical characteristics of the zinc ion, control the kinetics of its binding, and allow it to be either metabolically active or inert. Insights into all these functions are building an understanding of why zinc is so critical for such a multitude of life processes.

Gene Expression Regulation↗

High-resolution functional proteomics by active-site peptide profiling.

Characterization and functional annotation of the large number of proteins predicted from genome sequencing projects poses a major scientific challenge. Whereas several proteomics techniques have been developed to quantify the abundance of proteins, these methods provide little information regarding protein function. Here, we present a gel-free platform that permits ultrasensitive, quantitative, and high-resolution analyses of protein activities in proteomes, including highly problematic samples such as undiluted plasma. We demonstrate the value of this platform for the discovery of both disease-related enzyme activities and specific inhibitors that target these proteins.

Animals↗