PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Checks and balancers: balancer chromosomes to facilitate genome annotation.

Phenotype-driven mutagenesis screens are used to discover gene function in model organisms. Mutations that are induced by chemical mutagens can occur anywhere in the genome. However, the use of a balancer chromosome (where a phenotypically marked segment of a chromosome is inverted) in a mutagenesis screen enables mutations to be mapped in a defined region of the genome and maintained stably in a heterozygous state. Mouse balancer chromosomes can be engineered using Cre-loxP technology in selected regions of the genome. Balancer mutagenesis screens will provide a systematic functional analysis of the genes on mouse chromosomes, and consequently, will facilitate a functional annotation of the mammalian genome sequence.

Animals↗

AVID: an integrative framework for discovering functional relationships among proteins.

BACKGROUND: Determining the functions of uncharacterized proteins is one of the most pressing problems in the post-genomic era. Large scale protein-protein interaction assays, global mRNA expression analyses and systematic protein localization studies provide experimental information that can be used for this purpose. The data from such experiments contain many false positives and false negatives, but can be processed using computational methods to provide reliable information about protein-protein relationships and protein function. An outstanding and important goal is to predict detailed functional annotation for all uncharacterized proteins that is reliable enough to effectively guide experiments. RESULTS: We present AVID, a computational method that uses a multi-stage learning framework to integrate experimental results with sequence information, generating networks reflecting functional similarities among proteins. We illustrate use of the networks by making predictions of detailed Gene Ontology (GO) annotations in three categories: molecular function, biological process, and cellular component. Applied to the yeast Saccharomyces cerevisiae, AVID provides 37,451 pair-wise functional linkages between 4,191 proteins. These relationships are approximately 65-78% accurate, as assessed by cross-validation testing. Assignments of highly detailed functional descriptors to proteins, based on the networks, are estimated to be approximately 67% accurate for GO categories describing molecular function and cellular component and approximately 52% accurate for terms describing biological process. The predictions cover 1,490 proteins with no previous annotation in GO and also assign more detailed functions to many proteins annotated only with less descriptive terms. Predictions made by AVID are largely distinct from those made by other methods. Out of 37,451 predicted pair-wise relationships, the greatest number shared in common with another method is 3,413. CONCLUSION: AVID provides three networks reflecting functional associations among proteins. We use these networks to generate new, highly detailed functional predictions for roughly half of the yeast proteome that are reliable enough to drive targeted experimental investigations. The predictions suggest many specific, testable hypotheses. All of the data are available as downloadable files as well as through an interactive website at http://web.mit.edu/biology/keating/AVID. Thus, AVID will be a valuable resource for experimental biologists.

Algorithms↗

The fragment transformation method to detect the protein structural motifs.

To identify functional structural motifs from protein structures of unknown function becomes increasingly important in recent years due to the progress of the structural genomics initiatives. Although certain structural patterns such as the Asp-His-Ser catalytic triad are easy to detect because of their conserved residues and stringently constrained geometry, it is usually more challenging to detect a general structural motifs like, for example, the betabetaalpha-metal binding motif, which has a much more variable conformation and sequence. At present, the identification of these motifs usually relies on manual procedures based on different structure and sequence analysis tools. In this study, we develop a structural alignment algorithm combining both structural and sequence information to identify the local structure motifs. We applied our method to the following examples: the betabetaalpha-metal binding motif and the treble clef motif. The betabetaalpha-metal binding motif plays an important role in nonspecific DNA interactions and cleavage in host defense and apoptosis. The treble clef motif is a zinc-binding motif adaptable to diverse functions such as the binding of nucleic acid and hydrolysis of phosphodiester bonds. Our results are encouraging, indicating that we can effectively identify these structural motifs in an automatic fashion. Our method may provide a useful means for automatic functional annotation through detecting structural motifs associated with particular functions.

Amino Acid Motifs↗

Exploring the ecological drivers of bacteriophage diversity and functional viral potential in the skin of the axolotl Ambystoma altamirani.

Bacteriophages play important roles in shaping microbial community dynamics across diverse environments. In the amphibian skin, most microbiome studies have focused on bacteria and their interactions with the fungus Batrachochytrium dendrobatidis (Bd), leaving other microbial components, including viruses, largely unexplored. Here, we present the first characterization of the viral community in the amphibian skin microbiome, focusing on ecological drivers of bacteriophage diversity and functional potential in the axolotl Ambystoma altamirani. Using public shotgun metagenomes, we found that the viral fraction was dominated by bacteriophages of the class Caudoviricetes. Bacteriophage diversity was significantly associated with local physicochemical parameters at the time of sampling, and showed a strong positive correlation with bacterial diversity, whereas no significant associations were detected with the presence of Bd. In addition, seasonality influenced the composition and properties of bacteria-bacteriophage co-abundance networks. Functional annotation of assembled bacteriophage sequences revealed a diverse functional potential, including putative auxiliary metabolic genes, superinfection exclusion, toxin-antitoxin, and virulence factors. Overall, these findings highlight the ecological relevance of bacteriophages in amphibian skin microbiomes and underscore the need for further studies on their role in the amphibian host's health.

Animals↗

Large-scale analysis of the barley transcriptome based on expressed sequence tags.

To provide resources for barley genomics, 110,981 expressed sequence tags (ESTs) were generated from 22 cDNA libraries representing tissues at various developmental stages. This EST collection corresponds to approximately one-third of the 380,000 publicly available barley ESTs. Clustering and assembly resulted in 14,151 tentative consensi (TCs) and 11 073 singletons, altogether representing 25 224 putatively unique sequences. Of these, 17.5% showed no significant similarity to other barley ESTs present in dbEST. More than 41% of all barley genes are supposed to belong to multigene families and approximately 4% of the barley genes undergo alternative splicing. Based on the functional annotation of the set of unique sequences, the functional category 'Energy' was further analysed to reveal tissue- and stage-specific differences in gene expression. Hierarchical clustering of 362 differentially expressed TCs resulted in the identification of seven major clusters. The clusters reflect biochemical pathways predominantly activated in specific tissues and at various developmental stages. During seed germination glycolysis could be identified as the most predominant biochemical pathway. Germination-specific glycolysis is characterized by the coordinated expression of phosphoenolpyruvate carboxylase and phosphoenolpyruvate carboxykinase, whose antagonistic actions possibly regulate the flux of amino acids into protein biosynthesis and gluconeogenesis respectively. The expression of defence-related and antioxidant genes during germination might be controlled by the ethylene-signalling pathway as concluded from the coordinated expression of those genes and the transcription factors (TF) EIN3 and EREBPG. Moreover, because of their predominant expression in germinating seeds, TF of the AP2 and MYB type are presumably major regulators of germination.

Expressed Sequence Tags↗

microRNA target predictions across seven Drosophila species and comparison to mammalian targets.

microRNAs are small noncoding genes that regulate the protein production of genes by binding to partially complementary sites in the mRNAs of targeted genes. Here, using our algorithm PicTar, we exploit cross-species comparisons to predict, on average, 54 targeted genes per microRNA above noise in Drosophila melanogaster. Analysis of the functional annotation of target genes furthermore suggests specific biological functions for many microRNAs. We also predict combinatorial targets for clustered microRNAs and find that some clustered microRNAs are likely to coordinately regulate target genes. Furthermore, we compare microRNA regulation between insects and vertebrates. We find that the widespread extent of gene regulation by microRNAs is comparable between flies and mammals but that certain microRNAs may function in clade-specific modes of gene regulation. One of these microRNAs (miR-210) is predicted to contribute to the regulation of fly oogenesis. We also list specific regulatory relationships that appear to be conserved between flies and mammals. Our findings provide the most extensive microRNA target predictions in Drosophila to date, suggest specific functional roles for most microRNAs, indicate the existence of coordinate gene regulation executed by clustered microRNAs, and shed light on the evolution of microRNA function across large evolutionary distances. All predictions are freely accessible at our searchable Web site http://pictar.bio.nyu.edu.

Journal Article↗

Who tangos with GOA?-Use of Gene Ontology Annotation (GOA) for biological interpretation of '-omics' data and for validation of automatic annotation tools.

The number of large-scale experimental datasets generated from high-throughput technologies has grown rapidly. Biological knowledge resources such as the Gene Ontology Annotation (GOA) database, which provides high-quality functional annotation to proteins within the UniProt Knowledgebase, can play an important role in the analysis of such data. The integration of GOA with analytical tools has proved to aid the clustering, annotation and biological interpretation of such large expression datasets. GOA is also useful in the development and validation of automated annotation tools, in particular text-mining systems. The increasing interest in GOA highlights the great potential of this freely available resource to assist both the biological research and bioinformatics communities.

Animals↗

A procedure for assessing GO annotation consistency.

MOTIVATION: The Gene Ontology (GO) is widely used to annotate molecular attributes of genes and gene products. Multiple groups undertaking functional annotations of genomes contribute their annotation sets to the GO database resource and these data are subsequently used in comparative functional analysis research. Although GO curators adhere to the same protocols and standards while assigning GO annotations, the specific procedure followed by each annotation group can vary. Since differences in application of annotation standards would dilute the effectiveness of comparative analysis, methods for assessing annotation consistency are essential. The development of methodologies that are broadly applicable for the assessment of GO annotation consistency is an important issue for the comparative genomics community. RESULTS: We have developed a methodology for assessing the consistency of GO annotations provided by different annotation groups. The method is completely general and can be applied to compare any two sets of GO annotations. This is the first attempt to assess cross-species GO annotation consistency. Our method compares annotation sets utilizing the hierarchical structure of the GO to compare GO annotations between orthologous gene pairs. The method produces a report on the annotation consistency and inconsistency for each orthologous pair. We present results obtained by comparing GO annotations for mouse and human gene sets. AVAILABILITY: The complete current MGI_GOA GO annotation consistency report is available online at http://www.spatial.maine.edu/~mdolan/

Animals↗

Functional mapping and annotation of genetic associations with FUMA.

A main challenge in genome-wide association studies (GWAS) is to pinpoint possible causal variants. Results from GWAS typically do not directly translate into causal variants because the majority of hits are in non-coding or intergenic regions, and the presence of linkage disequilibrium leads to effects being statistically spread out across multiple variants. Post-GWAS annotation facilitates the selection of most likely causal variant(s). Multiple resources are available for post-GWAS annotation, yet these can be time consuming and do not provide integrated visual aids for data interpretation. We, therefore, develop FUMA: an integrative web-based platform using information from multiple biological resources to facilitate functional annotation of GWAS results, gene prioritization and interactive visualization. FUMA accommodates positional, expression quantitative trait loci (eQTL) and chromatin interaction mappings, and provides gene-based, pathway and tissue enrichment results. FUMA results directly aid in generating hypotheses that are testable in functional experiments aimed at proving causal relations.

Chromatin↗

Systematic discovery of pathogen effector functions across human pathogens and pathways.

Pathogens deploy effector proteins to exploit host cell biology, and most effector open reading frames (ORFs) are rapidly evolving and lack functional annotation. We developed the effector ORFeome (eORFeome), a scalable functional genomics platform encompassing 3,835 effector ORFs from diverse viruses, bacteria, and parasites. High-throughput barcoded screens across nuclear factor κB (NF-κB), apoptosis, p53, cGAS-STING, and major histocompatibility complex class I (MHC class I) pathways revealed novel pathway-modulating functions for hundreds of uncharacterized eORFs, unexpected activities of known effectors, and distinct pathway-specific functions encoded by single ORFs. Illustrating the power of this approach, we identified HHV6A U14 as a p53 antagonist, HHV7 U21 as a dual-function STING antagonist and MHC-I antigen display inhibitor, and adenoviral 13.6K/i-leader protein as a de novo-evolved TAP inhibitor that suppresses MHC-I display. These results establish a general framework for systematic effector annotation, uncover new mechanisms of host-pathogen interaction across kingdoms, and highlight pathogen effectors as a versatile toolkit for rewiring and probing human cellular pathways.

Humans↗

High precision multi-genome scale reannotation of enzyme function by EFICAz.

BACKGROUND: The functional annotation of most genes in newly sequenced genomes is inferred from similarity to previously characterized sequences, an annotation strategy that often leads to erroneous assignments. We have performed a reannotation of 245 genomes using an updated version of EFICAz, a highly precise method for enzyme function prediction. RESULTS: Based on our three-field EC number predictions, we have obtained lower-bound estimates for the average enzyme content in Archaea (29%), Bacteria (30%) and Eukarya (18%). Most annotations added in KEGG from 2005 to 2006 agree with EFICAz predictions made in 2005. The coverage of EFICAz predictions is significantly higher than that of KEGG, especially for eukaryotes. Thousands of our novel predictions correspond to hypothetical proteins. We have identified a subset of 64 hypothetical proteins with low sequence identity to EFICAz training enzymes, whose biochemical functions have been recently characterized and find that in 96% (84%) of the cases we correctly identified their three-field (four-field) EC numbers. For two of the 64 hypothetical proteins: PA1167 from Pseudomonas aeruginosa, an alginate lyase (EC 4.2.2.3) and Rv1700 of Mycobacterium tuberculosis H37Rv, an ADP-ribose diphosphatase (EC 3.6.1.13), we have detected annotation lag of more than two years in databases. Two examples are presented where EFICAz predictions act as hypothesis generators for understanding the functional roles of hypothetical proteins: FLJ11151, a human protein overexpressed in cancer that EFICAz identifies as an endopolyphosphatase (EC 3.6.1.10), and MW0119, a protein of Staphylococcus aureus strain MW2 that we propose as candidate virulence factor based on its EFICAz predicted activity, sphingomyelin phosphodiesterase (EC 3.1.4.12). CONCLUSION: Our results suggest that we have generated enzyme function annotations of high precision and recall. These predictions can be mined and correlated with other information sources to generate biologically significant hypotheses and can be useful for comparative genome analysis and automated metabolic pathway reconstruction.

Animals↗

A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).

Genomic sequencing is no longer a novelty, but gene function annotation remains a key challenge in modern biology. A variety of functional genomics experimental techniques are available, from classic methods such as affinity precipitation to advanced high-throughput techniques such as gene expression microarrays. In the future, more disparate methods will be developed, further increasing the need for integrated computational analysis of data generated by these studies. We address this problem with MAGIC (Multisource Association of Genes by Integration of Clusters), a general framework that uses formal Bayesian reasoning to integrate heterogeneous types of high-throughput biological data (such as large-scale two-hybrid screens and multiple microarray analyses) for accurate gene function prediction. The system formally incorporates expert knowledge about relative accuracies of data sources to combine them within a normative framework. MAGIC provides a belief level with its output that allows the user to vary the stringency of predictions. We applied MAGIC to Saccharomyces cerevisiae genetic and physical interactions, microarray, and transcription factor binding sites data and assessed the biological relevance of gene groupings using Gene Ontology annotations produced by the Saccharomyces Genome Database. We found that by creating functional groupings based on heterogeneous data types, MAGIC improved accuracy of the groupings compared with microarray analysis alone. We describe several of the biological gene groupings identified.

Algorithms↗

Proteomic analysis using an unfinished bacterial genome: the effects of subminimum inhibitory concentrations of antibiotics on Mannheimia haemolytica virulence factor expression.

Here we identify, using nonelectrophoretic proteomics, effects of subminimum inhibitory concentrations (subMIC) of two antibiotic preparations, chlortetracycline (CTC), and chlortetracycline-sulfamethazine (CTC + SMZ), on protein expression in the bovine respiratory pathogen Mannheimia haemolytica. The M. haemolytica genome is currently in draft form, and annotation is incomplete. Relying on the principle of gene sequence conservation across species, we used annotated genomes from closely related species to identify, confirm, and functionally annotate 495 M. haemolytica proteins. To conduct quantitative comparative proteomics, we developed a protein quantitation method based on the cross correlation function of the SEQUEST algorithm. When M. haemolytica was cultivated in the presence of 1/4 MIC of CTC and CTC + SMZ, expression of proteins involved in energy production, nucleotide metabolism, translation, and the bacterial stress response (chaperones) were affected. The most notable subMIC effect was a significant decrease in the expression of leukotoxin A, which is an important M. haemolytica virulence factor. Reduction in leukotoxin expression could be one of the molecular mechanisms responsible for the efficacy of these antibiotics against bovine respiratory disease.

Algorithms↗

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

Distinguishing enzyme structures from non-enzymes without alignments.

The ability to predict protein function from structure is becoming increasingly important as the number of structures resolved is growing more rapidly than our capacity to study function. Current methods for predicting protein function are mostly reliant on identifying a similar protein of known function. For proteins that are highly dissimilar or are only similar to proteins also lacking functional annotations, these methods fail. Here, we show that protein function can be predicted as enzymatic or not without resorting to alignments. We describe 1178 high-resolution proteins in a structurally non-redundant subset of the Protein Data Bank using simple features such as secondary-structure content, amino acid propensities, surface properties and ligands. The subset is split into two functional groupings, enzymes and non-enzymes. We use the support vector machine-learning algorithm to develop models that are capable of assigning the protein class. Validation of the method shows that the function can be predicted to an accuracy of 77% using 52 features to describe each protein. An adaptive search of possible subsets of features produces a simplified model based on 36 features that predicts at an accuracy of 80%. We compare the method to sequence-based methods that also avoid calculating alignments and predict a recently released set of unrelated proteins. The most useful features for distinguishing enzymes from non-enzymes are secondary-structure content, amino acid frequencies, number of disulphide bonds and size of the largest cleft. This method is applicable to any structure as it does not require the identification of sequence or structural similarity to a protein of known function.

Algorithms↗

Identifying active transcription factors and kinases from expression data using pathway queries.

MOTIVATION: Although progress has been made identifying regulatory relationships from expression data in general, only few methods have focused on detecting biological mechanisms like active pathways using a single measurement. This is of particular importance when only few measurements are available, e.g. if special cell types or conditions are under investigation. Here we present a method to test user specified hypotheses (pathway queries) on expression data where prior knowledge is given in the form of networks and functional annotations. Based on this method, we develop a scoring function to identify active transcription factors or kinases, thus making a first step toward explaining the measured expression data. RESULTS: We apply the algorithm to the Rosetta Yeast Compendium dataset, finding that in many cases the results are in concordance with biological knowledge. We were able to confirm that transcription factors and to a lesser degree, kinases identified by our method play an important role in the biological processes affected by the respective knockouts. Furthermore, we show that correlation of inferred activities can provide evidence for a physical interaction or cooperation of transcription factors where correlation of plain expression data fails to do so.

Algorithms↗

Improving the Annotations of JCVI-Syn3a Proteins.

The JCVI-Syn3 organism is a minimal organism derived from Mycoplasma mycoides capri, which is capable of self-replication. While the ancestor has 863 genes, the synthetic progeny has only 473, with 434 of these coding for proteins. Despite initial efforts to understand all functions of the organism, a significant number of these protein-coding genes still have unknown functions, and subsequent studies have been only partially successful in elucidating their roles. In this study, we employ our innovative method PROST to identify homologs and better understand these previously unidentified genes. PROST employs protein language embeddings and enables the identification of remote homologs with as low as 16% sequence identity. PROST successfully finds functionally annotated homologs for 93% of the minimal genome with a high level of accuracy, both confirming previously identified functions, as well as proposing new functions for others. The results of our study can be accessed at https://bit.ly/prost-syn3a .

Molecular Sequence Annotation↗

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676 bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria↗