PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Predicting gene function from patterns of annotation.

The Gene Ontology (GO) Consortium has produced a controlled vocabulary for annotation of gene function that is used in many organism-specific gene annotation databases. This allows the prediction of gene function based on patterns of annotation. For example, if annotations for two attributes tend to occur together in a database, then a gene holding one attribute is likely to hold the other as well. We modeled the relationships among GO attributes with decision trees and Bayesian networks, using the annotations in the Saccharomyces Genome Database (SGD) and in FlyBase as training data. We tested the models using cross-validation, and we manually assessed 100 gene-attribute associations that were predicted by the models but that were not present in the SGD or FlyBase databases. Of the 100 manually assessed associations, 41 were judged to be true, and another 42 were judged to be plausible.

Animals↗

Understanding the cell in terms of structure and function: insights from structural genomics.

Structural genomics programs are only now moving into the large-scale production phase, yet have already produced around 2000 protein structures. Through a widespread if not exclusive emphasis on structural novelty, our knowledge of the protein fold universe is improving rapidly. With this information comes the challenge of structure-based function annotation for the many target proteins about which little or nothing is known. Recent years have therefore seen the emergence of impressively diverse bioinformatics approaches to predict the function of a protein structure. Attention is now turning to means of combining these predictions with information from various other sources.

Computational Biology↗

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (∼150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ∼20% for AlphaMissense underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular protein trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article↗

Dental wastewater reveals a hidden reservoir of oral bacteriophage diversity.

Bacteriophages (phages) are being explored as alternatives or complements to antibiotics because of their ability to selectively kill bacterial pathogens. However, phages that infect many oral bacteria remain undiscovered. Here, we discovered that dental wastewater harbors previously underexplored phage diversity. Viral particles concentrated from dental wastewater displayed diverse morphologies, including abundant filamentous phage-like particles. Deep long-read metagenomic sequencing of concentrated viral particles generated 7.4 billion bases of sequence data and yielded 255 medium- to high-quality viral operational taxonomic units (vOTUs), including 46 predicted complete genomes. Comparison with large phage databases revealed that 63 of these 255 vOTUs had no detectable match, indicating that extensive sequencing of dental wastewater substantially expands the number of potential bacteriophages associated with the human oral microbiome. Host prediction linked many vOTUs to oral-associated bacterial taxa, including species with few or no previously reported phages, such as Porphyromonas gingivalis, Tannerella forsythia, and Candidatus Saccharibacteria. Functional annotation identified diverse genes associated with antiphage defense systems within a subset of vOTUs, suggesting that oral phages may contribute to the movement of genes encoding bacterial immune functions within the oral microbiome. Together, these findings expand the known oral phageome and show that dental wastewater contains a largely untapped diversity of phages.IMPORTANCEThe human oral cavity contains a diverse microbial community, but the bacteriophages (phages) that infect many oral bacteria remain poorly characterized. This gap limits our understanding of how phages shape oral microbial communities. Here, we show that dental wastewater is an underexplored source of oral phage diversity. Deep long-read metagenomic sequencing revealed 255 medium- to high-quality phage operational taxonomic units, many of which are not present in existing oral phage databases. These genomes include predicted phages of periodontal disease-associated bacteria and other oral taxa with few or no known phages. Dental wastewater therefore expands the known human oral phageome and reveals candidate phages linked to bacteria associated with oral health and disease.

Bacteriophages↗

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals↗

Genome-wide screening for gene function using RNAi in mammalian cells.

Mammalian genome sequencing has identified numerous genes requiring functional annotation. The discovery that dsRNA can direct gene-specific silencing in both model organisms and mammalian cells through RNA interference (RNAi) has provided a platform for dissecting the function of independent genes. The generation of large-scale RNAi libraries targeting all predicted genes within mouse, rat and human cells, combined with the large number of cell-based assays, provides a unique opportunity to perform high-throughput genetics in these complex cell systems. Many different formats exist for the generation of genome-wide RNAi libraries for use in mammalian cells. Furthermore, the use of these libraries in either genetic screens or genetic selections allows for the identification of known and novel genes involved in complex cellular phenotypes and biological processes, some of which underpin human disease. In this review, we examine genome-wide RNAi libraries used in model organisms and mammalian cells and provide examples of how these information rich reagents can be used for determining gene function, discovering novel therapeutic targets and dissecting signalling pathways, cellular processes and complex phenotypes.

Animals↗

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements↗

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important, promiscuous divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (~150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ~20% underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article↗

Microarray expression profiling resources for plant genomics.

Large volumes of genomic data have been generated for several plant species over the past decade, including structural sequence data and functional annotation at the genome level. Various technologies such as expressed sequence tags (ESTs), massively parallel signature sequencing (MPSS) and microarrays have been used to study gene expression and to provide functional data for many genes simultaneously. This review focuses on recent advances in the application of microarrays in plant genomic research and in gene expression databases available for plants. Large sets of Arabidopsis microarray data are publicly available. Recently developed array platforms are currently being used to generate genome-wide expression profiles for several crop species. Coupled to these platforms are public databases that provide access to these large-scale expression data, which can be used to aid the functional discovery of gene function.

Arabidopsis↗

Genes linked by fusion events are generally of the same functional category: a systematic analysis of 30 microbial genomes.

Recent work in computational genomics has shown that a functional association between two genes can be derived from the existence of a fusion of the two as one continuous sequence in another genome. For each of 30 completely sequenced microbial genomes, we established all such fusion links among its genes and determined the distribution of links within and among 15 broad functional categories. We found that 72% of all fusion links related genes of the same functional category. A comparison of the distribution of links to simulations on the basis of a random model further confirmed the significance of intracategory fusion links. Where a gene of annotated function is linked to an unclassified gene, the fusion link suggests that the two genes belong to the same functional category. The predictions based on fusion links are shown here for Methanobacterium thermoautotrophicum, and another 661 predictions are available at http://fusion.bu.edu.

Gene Expression Regulation, Bacterial↗

Localization of binding sites in protein structures by optimization of a composite scoring function.

The rise in the number of functionally uncharacterized protein structures is increasing the demand for structure-based methods for functional annotation. Here, we describe a method for predicting the location of a binding site of a given type on a target protein structure. The method begins by constructing a scoring function, followed by a Monte Carlo optimization, to find a good scoring patch on the protein surface. The scoring function is a weighted linear combination of the z-scores of various properties of protein structure and sequence, including amino acid residue conservation, compactness, protrusion, convexity, rigidity, hydrophobicity, and charge density; the weights are calculated from a set of previously identified instances of the binding-site type on known protein structures. The scoring function can easily incorporate different types of information useful in localization, thus increasing the applicability and accuracy of the approach. To test the method, 1008 known protein structures were split into 20 different groups according to the type of the bound ligand. For nonsugar ligands, such as various nucleotides, binding sites were correctly identified in 55%-73% of the cases. The method is completely automated (http://salilab.org/patcher) and can be applied on a large scale in a structural genomics setting.

Amino Acid Sequence↗

Genome-wide promoter extraction and analysis in human, mouse, and rat.

Large-scale and high-throughput genomics research needs reliable and comprehensive genome-wide promoter annotation resources. We have conducted a systematic investigation on how to improve mammalian promoter prediction by incorporating both transcript and conservation information. This enabled us to build a better multispecies promoter annotation pipeline and hence to create CSHLmpd (Cold Spring Harbor Laboratory Mammalian Promoter Database) for the biomedical research community, which can act as a starting reference system for more refined functional annotations.

Animals↗

ESTIMA, a tool for EST management in a multi-project environment.

BACKGROUND: Single-pass, partial sequencing of complementary DNA (cDNA) libraries generates thousands of chromatograms that are processed into high quality expressed sequence tags (ESTs), and then assembled into contigs representative of putative genes. Usually, to be of value, ESTs and contigs must be associated with meaningful annotations, and made available to end-users. RESULTS: A web application, Expressed Sequence Tag Information Management and Annotation (ESTIMA), has been created to meet the EST annotation and data management requirements of multiple high-throughput EST sequencing projects. It is anchored on individual ESTs and organized around different properties of ESTs including chromatograms, base-calling quality scores, structure of assembled transcripts, and multiple sources of comparison to infer functional annotation, Gene Ontology associations, and cDNA library information. ESTIMA consists of a relational database schema and a set of interactive query interfaces. These are integrated with a suite of web-based tools that allow a user to query and retrieve information. Further, query results are interconnected among the various EST properties. ESTIMA has several unique features. Users may run their own EST processing pipeline, search against preferred reference genomes, and use any clustering and assembly algorithm. The ESTIMA database schema is very flexible and accepts output from any EST processing and assembly pipeline. ESTIMA has been used for the management of EST projects of many species, including honeybee (Apis mellifera), cattle (Bos taurus), songbird (Taeniopygia guttata), corn rootworm (Diabrotica vergifera), catfish (Ictalurus punctatus, Ictalurus furcatus), and apple (Malus x domestica). The entire resource may be downloaded and used as is, or readily adapted to fit the unique needs of other cDNA sequencing projects. CONCLUSIONS: The scripts used to create the ESTIMA interface are freely available to academic users in an archived format from http://titan.biotec.uiuc.edu/ESTIMA/. The entity-relationship (E-R) diagrams and the programs used to generate the Oracle database tables are also available. We have also provided detailed installation instructions and a tutorial at the same website. Presently the chromatograms, EST databases and their annotations have been made available for cattle and honeybee brain EST projects. Non-academic users need to contact the W.M. Keck Center for Functional and Comparative Genomics, University of Illinois at Urbana-Champaign, Urbana, IL, for licensing information.

Animals↗

Identitag, a relational database for SAGE tag identification and interspecies comparison of SAGE libraries.

BACKGROUND: Serial Analysis of Gene Expression (SAGE) is a method of large-scale gene expression analysis that has the potential to generate the full list of mRNAs present within a cell population at a given time and their frequency. An essential step in SAGE library analysis is the unambiguous assignment of each 14 bp tag to the transcript from which it was derived. This process, called tag-to-gene mapping, represents a step that has to be improved in the analysis of SAGE libraries. Indeed, the existing web sites providing correspondence between tags and transcripts do not concern all species for which numerous EST and cDNA have already been sequenced. RESULTS: This is the reason why we designed and implemented a freely available tool called Identitag for tag identification that can be used in any species for which transcript sequences are available. Identitag is based on a relational database structure in order to allow rapid and easy storage and updating of data and, most importantly, in order to be able to precisely define identification parameters. This structure can be seen like three interconnected modules : the first one stores virtual tags extracted from a given list of transcript sequences, the second stores experimental tags observed in SAGE experiments, and the third allows the annotation of the transcript sequences used for virtual tag extraction. It therefore connects an observed tag to a virtual tag and to the sequence it comes from, and then to its functional annotation when available. Databases made from different species can be connected according to orthology relationship thus allowing the comparison of SAGE libraries between species. We successfully used Identitag to identify tags from our chicken SAGE libraries and for chicken to human SAGE tags interspecies comparison. Identitag sources are freely available on http://pbil.univ-lyon1.fr/software/identitag/ web site. CONCLUSIONS: Identitag is a flexible and powerful tool for tag identification in any single species and for interspecies comparison of SAGE libraries. It opens the way to comparative transcriptomic analysis, an emerging branch of biology.

Animals↗

Identification and characterization of protein subcomplexes in yeast.

Protein complexes are major components of cellular organization. Based on large-scale protein complex data, we present the first statistical procedure to find insightful substructures in protein complexes: we identify protein subcomplexes (SCs), i.e., multiprotein assemblies residing in different protein complexes. Four protein complex datasets with different origins and variable reliability are separately analyzed. Our method identifies well-characterized protein assemblies with known functions, thereby confirming the utility of the procedure. In addition, we also identify hitherto unknown functional entities consisting of either functionally unknown proteins or proteins with different functional annotation. We show that SCs represent more reliable protein assemblies than the original complexes. Finally, we demonstrate unique properties of subcomplex proteins that underline the distinct roles of SCs: (i) SCs are functionally and spatially more homogeneous than complete protein complexes (this fact is utilized to predict functional roles and subcellular localizations for so far unannotated proteins); (ii) the abundance of subcomplex proteins is less variable than the abundance of other proteins; (iii) SCs are enriched with essential and synthetic lethal proteins; and (iv) mutations in SC-proteins have higher fitness effects than mutations in other proteins.

Gene Deletion↗

CLIP: a method for identifying protein-RNA interaction sites in living cells.

Nucleic-acid binding proteins constitute nearly one-fourth of all functionally annotated human genes. Genome-wide analysis of protein-nucleic acid contacts has not yet been performed for most of these proteins, restricting attempts to establish a comprehensive understanding of protein function. UV cross-linking is a method typically used to determine the position of direct interactions between proteins and nucleic acids. We have developed the cross-linking and immunoprecipitation assay, which exploits the covalent protein-nucleic acid cross-linking to stringently purify a specific protein-RNA complex using immunoprecipitation followed by SDS-PAGE separation. In this way, the vast majority of non-specific contaminating RNA, which can bind to co-immunoprecipitated proteins or beads, can be removed. Here, we present an improved protocol that performs RNA linker ligation before the SDS-PAGE step, and describe its application to the specific purification and amplification of RNA ligands of Nova in neurons.

Animals↗

A text-mining analysis of the human phenome.

A number of large-scale efforts are underway to define the relationships between genes and proteins in various species. But, few attempts have been made to systematically classify all such relationships at the phenotype level. Also, it is unknown whether such a phenotype map would carry biologically meaningful information. We have used text mining to classify over 5000 human phenotypes contained in the Online Mendelian Inheritance in Man database. We find that similarity between phenotypes reflects biological modules of interacting functionally related genes. These similarities are positively correlated with a number of measures of gene function, including relatedness at the level of protein sequence, protein motifs, functional annotation, and direct protein-protein interaction. Phenotype grouping reflects the modular nature of human disease genetics. Thus, phenotype mapping may be used to predict candidate genes for diseases as well as functional relations between genes and proteins. Such predictions will further improve if a unified system of phenotype descriptors is developed. The phenotype similarity data are accessible through a web interface at http://www.cmbi.ru.nl/MimMiner/.

Chromosome Mapping↗

Prediction of function divergence in protein families using the substitution rate variation parameter alpha.

Protein families typically embody a range of related functions and may thus be decomposed into subfamilies with, for example, distinct substrate specificities. Detection of functionally divergent subfamilies is possible by methods for recognizing branches of adaptive evolution in a gene tree. As the number of genome sequences is growing rapidly, it is highly desirable to automatically detect subfamily function divergence. To this end, we here introduce a method for large-scale prediction of function divergence within protein families. It is called the alpha shift measure (ASM) as it is based on detecting a shift in the shape parameter (alpha [alpha]) of the substitution rate gamma distribution. Four different methods for estimating alpha were investigated. We benchmarked the accuracy of ASM using function annotation from Enzyme Commission numbers within Pfam protein families divided into subfamilies by the automatic tree-based method BETE. In a test using 563 subfamily pairs in 162 families, ASM outperformed functional site-based methods using rate or conservation shifting (rate shift measure [RSM] and conservation shift measure [CSM]). The best results were obtained using the "GZ-Gamma" method for estimating alpha. By combining ASM with RSM and CSM using linear discriminant analysis, the prediction accuracy was further improved.

Algorithms↗