PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

WILMA-automated annotation of protein sequences.

Large-scale annotation of sets of proteins is a frequently occurring task in association with genome sequencing projects. Here, we present an automated platform for the functional annotation of large sets of protein sequences. Various bioinformatics tools are used to achieve a comprehensive description of protein sequences and to link these results to standard Gene Ontology descriptors for molecular function, biological processes and cellular components. Access to the annotation is provided via a web-interface and database queries. These interfaces allow to formulate proteome wide queries as well as the investigation of details of individual results. WILMA annotations of the proteomes of Homo sapiens, Mus musculus, Arabidopsis thaliana and Caenorhabditis elegans are accessible at http://www.came.sbg.ac.at/wilma/

Amino Acid Sequence↗

Microarray expression profiling and functional characterization of AtTPS genes: duplicated Arabidopsis thaliana sesquiterpene synthase genes At4g13280 and At4g13300 encode root-specific and wound-inducible (Z)-gamma-bisabolene synthases.

The Arabidopsis thaliana genome contains at least 32 terpenoid synthase (AtTPS) genes [Aubourg et al., Mol. Genet. Genom. 267 (2002) 730] a few of which have recently been characterized. Based on hierarchical cluster analysis of AtTPS gene expression, measured by microarray profiling and validated with published expression data, we identified two groups of predominantly root expressed AtTPS genes containing five members with previously unknown biochemical functions (At4g13280, At4g13300, At5g48110, At1g33750, and At3g29410). Among the root expressed AtTPS genes, a pair of tandem-organized genes, At4g13280 (AtTPS12) and At4g13300 (AtTPS13), shares 91% predicted amino acid identity indicating recent gene duplication. Bacterial expression of cDNAs and enzyme assays showed that both At4g13280 and At4g13300 encode sesquiterpene synthases catalyzing the conversion of farnesyl diphosphate to (Z)-gamma-bisabolene and the additional minor products E-nerolidol and alpha-bisabolol. Expression of beta-glucuronidase (GUS) reporter gene fused to upstream genomic regions of At4g13280 or At4g13300 showed constitutive promoter activities in the cortex and sub-epidermal layers of Arabidopsis roots. In addition, highly localized promoter activities were found in leaf hydathodes and flower stigmata. Mechanical wounding of Arabidopsis leaves induced local expression of At4g13280 and At4g13300. The functional characterization of At4g13280 gene product AtTPS12 and At4g13230 gene product AtTPS13 as (Z)-gamma-bisabolene synthases, together with the recent characterization of two flower-specific AtTPS [At5g23960 and At5g44630; Tholl et al., Plant J. 42 (2005) 757], concludes the biochemical functional annotation of all four predicted Arabidopsis sesquiterpene synthase genes. Our data suggest biological functions for At4g13280 and At4g13300 in the rhizosphere with additional roles in aerial plant tissues.

Alkyl and Aryl Transferases↗

Structural genomics of proteins from conserved biochemical pathways and processes.

During the past year, X-ray crystallographers and solution NMR spectroscopists have made significant progress towards the complete structural characterization of conserved biochemical pathways and processes. Some of these advances were made in the context of nascent structural genomics programs, which promise to accelerate structural studies of biologically and medically important proteins. The results of high-throughput protein production, crystallization, structure determination, homology modeling and functional annotation published by two such programs have provided insight into the evolution and function of enzymes in the isoprenoid biosynthesis and ribulose monophosphate pathways.

Biological Evolution↗

The bioinformatics resource for oral pathogens.

Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.

Bacteria↗

Genome-wide gene expression in response to parasitoid attack in Drosophila.

BACKGROUND: Parasitoids are insect parasites whose larvae develop in the bodies of other insects. The main immune defense against parasitoids is encapsulation of the foreign body by blood cells, which subsequently often melanize. The capsule sequesters and kills the parasite. The molecular processes involved are still poorly understood, especially compared with insect humoral immunity. RESULTS: We explored the transcriptional response to parasitoid attack in Drosophila larvae at nine time points following parasitism, hybridizing five biologic replicates per time point to whole-genome microarrays for both parasitized and control larvae. We found significantly different expression profiles for 159 probe sets (representing genes), and we classified them into 16 clusters based on patterns of co-expression. A series of functional annotations were nonrandomly associated with different clusters, including several involving immunity and related functions. We also identified nonrandom associations of transcription factor binding sites for three main regulators of innate immune responses (GATA/srp-like, NF-kappaB/Rel-like and Stat), as well as a novel putative binding site for an unknown transcription factor. The appearance or absence of candidate genes previously associated with insect immunity in our differentially expressed gene set was surveyed. CONCLUSION: Most genes that exhibited altered expression following parasitoid attack differed from those induced during antimicrobial immune responses, and had not previously been associated with defense. Applying bioinformatic techniques contributed toward a description of the encapsulation response as an integrated system, identifying putative regulators of co-expressed and functionally related genes. Genome-wide studies such as ours are a powerful first approach to investigating novel genes involved in invertebrate immunity.

Animals↗

Whole-proteome prediction of protein function via graph-theoretic analysis of interaction maps.

MOTIVATION: Determining protein function is one of the most important problems in the post-genomic era. For the typical proteome, there are no functional annotations for one-third or more of its proteins. Recent high-throughput experiments have determined proteome-scale protein physical interaction maps for several organisms. These physical interactions are complemented by an abundance of data about other types of functional relationships between proteins, including genetic interactions, knowledge about co-expression and shared evolutionary history. Taken together, these pairwise linkages can be used to build whole-proteome protein interaction maps. RESULTS: We develop a network-flow based algorithm, FunctionalFlow, that exploits the underlying structure of protein interaction maps in order to predict protein function. In cross-validation testing on the yeast proteome, we show that FunctionalFlow has improved performance over previous methods in predicting the function of proteins with few (or no) annotated protein neighbors. By comparing several methods that use protein interaction maps to predict protein function, we demonstrate that FunctionalFlow performs well because it takes advantage of both network topology and some measure of locality. Finally, we show that performance can be improved substantially as we consider multiple data sources and use them to create weighted interaction networks. AVAILABILITY: http://compbio.cs.princeton.edu/function

Algorithms↗

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant↗

CDS annotation in full-length cDNA sequence.

The identification of coding sequences (CDS) is an important step in the functional annotation of genes. CDS prediction for mammalian genes from genomic sequence is complicated by the vast abundance of intergenic sequence in the genome, and provides little information about how different parts of potential CDS regions are expressed. In contrast, mammalian gene CDS prediction from cDNA sequence offers obvious advantages, yet encounters a different set of complexities when performed on high-throughput cDNA (HTC) sequences, such as the set of 60,770 cDNAs isolated from full-length enriched libraries of the FANTOM2 project. We developed a CDS annotation strategy that uses a variety of different CDS prediction programs to annotate the CDS regions of FANTOM2 cDNAs. These include rsCDS, which uses sequence similarity to known proteins; ProCrest; Longest-ORF and Truncated-ORF, which are ab initio based predictors; and finally, DECODER and NCBI CDS predictor, which use a combination of both principles. Aided by graphical displays of these CDS prediction results in the context of other sequence similarity results for each cDNA, FANTOM2 CDS inspection by curators and follow-up quality control procedures resulted in high quality CDS predictions for a total of 14,345 FANTOM2 clones.

Animals↗

Strong associations between gene function and codon usage.

The association between codon usage and gene function was analyzed in the complete genomes of Eschericia coli, Bacillus subtilis, Lactococcus lactis and Campylobacter jejuni, using the functional annotation provided by NCBI. Two distinctly different ways of quantifying codon usage were used in the analysis. By using contingency tables it was found that for most amino acids a highly significant association with gene function exists for all species, indicating that codon usage at the level of individual amino acids is generally closely coordinated with gene function. By computing the effective number of codons in the annotated genes and comparing the median values in groups of different gene functions it was shown for all species that codon bias gene by gene also differs.

Amino Acids↗

UTRdb: a specialized database of 5' and 3' untranslated regions of eukaryotic mRNAs.

The 5' and 3' untranslated regions of eukaryotic mRNAs may play a crucial role in the regulation of gene expression controlling mRNA localization, stability and translational efficiency. For this reason we developed UTRdb (http://bigarea.area.ba.cnr.it:8000/BioWWW/#U TRdb), a specialized database of 5' and 3' untranslated sequences of eukaryotic mRNAs cleaned from redundancy. UTRdb entries are enriched with specialized information not present in the primary databases including the presence of nucleotide sequence patterns already demonstrated by experimental analysis to have some functional role. All these patterns have been collected in the UTRsite database so that it is possible to search any input sequence for the presence of annotated functional motifs. Furthermore, UTRdb entries have been annotated for the presence of repetitive elements.

3' Untranslated Regions↗

CYGD: the Comprehensive Yeast Genome Database.

The Comprehensive Yeast Genome Database (CYGD) compiles a comprehensive data resource for information on the cellular functions of the yeast Saccharomyces cerevisiae and related species, chosen as the best understood model organism for eukaryotes. The database serves as a common resource generated by a European consortium, going beyond the provision of sequence information and functional annotations on individual genes and proteins. In addition, it provides information on the physical and functional interactions among proteins as well as other genetic elements. These cellular networks include metabolic and regulatory pathways, signal transduction and transport processes as well as co-regulated gene clusters. As more yeast genomes are published, their annotation becomes greatly facilitated using S.cerevisiae as a reference. CYGD provides a way of exploring related genomes with the aid of the S.cerevisiae genome as a backbone and SIMAP, the Similarity Matrix of Proteins. The comprehensive resource is available under http://mips.gsf.de/genre/proj/yeast/.

Binding Sites↗

Schistosoma mansoni: DNA microarray gene expression profiling during the miracidium-to-mother sporocyst transformation.

For the human blood fluke, Schistosoma mansoni, the developmental period that constitutes the transition from miracidium to sporocyst within the molluscan host involves major alterations in morphology and physiology. Although the genetic basis for this transformation process is not well understood, it is likely to be accompanied by changes in gene expression. In an effort to reveal genes involved in this process, we performed a DNA microarray analysis of expressed mRNAs between miracidial and 4 d old in vitro-cultured mother sporocyst stages of S. mansoni. Fluorescently labeled, dsDNA targets were synthesized from miracidia and sporocyst total RNA and hybridized to oligonucleotide DNA microarrays containing 7335 S. mansoni sequences. Fluorescence intensity ratios were statistically compared between five biologically replicated experiments to identify particular transcripts that displayed stage-associated expression within miracidial and sporocyst mRNA populations. A total of 361 sequences showed stage-associated expression in miracidia, while 273 probes displayed sporocyst-associated expression. Differentially expressed mRNAs were annotated with gene ontology terminology based on BLAST homology using high throughput gene ontology functional annotation toolkit (HT-GO-FAT) and clustered using the GOblet GO browser software. A subset of genes displaying stage-associated expression by microarray analyses was verified utilizing real-time quantitative PCR. The use of DNA microarrays for the profiling of gene expression in early-developing S. mansoni larvae provides a starting point for expanding our understanding of the genes that may be involved in the establishment of parasitism and maintenance of infection in these important life cycle stages.

Animals↗

Age-specific hormonal decline is accompanied by transcriptional changes in human sebocytes in vitro.

The importance of hormones in endogenous aging has been displayed by recent studies performed on animal models and humans. To decipher the molecular mechanisms involved in aging we maintained human sebocytes at defined hormone-substituted conditions that corresponded to average serum levels of females from 20 (f20) to 60 (f60) years of age. The corresponding hormone receptor expression was demonstrated by reverse transcription-polymerase chain reaction (RT-PCR), Western blotting and immunocytochemistry. Cells at f60 produced significantly lower lipids than at f20. Increased mRNA and protein levels of c-Myc and increased protein levels of FN1, which have been associated with aging, were detected in SZ95 sebocytes at f60 compared to those detected at f20 after 5 days of treatment. Expression profiling employing a cDNA microarray composed of 15 529 cDNAs identified 899 genes with altered expression levels at f20 vs. f60. Confirmation of gene regulation was performed by real-time RT-PCR. The functional annotation of these genes according to the Gene Ontology identified pathways related to mitochondrial function, oxidative stress, ubiquitin-mediated proteolysis, cell cycle, immune responses, steroid biosynthesis and phospholipid degradation - all hallmarks of aging. Twenty-five genes in common with those identified in aging kidneys and several genes involved in neurodegenerative diseases were also detected. This is the first report describing the transcriptome of human sebocytes and its modification by a cocktail of hormones administered in age-specific levels and provides an in vitro model system, which approximates some of the hormone-dependent changes in gene transcription that occur during aging in humans.

Aging↗

A first-draft human protein-interaction map.

BACKGROUND: Protein-interaction maps are powerful tools for suggesting the cellular functions of genes. Although large-scale protein-interaction maps have been generated for several invertebrate species, projects of a similar scale have not yet been described for any mammal. Because many physical interactions are conserved between species, it should be possible to infer information about human protein interactions (and hence protein function) using model organism protein-interaction datasets. RESULTS: Here we describe a network of over 70,000 predicted physical interactions between around 6,200 human proteins generated using the data from lower eukaryotic protein-interaction maps. The physiological relevance of this network is supported by its ability to preferentially connect human proteins that share the same functional annotations, and we show how the network can be used to successfully predict the functions of human proteins. We find that combining interaction datasets from a single organism (but generated using independent assays) and combining interaction datasets from two organisms (but generated using the same assay) are both very effective ways of further improving the accuracy of protein-interaction maps. CONCLUSIONS: The complete network predicts interactions for a third of human genes, including 448 human disease genes and 1,482 genes of unknown function, and so provides a rich framework for biomedical research.

Databases, Protein↗

Integrating genomic data to predict transcription factor binding.

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

Algorithms↗

Assigning new GO annotations to protein data bank sequences by combining structure and sequence homology.

Accompanying the discovery of an increasing number of proteins, there is the need to provide functional annotation that is both highly accurate and consistent. The Gene Ontology (GO) provides consistent annotation in a computer readable and usable form; hence, GO annotation (GOA) has been assigned to a large number of protein sequences based on direct experimental evidence and through inference determined by sequence homology. Here we show that this annotation can be extended and corrected for cases where protein structures are available. Specifically, using the Combinatorial Extension (CE) algorithm for structure comparison, we extend the protein annotation currently provided by GOA at the European Bioinformatics Institute (EBI) to further describe the contents of the Protein Data Bank (PDB). Specific cases of biologically interesting annotations derived by this method are given. Given that the relationship between sequence, structure, and function is complicated, we explore the impact of this relationship on assigning GOA. The effect of superfolds (folds with many functions) is considered and, by comparison to the Structural Classification of Proteins (SCOP), the individual effects of family, superfamily, and fold.

Algorithms↗

Comparative genomics approaches to identify genomic regions associated with the antimicrobial activity of Pseudomonas protegens PBL3.

The environmental bacterium Pseudomonas protegens PBL3 has antagonistic activity against the plant pathogenic bacterium Burkholderia glumae, an important pathogen in rice. The antimicrobial activity of P. protegens PBL3 was found in the bacteria-free secreted fraction (secretome), but the specific molecules, as well as the genetic basis of that activity, have not been identified. In this study, we integrated genomic information with antimicrobial assays on P. protegens PBL3 and additional six Pseudomonas spp. strains, to identify putative genomic regions in P. protegens PBL3 associated with antimicrobial activity. We hypothesized that Pseudomonas spp. strains with antimicrobial activity against B. glumae have conserved genes with P. protegens PBL3 that are absent in strains lacking activity. Comparative genomics analyses with anvi'o and progressiveMauve, and using P. protegens PBL3 as the reference genome, revealed 188 genes uniquely present in antimicrobial-producing strains. Seven of those genes were annotated as biosynthetic gene clusters predicted to encode secondary metabolites; additional genes were grouped into 25 contiguous clusters with functions annotated as secretion, signal transduction, regulation, transport/efflux, carbohydrate metabolism and one with an additional uncharacterized function. Altogether, this study uncovered a complex and multi-functional network of candidate genes, suggesting that the antimicrobial activity in P. protegens PBL3 is not limited to biosynthetic pathways but also involves additional regulatory, metabolic and export modules to synthesize and deploy antimicrobials.

Pseudomonas↗

Transcriptomal profiling of the cellular transformation induced by Rho subfamily GTPases.

We have used microarray technology to identify the transcriptional targets of Rho subfamily guanosine 5'-triphosphate (GTP)ases in NIH3T3 cells. This analysis indicated that murine fibroblasts transformed by these proteins show similar transcriptomal profiles. Functional annotation of the regulated genes indicate that Rho subfamily GTPases target a wide spectrum of functions, although loci encoding proteins linked to proliferation and DNA synthesis/transcription are upregulated preferentially. Rho proteins promote four main networks of interacting proteins nucleated around E2F, c-Jun, c-Myc and p53. Of those, E2F, c-Jun and c-Myc are essential for the maintenance of cell transformation. Inhibition of Rock, one of the main Rho GTPase targets, leads to small changes in the transcriptome of Rho-transformed cells. Rock inhibition decreases c-myc gene expression without affecting the E2F and c-Jun pathways. Loss-of-function studies demonstrate that c-Myc is important for the blockage of cell-contact inhibition rather than for promoting the proliferation of Rho-transformed cells. However, c-Myc overexpression does not bypass the inhibition of cell transformation induced by Rock blockage, indicating that c-Myc is essential, but not sufficient, for Rock-dependent transformation. These results reveal the complexity of the genetic program orchestrated by the Rho subfamily and pinpoint protein networks that mediate different aspects of the malignant phenotype of Rho-transformed cells.

Amino Acid Substitution↗