PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

Automated genome sequence analysis and annotation.

MOTIVATION: Large-scale genome projects generate a rapidly increasing number of sequences, most of them biochemically uncharacterized. Research in bioinformatics contributes to the development of methods for the computational characterization of these sequences. However, the installation and application of these methods require experience and are time consuming. RESULTS: We present here an automatic system for preliminary functional annotation of protein sequences that has been applied to the analysis of sets of sequences from complete genomes, both to refine overall performance and to make new discoveries comparable to those made by human experts. The GeneQuiz system includes a Web-based browser that allows examination of the evidence leading to an automatic annotation and offers additional information, views of the results, and links to biological databases that complement the automatic analysis. System structure and operating principles concerning the use of multiple sequence databases, underlying sequence analysis tools, lexical analyses of database annotations and decision criteria for functional assignments are detailed. The system makes automatic quality assessments of results based on prior experience with the underlying sequence analysis tools; overall error rates in functional assignment are estimated at 2.5-5% for cases annotated with highest reliability ('clear' cases). Sources of over-interpretation of results are discussed with proposals for improvement. A conservative definition for reporting 'new findings' that takes account of database maturity is presented along with examples of possible kinds of discoveries (new function, family and superfamily) made by the system. System performance in relation to sequence database coverage, database dynamics and database search methods is analysed, demonstrating the inherent advantages of an integrated automatic approach using multiple databases and search methods applied in an objective and repeatable manner. AVAILABILITY: The GeneQuiz system is publicly available for analysis of protein sequences through a Web server at http://www.sander.ebi.ac. uk/gqsrv/submit

Amino Acid Sequence↗

Gene expression profiling of bone marrow stromal cells from juvenile, adult, aged and osteoporotic rats: with an emphasis on osteoporosis.

PURPOSE: Osteoporosis is a multi-factorial, age-related disease with a complex etiology and mode of regulation involving a large numbers of genes. To better understand the possible relationships among genes, we fingerprinted genes in a rat model induced by ovariectomy to determine differences among osteoporotic, non-osteoporotic, aged and juvenile rats. METHODS: We applied genome wide cDNA microarray technology to analyze genes expressed in bone marrow mesenchymal stromal cells (BMSC) and compared non-osteoporotic adult vs. osteoporotic, non-osteoporotic adult vs. aged, and non-osteoporotic adult vs. juvenile. Rigorous statistical analysis of functional annotation (EASE program) identified over-represented biological and molecular functions with significant group wide changes (p< or =0.05). Some of the expressed genes were further confirmed by quantitative RT-PCR (reverse transcription-polymerase chain reaction). RESULTS: Differences in gene expression were observed by identifying transcripts selected by t-test that were consistently changed by a minimum of two-fold. There were 195 transcripts that showed an increased expression and 109 transcripts that showed decreased expression relative to the osteoporotic condition. Of these, 75% transcripts were unknown gene products or ESTs (expressed sequence tag). A number of genes found in the aged and juvenile groups were not present in the osteoporotic rats. Functional clustering of the genes using the EASE bioinformatics program revealed that transcripts in osteoporosis were associated with signal transduction, lipid metabolism, protein metabolism, ionic and protein transport, neuropeptide and G protein signaling pathways. Although some of the genes have previously been shown to play a key role in osteoporosis, several genes were uniquely identified in this study and likely play a role in developing aged related osteoporosis that could have compelling implications in the development of new diagnostic strategies and therapeutics for osteoporosis. CONCLUSIONS: These data suggest that osteoporosis is associated with changes of multiple novel gene expression and that numerous pathways could play important roles in osteoporosis pathogenesis.

Age Factors↗

A genome-wide map of conserved microRNA targets in C. elegans.

BACKGROUND: Metazoan miRNAs regulate protein-coding genes by binding the 3' UTR of cognate mRNAs. Identifying targets for the 115 known C. elegans miRNAs is essential for understanding their function. RESULTS: By using a new version of PicTar and sequence alignments of three nematodes, we predict that miRNAs regulate at least 10% of C. elegans genes through conserved interactions. We have developed a new experimental pipeline to assay 3' UTR-mediated posttranscriptional gene regulation via an endogenous reporter expression system amenable to high-throughput cloning, demonstrating the utility of this system using one of the most intensely studied miRNAs, let-7. Our expression analyses uncover several new potential let-7 targets and suggest a new let-7 activity in head muscle and neurons. To explore genome-wide trends in miRNA function, we analyzed functional categories of predicted target genes, finding that one-third of C. elegans miRNAs target gene sets are enriched for specific functional annotations. We have also integrated miRNA target predictions with other functional genomic data from C. elegans. CONCLUSIONS: At least 10% of C. elegans genes are predicted miRNA targets, and a number of nematode miRNAs seem to regulate biological processes by targeting functionally related genes. We have also developed and successfully utilized an in vivo system for testing miRNA target predictions in likely endogenous expression domains. The thousands of genome-wide miRNA target predictions for nematodes, humans, and flies are available from the PicTar website and are linked to an accessible graphical network-browsing tool allowing exploration of miRNA target predictions in the context of various functional genomic data resources.

Animals↗

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76&#x2009;Gb in length with a contig N50 of 33.53&#x2009;Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals↗

NIFAS: visual analysis of domain evolution in proteins.

MOTIVATION: Multi-domain proteins have evolved by insertions or deletions of distinct protein domains. Tracing the history of a certain domain combination can be important for functional annotation of multi-domain proteins, and for understanding the function of individual domains. In order to analyze the evolutionary history of the domains in modular proteins it is desirable to inspect a phylogenetic tree based on sequence divergence with the modular architecture of the sequences superimposed on the tree. RESULT: A Java applet, NIFAS, that integrates graphical domain schematics for each sequence in an evolutionary tree was developed. NIFAS retrieves domain information from the Pfam database and uses CLUSTAL W to calculate a tree for a given Pfam domain. The tree can be displayed with symbolic bootstrap values, and to allow the user to focus on a part of the tree, the layout can be altered by swapping nodes, changing the outgroup, and showing/collapsing subtrees. NIFAS is integrated with the Pfam database and is accessible over the internet (http://www.cgr.ki.se/Pfam). As an example, we use NIFAS to analyze the evolution of domains in Protein Kinases C.

Computer Graphics↗

Enhancer detection and gene trapping as tools for functional genomics in plants.

Although more than 25,000 genes of Arabidopsis thaliana have been sequenced and mapped, adequate expression or functional information is available for less than 15% of them. In the case of Oryza sativa (rice), about half of more than 55,000 predicted genes have been assigned to a vague functional category on the basis of their sequence, but fewer than 100 have been ascribed a precise, verified function after the identification of a mutant phenotype caused by the molecular disruption of the corresponding gene. Enhancer detection and gene trapping represent insertional mutagenesis strategies that report random expression of many genes and often generate loss-of-function mutations. Several trapping vectors have been designed in a limited number of species, and large-scale enhancer detection and gene trap screens that aim to generate a wide range of spatially and temporally restricted expression patterns have been initiated in both Arabidopsis and rice. These strategies are proving to be essential to the functional annotation of completely sequenced genomes, enabling the analysis of gene function in the context of the entire plant life cycle and substantially expanding our understanding of plant growth and development.

Arabidopsis↗

Molecular characterization of the developmental gene in eyes: through data-mining on integrated transcriptome databases.

OBJECTIVES: Our aim was to utilize publicly available and proprietary sources to discover candidate genes important for ocular development. DESIGN AND METHODS: The collated information on our 5092 non-redundant clusters was grouped and functional annotation was conducted using gene ontology (FatiGO) for categorizing them with respect to molecular function. The web-based viewer technological platform (H-InvDB) was employed for transcription analyses of in-house high quality fetal eye Expressed Sequence Tags (ESTs). Eye-specific ESTs were also analyzed across species by using EMBEST. RESULTS: According to adult eye cDNA libraries, nucleic acid binding and cell structure/cytoskeletal protein genes were the most abundant among the ESTs of fetal eyes. Using cDNA assembly in H-InvDB, 20 (80%) of the 25 most commonly expressed genes in the human eye are also expressed in extraocular tissues. The crystalline gamma S gene is highly expressed in the eye, but not in other tissues. We used EMBEST to compare human fetal eye and octopus eye ESTs and the expression similarity was low (1.6%). This indicated that our fetal eye library contains genes necessary for the developmental process and biological function of the eye, which may not be expressed in the fully developed octopus eyes. The human fetal eye cDNA library also contained highly abundant eye tissue genes, including alphaA-crystallin, eukaryotic translation elongation factor 1 alpha 1 (EEF1A1), bestrophin (VMD2), cystatin C, and transforming growth factor, beta-induced (BIGH3). CONCLUSIONS: Our annotated EST set provides a valuable resource for gene discovery and functional genomic analysis. This display will help to appreciate the strengths and weaknesses of the different technological platforms, so that in future studies the maximum amount of beneficial information can be derived from the appropriate use of each method.

Animals↗

Knowledge-assisted recognition of cluster boundaries in gene expression data.

BACKGROUND AND MOTIVATION: DNA microarray technology has made it possible to determine the expression levels of thousands of genes in parallel under multiple experimental conditions. Genome-wide analyses using DNA microarrays make a great contribution to the exploration of the dynamic state of genetic networks, and further lead to the development of new disease diagnosis technologies. An important step in the analysis of gene expression data is to classify genes with similar expression patterns into the same groups. To this end, hierarchical clustering algorithms have been widely used. Major advantages of hierarchical clustering algorithms are that investigators do not need to specify the number of clusters in advance and results are presented visually in the form of a dendrogram. However, since traditional hierarchical clustering methods simply provide results on the statistical characteristics of expression data, biological interpretations of the resulting clusters are not easy, and it requires laborious tasks to unveil hidden biological processes regulated by members in the clusters. Therefore, it has been a very difficult routine for experts. OBJECTIVE: Here, we propose a novel algorithm in which cluster boundaries are determined by referring to functional annotations stored in genome databases. MATERIALS AND METHODS: The algorithm first performs hierarchical clustering of gene expression profiles. Then, the cluster boundaries are determined by the Variance Inflation Factor among the Gene Function Vectors, which represents distributions of gene functions in each cluster. Our algorithm automatically specifies a cutoff that leads to functionally independent agglomerations of genes on the dendrogram derived from similarities among gene expression patterns. Finally, each cluster is annotated according to dominant gene functions within the respective cluster. RESULTS AND CONCLUSIONS: In this paper, we apply our algorithm to two gene expression datasets related to cell cycle and cold stress response in budding yeast Saccharomyces cerevisiae. As a result, we show that the algorithm enables us to recognize cluster boundaries characterizing fundamental biological processes such as the Early G1, Late G1, S, G2 and M phases in cell cycles, and also provides novel annotation information that has not been obtained by traditional hierarchical clustering methods. In addition, using formal cluster validity indices, high validity of our algorithm is verified by the comparison through other popular clustering algorithms, K-means, self-organizing map and AutoClass.

Algorithms↗

Pattern similarity study of functional sites in protein sequences: lysozymes and cystatins.

BACKGROUND: Although it is generally agreed that topography is more conserved than sequences, proteins sharing the same fold can have different functions, while there are protein families with low sequence similarity. An alternative method for profile analysis of characteristic conserved positions of the motifs within the 3D structures may be needed for functional annotation of protein sequences. Using the approach of quantitative structure-activity relationships (QSAR), we have proposed a new algorithm for postulating functional mechanisms on the basis of pattern similarity and average of property values of side-chains in segments within sequences. This approach was used to search for functional sites of proteins belonging to the lysozyme and cystatin families. RESULTS: Hydrophobicity and beta-turn propensity of reference segments with 3-7 residues were used for the homology similarity search (HSS) for active sites. Hydrogen bonding was used as the side-chain property for searching the binding sites of lysozymes. The profiles of similarity constants and average values of these parameters as functions of their positions in the sequences could identify both active and substrate binding sites of the lysozyme of Streptomyces coelicolor, which has been reported as a new fold enzyme (Cellosyl). The same approach was successfully applied to cystatins, especially for postulating the mechanisms of amyloidosis of human cystatin C as well as human lysozyme. CONCLUSION: Pattern similarity and average index values of structure-related properties of side chains in short segments of three residues or longer were, for the first time, successfully applied for predicting functional sites in sequences. This new approach may be applicable to studying functional sites in un-annotated proteins, for which complete 3D structures are not yet available.

Amino Acid Sequence↗

Automatic detection of subsystem/pathway variants in genome analysis.

MOTIVATION: Proteins work together in pathways and networks, collectively comprising the cellular machinery. A subsystem (a generalization of pathway concept) is a group of related functional roles (such as enzymes) jointly involved in a specific aspect of the cellular machinery. Subsystems provide a natural framework for comparative genome analysis and functional annotation. A subsystem may be implemented in a number of different functional variants in individual species. In order to reliably project functional assignments across multiple genomes, we have to be able to identify the variants implemented in each genome. The analysis of such variants across diverse species is an interesting problem by itself and may provide new evolutionary insights. However, no computational techniques are presently available for an automated detection and analysis of subsystem variants. RESULTS: Here we formulate the subsystem variant detection problem as finding the minimum number of subgraphs of a subsystem, which is represented as a graph, and solve the optimization problem by integer programming approach. The performance of our method was tested on subsystems encoded in the SEED, a genomic integration platform developed by the Fellowship for Interpretation of Genomes as a component of a large-scale effort on comparative analysis and annotation of multiple diverse genomes. Here we illustrate the results obtained for two expert-encoded subsystems of the biosynthesis of Coenzyme A and FMN/FAD cofactors. Applications of variant detection, to support genomic annotations and to assess divergence of species, are briefly discussed in the context of these universally conserved and essential metabolic subsystems. SUPPLEMENTARY INFORMATION: The details of the variant detection results are available at http://ffas.burnham.org/svar/supp.html.

Animals↗

UTRdb and UTRsite: specialized databases of sequences and functional elements of 5' and 3' untranslated regions of eukaryotic mRNAs.

The 5' and 3' untranslated regions of eukaryotic mRNAs may play a crucial role in the regulation of gene expression controlling mRNA localization, stability and translational efficiency. For this reason we developed UTRdb, a specialized database of 5' and 3' untranslated sequences of eukaryotic mRNAs cleaned from redundancy. UTRdb entries are enriched with specialized information not present in the primary databases including the presence of nucleotide sequence patterns already demonstrated by experimental analysis to have some functional role. All these patterns have been collected in the UTRsite database so that it is possible to search any input sequence for the presence of annotated functional motifs. Furthermore, UTRdb entries have been annotated for the presence of repetitive elements. All internet resources implemented for retrieval and functional analysis of 5' and 3' untranslated regions of eukaryotic mRNAs are accessible at http://bigarea.area.ba.cnr.it:8000/EmbIT/UTRH ome/

3' Untranslated Regions↗

ASAP: automated sequence annotation pipeline for web-based updating of sequence information with a local dynamic database.

The automated sequence annotation pipeline (ASAP) is designed to ease routine investigation of new functional annotations on unknown sequences, such as expressed sequence tags (ESTs), through querying of web-accessible resources and maintenance of a local database. The system allows easy use of the output from one search as the input for a new search, as well as the filtering of results. The database is used to store formats and parameters and information for parsing data from web sites. The database permits easy updating of format information should a site modify the format of a query or of a returned web page.

Database Management Systems↗

AnaGram: protein function assignment.

SUMMARY: AnaGram is a web service for protein function assignment based on identity detection of small significant fragments (protomotifs) that can act as modular pieces in peptide construction. The system is able to assign function by finding correlations between protomotifs and functional annotations contained in SWISS-PROT and Medline databases. In addition, function ontologies are used for hierarchical organization of the predicted functions. Extensive tests have been carried out to evaluate the accuracy and performance of the system. AVAILABILITY: http://jaguar.genetica.uma.es/anagram.htm

Algorithms↗

Prediction of the coding sequences of unidentified human genes. XVI. The complete sequences of 150 new cDNA clones from brain which code for large proteins in vitro.

We have carried out a human cDNA sequencing project to accumulate information regarding the coding sequences of unidentified human genes. As an extension of the preceding reports, we herein present the entire sequences of 150 cDNA clones of unknown human genes, named KIAA1294 to KIAA1443, from two sets of size-fractionated human adult and fetal brain cDNA libraries. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here reached 4.8 kb and 2.7 kb (910 amino acid residues), respectively. From sequence similarities and protein motifs, 73 predicted gene products were functionally annotated and 97% of them were classified into the following four functional categories: cell signaling/communication, nucleic acid management, cell structure/motility and protein management. Additionally, the chromosomal loci of the genes were assigned by using human-rodent hybrid panels for those genes whose mapping data were not available in the public databases. The expression profiles of the genes were also studied in 10 human tissues, 8 brain regions, spinal cord, fetal brain and fetal liver by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

MIPS: analysis and annotation of proteins from whole genomes.

The Munich Information Center for Protein Sequences (MIPS-GSF), Neuherberg, Germany, provides protein sequence-related information based on whole-genome analysis. The main focus of the work is directed toward the systematic organization of sequence-related attributes as gathered by a variety of algorithms, primary information from experimental data together with information compiled from the scientific literature. MIPS maintains automatically generated and manually annotated genome-specific databases, develops systematic classification schemes for the functional annotation of protein sequences and provides tools for the comprehensive analysis of protein sequences. This report updates the information on the yeast genome (CYGD), the Neurospora crassa genome (MNCDB), the database of complete cDNAs (German Human Genome Project, NGFN), the database of mammalian protein-protein interactions (MPPI), the database of FASTA homologies (SIMAP), and the interface for the fast retrieval of protein-associated information (QUIPOS). The Arabidopsis thaliana database, the rice database, the plant EST databases (MATDB, MOsDB, SPUTNIK), as well as the databases for the comprehensive set of genomes (PEDANT genomes) are described elsewhere in the 2003 and 2004 NAR database issues, respectively. All databases described, and the detailed descriptions of our projects can be accessed through the MIPS web server (http://mips.gsf.de).

Animals↗

ProSAT2--Protein Structure Annotation Server.

ProSAT2 is a server to facilitate interactive visualization of sequence-based, residue-specific annotations mapped onto 3D protein structures. As the successor of ProSAT (Protein Structure Annotation Tool), it includes its features for visualizing SwissProt and PROSITE functional annotations. Currently, the ProSAT2 server can perform automated mapping of information on variants and mutations from the UniProt KnowledgeBase and the BRENDA enzyme information system onto protein structures. It also accepts and maps user-prepared annotations. By means of an annotation selector, the user can interactively select and group residue-based information according to criteria such as whether a mutation affects enzyme activity. The visualization of the protein structures is based on the WebMol Java molecular viewer and permits simultaneous highlighting of annotated residues and viewing of the corresponding descriptive texts. ProSAT2 is available at http://projects.villa-bosch.de/mcm/database/prosat2/.

Amino Acids↗

TomatEST database: in silico exploitation of EST data to explore expression patterns in tomato species.

TomatEST is a secondary database integrating expressed sequence tag (EST)/cDNA sequence information from different libraries of multiple tomato species. Redundant EST collections from each species are organized into clusters (gene indices). A cluster consists of one or multiple contigs. Multiple contigs in a cluster represent alternatively transcribed forms of a gene. The set of stand-alone EST sequences (singletons) and contigs, representing all the computationally defined 'Transcript Indices', are annotated according to similarity versus protein and RNA family databases. Sequence function description is integrated with the Gene Ontologies and the Enzyme Commission identifiers for a standard classification of gene products and for the mapping of the expressed sequences onto metabolic pathways. Information on the origin of the ESTs, on their structural features, on clusters and contigs, as well as on functional annotations are accessible via a user-friendly web interface. Specific facilities in the database allow Transcript Indices from a query be automatically classified in Enzyme classes and in metabolic pathways. The 'on the fly' mapping onto the metabolic maps is integrated in the analytical tools. The TomatEST database website is freely available at http://biosrv.cab.unina.it/tomatestdb.

Computational Biology↗