PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Genome-wide associations of gene expression variation in humans.

The exploration of quantitative variation in human populations has become one of the major priorities for medical genetics. The successful identification of variants that contribute to complex traits is highly dependent on reliable assays and genetic maps. We have performed a genome-wide quantitative trait analysis of 630 genes in 60 unrelated Utah residents with ancestry from Northern and Western Europe using the publicly available phase I data of the International HapMap project. The genes are located in regions of the human genome with elevated functional annotation and disease interest including the ENCODE regions spanning 1% of the genome, Chromosome 21 and Chromosome 20q12-13.2. We apply three different methods of multiple test correction, including Bonferroni, false discovery rate, and permutations. For the 374 expressed genes, we find many regions with statistically significant association of single nucleotide polymorphisms (SNPs) with expression variation in lymphoblastoid cell lines after correcting for multiple tests. Based on our analyses, the signal proximal (cis-) to the genes of interest is more abundant and more stable than distal and trans across statistical methodologies. Our results suggest that regulatory polymorphism is widespread in the human genome and show that the 5-kb (phase I) HapMap has sufficient density to enable linkage disequilibrium mapping in humans. Such studies will significantly enhance our ability to annotate the non-coding part of the genome and interpret functional variation. In addition, we demonstrate that the HapMap cell lines themselves may serve as a useful resource for quantitative measurements at the cellular level.

Chromosome Mapping↗

The GATO gene annotation tool for research laboratories.

Large-scale genome projects have generated a rapidly increasing number of DNA sequences. Therefore, development of computational methods to rapidly analyze these sequences is essential for progress in genomic research. Here we present an automatic annotation system for preliminary analysis of DNA sequences. The gene annotation tool (GATO) is a Bioinformatics pipeline designed to facilitate routine functional annotation and easy access to annotated genes. It was designed in view of the frequent need of genomic researchers to access data pertaining to a common set of genes. In the GATO system, annotation is generated by querying some of the Web-accessible resources and the information is stored in a local database, which keeps a record of all previous annotation results. GATO may be accessed from everywhere through the internet or may be run locally if a large number of sequences are going to be annotated. It is implemented in PHP and Perl and may be run on any suitable Web server. Usually, installation and application of annotation systems require experience and are time consuming, but GATO is simple and practical, allowing anyone with basic skills in informatics to access it without any special training. GATO can be downloaded at [http://mariwork.iq.usp.br/gato/]. Minimum computer free space required is 2 MB.

Biomedical Research↗

From immunogenetics to immunomics: functional prospecting of genes and transcripts.

Human and mouse genome and transcriptome projects have expanded the field of 'immunogenetics' beyond the traditional study of the genetics and evolution of MHC, TCR and Ig loci into the new interdisciplinary area of 'immunomics'. Immunomics is the study of the molecular functions associated with all immune-related coding and non-coding mRNA transcripts. To unravel the function, regulation and diversity of the immunome requires that we identify and correctly categorize all immune-related transcripts. The importance of intercalated genes, antisense transcripts and non-coding RNAs and their potential role in regulation of immune development and function are only just starting to be appreciated. To better understand immune function and regulation, transcriptome projects (e.g. Functional Annotation of the Mouse, FANTOM), that focus on sequencing full-length transcripts from multiple tissue sources, ideally should include specific immune cells (e.g. T cell, B cells, macrophages, dendritic cells) at various states of development, in activated and unactivated states and in different disease contexts. Progress in deciphering immune regulatory networks will require the cooperative efforts of immunologists, immunogeneticists, molecular biologists and bioinformaticians. Although primary sequence analysis remains useful for annotation of new transcripts it is less useful for identifying novel functions of known transcripts in a new context (protein interaction network or pathway). The most efficient approach to mine useful information from the vast a priori knowledge contained in biological databases and the scientific literature, is to use a combination of computational and expert-driven knowledge discovery strategies. This paper will illustrate the challenges posed in attempts to functionally infer transcriptional regulation and interaction of immune-related genes from text and sequence-based data sources.

Alternative Splicing↗

UTRdb and UTRsite: specialized databases of sequences and functional elements of 5' and 3' untranslated regions of eukaryotic mRNAs. Update 2002.

The 5'- and 3'-untranslated regions (5'- and 3'-UTRs) of eukaryotic mRNAs are known to play a crucial role in post-transcriptional regulation of gene expression modulating nucleo-cytoplasmic mRNA transport, translation efficiency, subcellular localization and stability. UTRdb is a specialized database of 5' and 3' untranslated sequences of eukaryotic mRNAs cleaned from redundancy. UTRdb entries are enriched with specialized information not present in the primary databases including the presence of nucleotide sequence patterns already demonstrated by experimental analysis to have some functional role. All these patterns have been collected in the UTRsite database so that it is possible to search any input sequence for the presence of annotated functional motifs. Furthermore, UTRdb entries have been annotated for the presence of repetitive elements. All Internet resources we implemented for retrieval and functional analysis of 5'- and 3'-UTRs of eukaryotic mRNAs are accessible at http://bighost.area.ba.cnr.it/BIG/UTRHome/.

3' Untranslated Regions↗

Assembly, annotation, and integration of UNIGENE clusters into the human genome draft.

The recent release of the first draft of the human genome provides an unprecedented opportunity to integrate human genes and their functions in a complete positional context. However, at least three significant technical hurdles remain: first, to assemble a complete and nonredundant human transcript index; second, to accurately place the individual transcript indices on the human genome; and third, to functionally annotate all human genes. Here, we report the extension of the UNIGENE database through the assembly of its sequence clusters into nonredundant sequence contigs. Each resulting consensus was aligned to the human genome draft. A unique location for each transcript within the human genome was determined by the integration of the restriction fingerprint, assembled genomic contig, and radiation hybrid (RH) maps. A total of 59,500 UNIGENE clusters were mapped on the basis of at least three independent criteria as compared with the 30,000 human genes/ESTs currently mapped in Genemap'99. Finally, the extension of the human transcript consensus in this study enabled a greater number of putative functional assignments than the 11,000 annotated entries in UNIGENE. This study reports a draft physical map with annotations for a majority of the human transcripts, called the Human Index of Nonredundant Transcripts (HINT). Such information can be immediately applied to the discovery of new genes and the identification of candidate genes for positional cloning.

Alleles↗

Discovery of immune-related genes expressed in hemocytes of the tarantula spider Acanthoscurria gomesiana.

The present study reports the identification of immune related transcripts from hemocytes of the spider Acanthoscurria gomesiana by high throughput sequencing of expressed sequence tags (ESTs). To generate ESTs from hemocytes, two cDNA libraries were prepared: one by directional cloning (primary) and the other by the normalization of the first (normalized). A total of 7584 clones were sequenced and the identical ESTs were clustered, resulting in 3723 assembled sequences (AS). At least 20% of these sequences are putative novel genes. The automatic functional annotation of AS based on Gene Ontology revealed several abundant transcripts related to the following functional classes: hemocyanin, lectin, and structural constituents of ribosome and cytoskeleton. From this annotation, 73 transcripts possibly involved in immune response were also identified, suggesting the existence of several molecular processes not previously described for spiders, such as: pathogen recognition, coagulation, complement activation, cell adhesion and intracellular signaling pathway for the activation of cellular defenses.

Amino Acid Sequence↗

Structural proteomics: a tool for genome annotation.

In any newly sequenced genome, 30% to 50% of genes encode proteins with unknown molecular or cellular function. Fortunately, structural genomics is emerging as a powerful approach of functional annotation. Because of recent developments in high-throughput technologies, ongoing structural genomics projects are generating new structures at an unprecedented rate. In the past year, structural studies have identified many new structural motifs involved in enzymatic catalysis or in binding ligands or other macromolecules (DNA, RNA, protein). The efficiency by which function is deduced from structure can be further improved by the integration of structure with bioinformatics and other experimental approaches, such as screening for enzymatic activity or ligand binding.

DNA↗

Protein encoding genes in an ancient plant: analysis of codon usage, retained genes and splice sites in a moss, Physcomitrella patens.

BACKGROUND: The moss Physcomitrella patens is an emerging plant model system due to its high rate of homologous recombination, haploidy, simple body plan, physiological properties as well as phylogenetic position. Available EST data was clustered and assembled, and provided the basis for a genome-wide analysis of protein encoding genes. RESULTS: We have clustered and assembled Physcomitrella patens EST and CDS data in order to represent the transcriptome of this non-seed plant. Clustering of the publicly available data and subsequent prediction resulted in a total of 19,081 non-redundant ORF. Of these putative transcripts, approximately 30% have a homolog in both rice and Arabidopsis transcriptome. More than 130 transcripts are not present in seed plants but can be found in other kingdoms. These potential "retained genes" might have been lost during seed plant evolution. Functional annotation of these genes reveals unequal distribution among taxonomic groups and intriguing putative functions such as cytotoxicity and nucleic acid repair. Whereas introns in the moss are larger on average than in the seed plant Arabidopsis thaliana, position and amount of introns are approximately the same. Contrary to Arabidopsis, where CDS contain on average 44% G/C, in Physcomitrella the average G/C content is 50%. Interestingly, moss orthologs of Arabidopsis genes show a significant drift of codon fraction usage, towards the seed plant. While averaged codon bias is the same in Physcomitrella and Arabidopsis, the distribution pattern is different, with 15% of moss genes being unbiased. Species-specific, sensitive and selective splice site prediction for Physcomitrella has been developed using a dataset of 368 donor and acceptor sites, utilizing a support vector machine. The prediction accuracy is better than those achieved with tools trained on Arabidopsis data. CONCLUSION: Analysis of the moss transcriptome displays differences in gene structure, codon and splice site usage in comparison with the seed plant Arabidopsis. Putative retained genes exhibit possible functions that might explain the peculiar physiological properties of mosses. Both the transcriptome representation (including a BLAST and retrieval service) and splice site prediction have been made available on http://www.cosmoss.org, setting the basis for assembly and annotation of the Physcomitrella genome, of which draft shotgun sequences will become available in 2005.

Alternative Splicing↗

Gene profiling and bioinformatic analysis of Schwann cell embryonic development and myelination.

To elucidate the molecular mechanisms involved in Schwann cell development, we profiled gene expression in the developing and injured rat sciatic nerve. The genes that showed significant changes in expression in developing and dedifferentiated nerve were validated with RT-PCR, in situ hybridisation, Western blot and immunofluorescence. A comprehensive approach to annotating micro-array probes and their associated transcripts was performed using Biopendium, a database of sequence and structural annotation. This approach significantly increased the number of genes for which a functional insight could be found. The analysis implicates agrin and two members of the collapsin response-mediated protein (CRMP) family in the switch from precursors to Schwann cells, and synuclein-1 and alphaB-crystallin in peripheral nerve myelination. We also identified a group of genes typically related to chondrogenesis and cartilage/bone development, including type II collagen, that were expressed in a manner similar to that of myelin-associated genes. The comprehensive function annotation also identified, among the genes regulated during nerve development or after nerve injury, proteins belonging to high-interest families, such as cytokines and kinases, and should therefore provide a uniquely valuable resource for future research.

Agrin↗

The use of evolutionary biology concepts for genome annotation.

The past decade has seen the completion of numerous whole-genome sequencing projects, began with bacterial genomes and continued with eukaryotic species from different phyla: fungi, plants and animals. Besides, more biological information are produced and are shared thanks to information exchange systems, and more biological concepts, as well as more bioinformatics tools, are available. In this article, we will describe how the evolutionary biology concepts, as well as computer science, are useful for a better understanding of biology in general and genome annotation in particular. The genome annotation process consists of taking the raw DNA produced, for example, by the genome sequencing projects, adding the layers of analysis and interpretation necessary to extract its biological significance and placing it in the context of our understanding of biological processes. Genome annotation is a multistep process falling into two broad categories: structural and functional annotation.

Amino Acid Sequence↗

Cardiovascular-related proteins identified in human plasma by the HUPO Plasma Proteome Project pilot phase.

Proteomic profiling of accessible bodily fluids, such as plasma, has the potential to accelerate biomarker/biosignature development for human diseases. The HUPO Plasma Proteome Project pilot phase examined human plasma with distinct proteomic approaches across multiple laboratories worldwide. Through this effort, we confidently identified 3020 proteins, each requiring a minimum of two high-scoring MS/MS spectra. A critical step subsequent to protein identification is functional annotation, in particular with regard to organ systems and disease. Performing exhaustive literature searches, we have manually annotated a subset of these 3020 proteins that have cardiovascular-related functions on the basis of an existing body of published information. These cardiovascular-related proteins can be organized into eight groups: markers of inflammation and/or cardiovascular disease, vascular and coagulation, signaling, growth and differentiation, cytoskeletal, transcription factors, channels/receptors and heart failure and remodeling. In addition, analysis of the peptide per protein ratio for MS/MS identification reveals group-specific trends. These findings serve as a resource to interrogate the functions of plasma proteins, and moreover, the list of cardiovascular-related proteins in plasma constitutes a baseline proteomic blueprint for the future development of biosignatures for diseases such as myocardial ischemia and atherosclerosis.

Arteriosclerosis↗

Functional inferences from blind ab initio protein structure predictions.

Ab initio protein structure prediction methods have improved dramatically in the past several years. Because these methods require only the sequence of the protein of interest, they are potentially applicable to the open reading frames in the many organisms whose sequences have been and will be determined. Ab initio methods cannot currently produce models of high enough resolution for use in rational drug design, but there is an exciting potential for using the methods for functional annotation of protein sequences on a genomic scale. Here we illustrate how functional insights can be obtained from low-resolution predicted structures using examples from blind ab initio structure predictions from the third and fourth critical assessment of structure prediction (CASP3, CASP4) experiments.

Computational Biology↗

Gene expression profiling in porcine mammary gland during lactation and identification of breed- and developmental-stage-specific genes.

A total of 28941 ESTs were sequenced from five 5'-directed non-normalized cDNA libraries, which were assembled into 2212 contigs and 5642 singlets using CAP3. These sequences were annotated and clustered into 6857 unique genes, 2072 of which having no functional annotations were considered as novel genes. These genes were further classified into Gene Ontology categories. By comparing the expression profiles, we identified some breed- and developmental-stage-specific gene groups. These genes may be relative to reproductive performance or play important roles in milk synthesis, secretion and mammary involution. The unknown EST sequences and expression profiles at different developmental stages and breeds are very important resources for further research.

Animals↗

Chromosome-level genome assembly of Cheilinus chlorourus (Bloch, 1791) (Perciformes: Labridae).

In the classification of marine fish, the Labridae family ranks second in terms of species diversity and plays a vital role in coral reef ecosystems, comprising over 600 species across 82 genera. Despite its significance for ecological and evolutionary studies, genomic research on this group has lagged, resulting in a shortage of data, particularly regarding high-quality chromosome-level genome assemblies. To address this gap, this study focused on Cheilinus chlorourus from the Labridae family and successfully achieved a chromosome-level genome assembly. By integrating Illumina, PacBio, and Hi-C sequencing data, we assembled a genome measuring 940.36 Mb, with 926.86 Mb (98.56%) of the gene assembly organized into 21 chromosomes. A total of 29,213 protein-coding genes (PCGs) were identified, and 79.93% of these genes were functionally annotated. With this high-quality genome assembly, future investigations into the functional genomics and ecology of C. chlorourus will have a solid scientific foundation.

Animals↗

Crystallization and preliminary X-ray diffraction studies on the bicupin YwfC from Bacillus subtilis.

A central tenet of evolutionary biology is that proteins with diverse biochemical functions evolved from a single ancestral protein. A variation on this theme is that the functional repertoire of proteins in a living organism is enhanced by the evolution of single-chain multidomain polypeptides by gene-fusion or gene-duplication events. Proteins with a double-stranded beta-helix (cupin) scaffold perform a diverse range of functions. Bicupins are proteins with two cupin domains. There are four bicupins in Bacillus subtilis, encoded by the genes yvrK, yoaN, yxaG and ywfC. The extensive phylogenetic information on these four proteins makes them a good model system to study the evolution of function. The proteins YvrK and YoaN are oxalate decarboxylases, whereas YxaG is a quercetin dioxygenase. In an effort to aid the functional annotation of YwfC as well as to obtain a complete structure-function data set of bicupins, it was proposed to determine the crystal structure of YwfC. The bicupin YwfC was crystallized in two crystal forms. Preliminary crystallographic studies were performed on the diamond-shaped crystals, which belonged to the tetragonal space group P422. These crystals were grown using the microbatch method at 298 K. Native X-ray diffraction data from these crystals were collected to 2.2 A resolution on a home source. These crystals have unit-cell parameters a = b = 68.7, c = 211.5 A. Assuming the presence of two molecules per asymmetric unit, the V(M) value was 2.3 A3 Da(-1) and the solvent content was approximately 45%. Although the crystals appeared less frequently than the tetragonal form, YwfC also crystallizes in the monoclinic space group P2(1), with unit-cell parameters a = 46.7, b = 106.3, c = 48.7 A, beta = 92.7 degrees.

Amino Acid Sequence↗

MILANO--custom annotation of microarray results using automatic literature searches.

BACKGROUND: High-throughput genomic research tools are becoming standard in the biologist's toolbox. After processing the genomic data with one of the many available statistical algorithms to identify statistically significant genes, these genes need to be further analyzed for biological significance in light of all the existing knowledge. Literature mining--the process of representing literature data in a fashion that is easy to relate to genomic data--is one solution to this problem. RESULTS: We present a web-based tool, MILANO (Microarray Literature-based Annotation), that allows annotation of lists of genes derived from microarray results by user defined terms. Our annotation strategy is based on counting the number of literature co-occurrences of each gene on the list with a user defined term. This strategy allows the customization of the annotation procedure and thus overcomes one of the major limitations of the functional annotations usually provided with microarray results. MILANO expands the gene names to include all their informative synonyms while filtering out gene symbols that are likely to be less informative as literature searching terms. MILANO supports searching two literature databases: GeneRIF and Medline (through PubMed), allowing retrieval of both quick and comprehensive results. We demonstrate MILANO's ability to improve microarray analysis by analyzing a list of 150 genes that were affected by p53 overproduction. This analysis reveals that MILANO enables immediate identification of known p53 target genes on this list and assists in sorting the list into genes known to be involved in p53 related pathways, apoptosis and cell cycle arrest. CONCLUSIONS: MILANO provides a useful tool for the automatic custom annotation of microarray results which is based on all the available literature. MILANO has two major advances over similar tools: the ability to expand gene names to include all their informative synonyms while removing synonyms that are not informative and access to the GeneRIF database which provides short summaries of curated articles relevant to known genes. MILANO is available at http://milano.md.huji.ac.il.

Algorithms↗

Dynamic Protein Structure Paradox: An Integrative Framework for Endpoint-Conditioned Evidentiary Sufficiency in Structure-to-Function Claims.

Accurate coordinates for a represented protein state do not, by themselves, establish activity or any other condition-specific function. This article defines the Dynamic Protein Structure Paradox (DPSP) as the apparent conflict between structural accuracy and functional underdetermination and develops it as an integrative evidentiary assessment framework rather than a new theory or paradigm. The underlying problem has been longstanding, since structural genomics, function annotation, allostery, and disorder research each established that fold does not determine function and that function does not determine fold. DPSP consolidates those results into one endpoint-conditioned rule. Once a measurable endpoint is defined, it assesses four coupled dimensions: relevant-state completeness, context completeness, ensemble or kinetic dependence, and chemical dependence. A rubric rates each dimension as adequate, uncertain, or missing, and a materiality test determines which gaps influence the stated decision. The outcome is one of three mutually exclusive modes of utilization: geometry-led, conditional, or function-measured. The deliverable is a concise evidence statement delineating what the structure supports, which decisive variable remains unmeasured, and what corroboration is necessary. DPSP complements, rather than replaces, existing structural, ensemble, and computational approaches. The framework remains unvalidated, its thresholds are provisional, and the studies necessary to confirm or refute it are specified.

Proteins↗

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article↗