PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Joint Task Force Algorithm and Annotations for Diagnosis and Management of Rhinitis.

The algorithm and text annotations in this document are intended to assist clinical decision making about patients who present with symptoms of rhinitis. This document complements the Executive Summary of Joint Task Force Practice Parameters for Diagnosis and Management of Rhinitis (Ann Allergy, Asthma, Immunol 1998; 81:463-468) and Diagnosis and Management of Rhinitis: Complete Guidelines of the Joint Task Force on Practice Parameters in Allergy, Asthma and Immunology (Ann Allergy, Asthma, Immunol 1998;81:478-578). The Joint Task Force on Practice Parameters in Allergy, Asthma and Immunology is co-sponsored by the American Academy of Allergy, Asthma and Immunology, the American College of Allergy, Asthma and Immunology and the Joint Council of Allergy, Asthma and Immunology.

Algorithms↗

Beyond annotation transfer by homology: novel protein-function prediction methods to assist drug discovery.

Every entirely sequenced genome reveals 100 s to 1000 s of protein sequences for which the only annotation available is 'hypothetical protein'. Thus, in the human genome and in the genomes of pathogenic agents there could be 1000 s of potential, unexplored drug targets. Computational prediction of protein function can play a role in studying these targets. We shall review the challenges, research approaches and recently developed tools in the field of computational function-prediction and we will discuss the ways these issues can change the process of drug discovery.

Computational Biology↗

Comparison of rice and Arabidopsis annotation.

Several versions of the rice genome were published in 2002, providing a first overview of the genome content of this model monocot. At the same time, the genome of the model dicot, Arabidopsis thaliana, reached a new level of annotation as thousands of full-length cDNA sequences were integrated with the genome sequence.

Arabidopsis↗

Semantic annotation for concept-based cross-language medical information retrieval.

We present a framework for concept-based cross-language information retrieval in the medical domain, which is under development in the MUCHMORE project. Our approach is based on using the Unified Medical Language System (UMLS) as the primary source of semantic data. Documents and queries are annotated with multiple layers of linguistic information. Linguistic processing includes part-of-speech tagging, morphological analysis, phrase recognition and the identification of medical terms and semantic relations between them. The paper describes experiments in monolingual and cross-language document retrieval, performed on a corpus of medical abstracts. Results show that linguistic processing, especially lemmatization and compound analysis for German, is a crucial step in achieving a good baseline performance. On the other hand, they show that semantic information, specifically the combined use of concepts and relations, increases the performance in monolingual and cross-language retrieval.

Humans↗

Annotating enzymes of unknown function: N-formimino-L-glutamate deiminase is a member of the amidohydrolase superfamily.

The functional assignment of enzymes that catalyze unknown chemical transformations is a difficult problem. The protein Pa5106 from Pseudomonas aeruginosa has been identified as a member of the amidohydrolase superfamily by a comprehensive amino acid sequence comparison with structurally authenticated members of this superfamily. The function of Pa5106 has been annotated as a probablechlorohydrolase or cytosine deaminase. A close examination of the genomic content of P. aeruginosa reveals that the gene for this protein is in close proximity to genes included in the histidine degradation pathway. The first three steps for the degradation of histidine include the action of HutH, HutU, and HutI to convert L-histidine to N-formimino-L-glutamate. The degradation of N-formimino-L-glutamate to L-glutamate can occur by three different pathways. Three proteins in P. aeruginosa have been identified that catalyze two of the three possible pathways for the degradation of N-formimino-L-glutamate. The protein Pa5106 was shown to catalyze the deimination of N-formimino-L-glutamate to ammonia and N-formyl-L-glutamate, while Pa5091 catalyzed the hydrolysis of N-formyl-L-glutamate to formate and L-glutamate. The protein Pa3175 is dislocated from the hut operon and was shown to catalyze the hydrolysis of N-formimino-L-glutamate to formamide and L-glutamate. The reason for the coexistence of two alternative pathways for the degradation of N-formimino-L-glutamate in P. aeruginosa is unknown.

Amidohydrolases↗

sc-PDB: an annotated database of druggable binding sites from the Protein Data Bank.

The sc-PDB is a collection of 6 415 three-dimensional structures of binding sites found in the Protein Data Bank (PDB). Binding sites were extracted from all high-resolution crystal structures in which a complex between a protein cavity and a small-molecular-weight ligand could be identified. Importantly, ligands are considered from a pharmacological and not a structural point of view. Therefore, solvents, detergents, and most metal ions are not stored in the sc-PDB. Ligands are classified into four main categories: nucleotides (< 4-mer), peptides (< 9-mer), cofactors, and organic compounds. The corresponding binding site is formed by all protein residues (including amino acids, cofactors, and important metal ions) with at least one atom within 6.5 angstroms of any ligand atom. The database was carefully annotated by browsing several protein databases (PDB, UniProt, and GO) and storing, for every sc-PDB entry, the following features: protein name, function, source, domain and mutations, ligand name, and structure. The repository of ligands has also been archived by diversity analysis of molecular scaffolds, and several chemoinformatics descriptors were computed to better understand the chemical space covered by stored ligands. The sc-PDB may be used for several purposes: (i) screening a collection of binding sites for predicting the most likely target(s) of any ligand, (ii) analyzing the molecular similarity between different cavities, and (iii) deriving rules that describe the relationship between ligand pharmacophoric points and active-site properties. The database is periodically updated and accessible on the web at http://bioinfo-pharma.u-strasbg.fr/scPDB/.

Algorithms↗

Shotgun annotation of histone modifications: a new approach for streamlined characterization of proteins by top down mass spectrometry.

Eukaryotic histones serve as prototypical examples of posttranslational complexity with diverse modifications (PTMs) on many different residues that comprise a "Histone Code". To help crack this code more efficiently, we demonstrate a new strategy for protein characterization wherein complete PTM descriptions are obtained by database retrieval instead of manual interpretation of information-rich data from high-resolution tandem mass spectrometry (MS/MS). A database of nearly 50 000 modified histone H4 sequences was created and queried with 91 fragment ions from electron capture dissociation of a histone form +112 Da (versus unmodified mass) selectively accumulated in a quadrupole Fourier transform hybrid mass spectrometer. The correct form atop the retrieval list indicated dimethylation at Lys20, acetylation at the N terminus, and acetylation at Lys16 (resolved from trimethylation, Deltam = 0.036 Da). A statistical evaluation reveals the critical role of mass accuracy and that PTM "isomers" are retrieved as next-best matches. The applicability of shotgun annotation to forms of H4 with up to six PTMs is demonstrated, with extensibility to other histones (e.g., H2A, H2B, H3) and other protein classes projected.

Amino Acid Sequence↗

FAST-NMR: functional annotation screening technology using NMR spectroscopy.

An abundance of protein structures emerging from structural genomics and the Protein Structure Initiative (PSI) are not amenable to ready functional assignment because of a lack of sequence and structural homology to proteins of known function. We describe a high-throughput NMR methodology (FAST-NMR) to annotate the biological function of novel proteins through the structural and sequence analysis of protein-ligand interactions. This is based on basic tenets of biochemistry where proteins with similar functions will have similar active sites and exhibit similar ligand binding interactions, despite global differences in sequence and structure. Protein-ligand interactions are determined through a tiered NMR screen using a library composed of compounds with known biological activity. A rapid co-structure is determined by combining the experimental identification of the ligand binding site from NMR chemical shift perturbations with the protein-ligand docking program AutoDock. Our CPASS (Comparison of Protein Active Site Structures) software and database are then used to compare this active site with proteins of known function. The methodology is demonstrated using unannotated protein SAV1430 from Staphylococcus aureus.

Amino Acid Sequence↗

Re-evaluation and in silico annotation of the Tupaia herpesvirus proteins.

Herpesviruses represent an exceptionally suitable model to analyze evolutionary old pathogens, their competency to adapt to existing and changing molecular niches in host species, and the modulation of the gene content and function to comply with the requirements of life. The basis for numerous studies dealing with these questions are reliable statements about the gene content of herpesviral genomes and the functions of viral proteins. The recent determination of the coding strategy of the chimpanzee cytomegalovirus genome and the re-evaluation of the gene content of the human cytomegalovirus genome made it also necessary to restructure the putative transcription map of the Tupaia herpesvirus (THV) genome. Twenty-three THV-specific ORFs formerly predicted to be coding for viral proteins were deleted from the THV transcription map resulting in a gene layout that is now characterized by the presence of conserved genes in the genome center, that probably reflect the genome structure of common herpesviral ancestors, and species-specific genes at the termini. The conserved regions in the THV genome are characterized by high G + C contents between 60% and 80%, a high CpG dinucleotide frequency, and the presence of densely packed putative CpG islands. The genome termini seem to provide the requirements of large scale rearrangements and complements of the gene content to adapt to new environmental demands. With the help of the recently designed method of dictionary-driven, pattern-based protein annotation it was possible to assign putative functions to almost all potential THV proteins, e.g. 123 were found to be putative membrane or secreted proteins, putative signal domains were identified in 69, and 29 proteins were predicted to be glycosylated. The present study adds new aspects to the knowledge about the precise gene composition of herpesvirus genomes and viral protein functions that are of exceptional importance for studies dealing with the phylogeny, the evolution, vaccine vector development, virus-host interactions, pathogenesis and the determination of protein functions of herpesviruses.

Animals↗

Types of corrosion in removable appliances: annotated cases and preventative measures.

Three different types of electrochemical corrosion cells are discussed - composition, concentration, and stress cells. Annotated cases are presented for each type of corrosion cell in removable appliances using conventional photography, scanning electron microscopy, and X-ray elemental analyses. The complexity that corrosion can attain is illustrated via the case of a palatal expansion device. Specific examples of galvanic cells, which are possible in orthodontic practice, are summarized in terms of the anodic and cathodic half-cells. These include material differences, oxygen gradients, presence of debris, changes in pH, induced stress levels, changes in stress states, and variations in as-received history. A galvanic series is constructed from the orthodontist's perspective that includes titanium, nickel-titanium, cobalt-chromium, and stainless steel alloys, as well as glass-fiber-reinforced polymer-matrix composites. Preventative measures are presented from the perspectives of the raw material manufacturers, the dental laboratories, the chair-side practitioner, and the patient. A list of materials is provided, which are known to promote corrosion in the event that stainless steels are sensitized as a consequence of improper handling. In the final analysis, however, retainers and removable appliances are reliable if corrosion cells are conscientiously averted and preventative measures are routinely employed.

Journal Article↗

Analysis of the mouse transcriptome based on functional annotation of 60,770 full-length cDNAs.

Only a small proportion of the mouse genome is transcribed into mature messenger RNA transcripts. There is an international collaborative effort to identify all full-length mRNA transcripts from the mouse, and to ensure that each is represented in a physical collection of clones. Here we report the manual annotation of 60,770 full-length mouse complementary DNA sequences. These are clustered into 33,409 'transcriptional units', contributing 90.1% of a newly established mouse transcriptome database. Of these transcriptional units, 4,258 are new protein-coding and 11,665 are new non-coding messages, indicating that non-coding RNA is a major component of the transcriptome. 41% of all transcriptional units showed evidence of alternative splicing. In protein-coding transcripts, 79% of splice variations altered the protein product. Whole-transcriptome analyses resulted in the identification of 2,431 sense-antisense pairs. The present work, completely supported by physical clones, provides the most comprehensive survey of a mammalian transcriptome so far, and is a valuable resource for functional genomics.

Alternative Splicing↗

The DNA sequence, annotation and analysis of human chromosome 3.

After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

Animals↗

Using the transcriptome to annotate the genome.

A remaining challenge for the human genome project involves the identification and annotation of expressed genes. The public and private sequencing efforts have identified approximately 15,000 sequences that meet stringent criteria for genes, such as correspondence with known genes from humans or other species, and have made another approximately 10,000-20,000 gene predictions of lower confidence, supported by various types of in silico evidence, including homology studies, domain searches, and ab initio gene predictions. These computational methods have limitations, both because they are unable to identify a significant fraction of genes and exons and because they are unable to provide definitive evidence about whether a hypothetical gene is actually expressed. As the in silico approaches identified a smaller number of genes than anticipated, we wondered whether high-throughput experimental analyses could be used to provide evidence for the expression of hypothetical genes and to reveal previously undiscovered genes. We describe here the development of such a method--called long serial analysis of gene expression (LongSAGE), an adaption of the original SAGE approach--that can be used to rapidly identify novel genes and exons.

DNA, Complementary↗

Minimum information requested in the annotation of biochemical models (MIRIAM).

Most of the published quantitative models in biology are lost for the community because they are either not made available or they are insufficiently characterized to allow them to be reused. The lack of a standard description format, lack of stringent reviewing and authors' carelessness are the main causes for incomplete model descriptions. With today's increased interest in detailed biochemical models, it is necessary to define a minimum quality standard for the encoding of those models. We propose a set of rules for curating quantitative models of biological systems. These rules define procedures for encoding and annotating models represented in machine-readable form. We believe their application will enable users to (i) have confidence that curated models are an accurate reflection of their associated reference descriptions, (ii) search collections of curated models with precision, (iii) quickly identify the biological phenomena that a given curated model or model constituent represents and (iv) facilitate model reuse and composition into large subcellular models.

Biochemistry↗

C. elegans ORFeome version 1.1: experimental verification of the genome annotation and resource for proteome-scale protein expression.

To verify the genome annotation and to create a resource to functionally characterize the proteome, we attempted to Gateway-clone all predicted protein-encoding open reading frames (ORFs), or the 'ORFeome,' of Caenorhabditis elegans. We successfully cloned approximately 12,000 ORFs (ORFeome 1.1), of which roughly 4,000 correspond to genes that are untouched by any cDNA or expressed-sequence tag (EST). More than 50% of predicted genes needed corrections in their intron-exon structures. Notably, approximately 11,000 C. elegans proteins can now be expressed under many conditions and characterized using various high-throughput strategies, including large-scale interactome mapping. We suggest that similar ORFeome projects will be valuable for other organisms, including humans.

Alternative Splicing↗

Gene identification signature (GIS) analysis for transcriptome characterization and genome annotation.

We have developed a DNA tag sequencing and mapping strategy called gene identification signature (GIS) analysis, in which 5' and 3' signatures of full-length cDNAs are accurately extracted into paired-end ditags (PETs) that are concatenated for efficient sequencing and mapped to genome sequences to demarcate the transcription boundaries of every gene. GIS analysis is potentially 30-fold more efficient than standard cDNA sequencing approaches for transcriptome characterization. We demonstrated this approach with 116,252 PET sequences derived from mouse embryonic stem cells. Initial analysis of this dataset identified hundreds of previously uncharacterized transcripts, including alternative transcripts of known genes. We also uncovered several intergenically spliced and unusual fusion transcripts, one of which was confirmed as a trans-splicing event and was differentially expressed. The concept of paired-end ditagging described here for transcriptome analysis can also be applied to whole-genome analysis of cis-regulatory and other DNA elements and represents an important technological advance for genome annotation.

5' Flanking Region↗

Correlations between causal effect sizes of proximal SNPs vary with functional annotations and implicate stabilizing selection.

Causal disease effect sizes of proximal single-nucleotide polymorphisms (SNPs) are widely assumed to be independent but could be correlated. Here we introduce a new method, linkage disequilibrium SNP-pair effect correlation regression (LDSPEC), to estimate the correlation of causal disease effect sizes of derived alleles between proximal SNPs; LDSPEC produced robust estimates in simulations. Analyzing 70 UK Biobank diseases and traits (average N&#x2009;=&#x2009;305,646), we detected significantly non-zero SNP-pair effect correlations (for example, -0.37 &#xb1; 0.09 for low-frequency positive linkage disequilibrium 0-100-bp SNP pairs) that decayed with distance and varied with allele frequency and linkage disequilibrium between SNPs. SNP pairs with shared functions had stronger effect correlations that spanned longer genomic distances. Consequently, SNP heritability estimates were smaller than estimates of the sum of causal effect size variances across SNPs, particularly for certain functional annotations. We recapitulated our findings via forward simulations involving stabilizing selection, implicating the action of linkage masking, whereby haplotypes containing linked SNPs with opposite effects on disease have reduced effects on fitness and escape negative selection.

Polymorphism, Single Nucleotide↗