PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

SORTEZ: a relational translator for NCBI's ASN.1 database.

The National Center for Biotechnology Information (NCBI) has created a database collection that includes several protein and nucleic acid sequence databases, a biosequence-specific subset of MEDLINE, as well as value-added information such as links between similar sequences. Information in the NCBI database is modeled in Abstract Syntax Notation 1 (ASN.1) an Open Systems Interconnection protocol designed for the purpose of exchanging structured data between software applications rather than as a data model for database systems. While the NCBI database is distributed with an easy-to-use information retrieval system, ENTREZ, the ASN.1 data model currently lacks an ad hoc query language for general-purpose data access. For that reason, we have developed a software package, SORTEZ, that transforms the ASN.1 database (or other databases with nested data structures) to a relational data model and subsequently to a relational database management system (Sybase) where information can be accessed through the relational query language, SQL. Because the need to transform data from one data model and schema to another arises naturally in several important contexts, including efficient execution of specific applications, access to multiple databases and adaptation to database evolution this work also serves as a practical study of the issues involved in the various stages of database transformation. We show that transformation from the ASN.1 data model to a relational data model can be largely automated, but that schema transformation and data conversion require considerable domain expertise and would greatly benefit from additional support tools.

Algorithms↗

Identification and distribution of protein families in 120 completed genomes using Gene3D.

Using a new protocol, PFscape, we undertake a systematic identification of protein families and domain architectures in 120 complete genomes. PFscape clusters sequences into protein families using a Markov clustering algorithm (Enright et al., Nucleic Acids Res 2002;30:1575-1584) followed by complete linkage clustering according to sequence identity. Within each protein family, domains are recognized using a library of hidden Markov models comprising CATH structural and Pfam functional domains. Domain architectures are then determined using DomainFinder (Pearl et al., Protein Sci 2002;11:233-244) and the protein family and domain architecture data are amalgamated in the Gene3D database (Buchan et al., Genome Res 2002;12:503-514). Using Gene3D, we have investigated protein sequence space, the extent of structural annotation, and the distribution of different domain architectures in completed genomes from all kingdoms of life. As with earlier studies by other researchers, the distribution of domain families shows power-law behavior such that the largest 2,000 domain families can be mapped to approximately 70% of nonsingleton genome sequences; the remaining sequences are assigned to much smaller families. While approximately 50% of domain annotations within a genome are assigned to 219 universal domain families, a much smaller proportion (< 10%) of protein sequences are assigned to universal protein families. This supports the mosaic theory of evolution whereby domain duplication followed by domain shuffling gives rise to novel domain architectures that can expand the protein functional repertoire of an organism. Functional data (e.g. COG/KEGG/GO) integrated within Gene3D result in a comprehensive resource that is currently being used in structure genomics initiatives and can be accessed via http://www.biochem.ucl.ac.uk/bsm/cath/Gene3D/.

Amino Acid Sequence↗

SNP databases and pharmacogenetics: great start, but a long way to go.

With the recent publication of the human genome project there has been an explosion of data available for pharmacogenetic research. Web-based databases containing information on single nucleotide polymorphisms (SNPs) are readily accessible to researchers, but there has been little comment on their utility. We used seven major international databases to identify SNPs in 74 genes involved in drug pathways. Very little overlap was seen among the databases, with only eight out of a putative 893 SNPs ( approximately 1%) common to the most commonly used databases. Problems with false positives, secondary to a high degree of homology in gene families, were also observed. These studies suggest researchers limiting their studies to one database would miss a great deal of information. Effort to update compilation databases, such as HGVbase, GeneSNP, PharmGKB, and HOWDY, and the aggressive removal of false positives from all databases is required if these resources are to facilitate the intended growth in pharmacogenetics research.

Amino Acid Sequence↗

Rates of molecular evolution in RNA viruses: a quantitative phylogenetic analysis.

The study of rates of nucleotide substitution in RNA viruses is central to our understanding of their evolution. Herein we report a comprehensive analysis of substitution rates in 50 RNA viruses using a recently developed maximum likelihood phylogenetic method. This analysis revealed a significant relationship between genetic divergence and isolation time for an extensive array of RNA viruses, although more rate variation was usually present among lineages than would be expected under the constraints of a molecular clock. Despite the lack of a molecular clock, the range of statistically significant variation in overall substitution rates was surprisingly narrow for those viruses where a significant relationship between genetic divergence and time was found, as was the case when synonymous sites were considered alone, where the molecular clock was rejected less frequently. An analysis of the ecological and genetic factors that might explain this rate variation revealed some evidence of significantly lower substitution rates in vector-borne viruses, as well as a weak correlation between rate and genome length. Finally, a simulation study revealed that our maximum likelihood estimates of substitution rates are valid, even if the molecular clock is rejected, provided that sufficiently large data sets are analyzed.

Amino Acid Substitution↗

Presence of dUTPase in the various human endogenous retrovirus K (HERV-K) families.

Various retroviruses have been shown to encode dUTPase. The overall phylogeny of dUTPase is unclear, though. The human genome contains a significant amount of human endogenous retroviruses (HERV) representing fossilized sequences of ancient exogenous retroviruses. A few HERV families have been reported to harbor dUTPase domains. We surveyed the various HERV families for the presence of dUTPase and found that ancestors of all HERV-K families but one encoded dUTPase. With two exceptions phylogenetic analysis shows a monophyletic origin of dUTPase for the different HERV-K dUTPases. Sequences of consensus dUTPase domains suggest that the various exogenous ancestors of HERV-K once encoded active enzymes. Our analysis provides informations on dUTPase phylogeny and further shows that endogenous retroviruses provide important informations regarding retrovirus evolution.

Amino Acid Sequence↗

Molecular adaptation in plant hemoglobin, a duplicated gene involved in plant-bacteria symbiosis.

The evolutionary history of the hemoglobin gene family in angiosperms is unusual in that it involves two mechanisms known for potentially generating molecular adaptation: gene duplication and among-species interaction. In plants able to achieve symbiosis with nitrogen-fixing bacteria, class 2 hemoglobin is expressed at high concentrations in nodules and appears to be a key factor for the achievement and regulation of the symbiotic exchange. In this study, we make use of codon models of DNA sequence evolution with the goal of determining the nature of the selective forces which have driven the evolution of this gene. Our results suggest that adaptive evolution occurred during the period of time following the duplication event (functional divergence) and that a change in the selective pressures arose in class 2 hemoglobin in relation to the acquisition of a symbiotic function.

Adaptation, Biological↗

Genomics and variation of ionotropic glutamate receptors: implications for neuroplasticity.

We used two approaches to identify sequence variants in ionotropic glutamate receptor (IGR) genes: high-throughput screening and resequencing techniques, and "information mining" of public (e.g. dbSNP, ENSEMBL) and private (i.e. Celera Discovery System) sequence databases. Each of the 16 known IGRs is represented in these databases, their positions on a canonical physical map are established. Comparisons of mouse, rat, and human sequences revealed substantial conservation among these genes, which are located on different chromosomes but found within syntenic groups of genes. The IGRs are members of a phylogenetically ancient gene family, sharing similarities with glutamate-like receptors in plants. Parsimony analysis of amino acid sequences groups the IGRs into three distinct clades based on ligand-binding specificity and structural features, such as the channel pore and membrane spanning domains. A collection of 38 variants with amino acid changes was obtained by combining screening, resequencing, and informatics approaches for several of the IGR genes. This represents only a fraction of the sequence variation across these genes, but in fact these may constitute a large fraction of the common polymorphisms at these genes and these polymorphisms are a starting point for understanding the role of these variants in function. Genetically influenced human neurobehavioral phenotypes are likely to be linked to IGR genetic variants. Because ionotropic glutamate receptor activation leads to calcium entry, which is fundamental in brain development and in forms of synaptic plasticity essential for learning and memory and is essential for neuronal survival, it is likely that sequence variants in IGR genes may have profound functional roles in neuronal activation and survival mechanisms.

Amino Acid Substitution↗

Prediction of beta-strand packing interactions using the signature product.

The prediction of beta-sheet topology requires the consideration of long-range interactions between beta-strands that are not necessarily consecutive in sequence. Since these interactions are difficult to simulate using ab initio methods, we propose a supplementary method able to assign beta-sheet topology using only sequence information. We envision using the results of our method to reduce the three-dimensional search space of ab initio methods. Our method is based on the signature molecular descriptor, which has been used previously to predict protein-protein interactions successfully, and to develop quantitative structure-activity relationships for small organic drugs and peptide inhibitors. Here, we show how the signature descriptor can be used in a Support Vector Machine to predict whether or not two beta-strands will pack adjacently within a protein. We then show how these predictions can be used to order beta-strands within beta-sheets. Using the entire PDB database with ten-fold cross-validation, we have achieved 74.0% accuracy in packing prediction and 75.6% accuracy in the prediction of edge strands. For the case of beta-strand ordering, we are able to predict the correct ordering accurately for 51.3% of the beta-sheets. Furthermore, using a simple confidence metric, we can determine those sheets for which accurate predictions can be obtained. For the top 25% highest confidence predictions, we are able to achieve 95.7% accuracy in beta-strand ordering. [Figure: see text].

Amino Acid Sequence↗

A genetic and structural analysis of the N-glycosylation capabilities.

The recent draft sequencing of the rice (Oryza sativa) genome has enabled a genetic analysis of the glycosylation capabilities of an agroeconomically important group of plants, the monocotyledons. In this study, we have not only identified genes putatively encoding enzymes involved in N-glycosylation, but have examined by MALDI-TOF MS the structures of the N-glycans of rice and other monocotyledons (maize, wheat and dates; Zea mays, Triticum aestivum and Phoenix dactylifera); these data show that within the plant kingdom the types of N-glycans found are very similar between monocotyledons, dicotyledons and gymnosperms. Subsequently, we constructed expression vectors for the key enzymes forming plant-typical structures in rice, N-acetylglucosaminyltransferase I (GlcNAc-TI; EC 2.4.1.101), core alpha1,3-fucosyltransferase (FucTA; EC 2.4.1.214) and beta1,2-xylosyltransferase (EC 2.4.2.38) and successfully expressed them in Pichia pastoris. Rice GlcNAc-TI, FucTA and xylosyltransferase are therefore the first monocotyledon glycosyltransferases involved in N-glycan biosynthesis to be characterised in a recombinant form.

Amino Acid Sequence↗

Molecular cloning and characterization of a full-length flavin-dependent monooxygenase from yeast.

Eucaryotes contain a class of enzymes called flavin-dependent monooxygenases (FMOs). Unlike mammals, yeast have only a single isoform-yFMO. Deletion mutants suggested that yFMO may play a role in folding proteins which contain disulfide bonds. Recently we detected two nucleotide errors in the GenBank sequences attributed to the yFMO gene. This previously led us to express and characterize a 373-residue catalytically active protein instead of the correct 432-residue enzyme. Here we report the sequencing, expression, and enzyme characterization of the full-length form of yFMO. Comparison of the two forms of yFMO showed similar pH profiles and K(m), K(cat), and V(max) values using glutathione as a substrate. These results indicate that the full-length yeast FMO has biochemical and catalytic properties similar to those of the truncated protein. Therefore, it is likely that the hypotheses concerning the enzyme's function proposed earlier are still valid.

Amino Acid Sequence↗

A structural and primary sequence comparison of the viral RNA-dependent RNA polymerases.

A systematic bioinformatic approach to identifying the evolutionarily conserved regions of proteins has verified the universality of a newly described conserved motif in RNA-dependent RNA polymerases (motif F). In combination with structural comparisons, this approach has defined two regions that may be involved in unwinding double-stranded RNA (dsRNA) for transcription. One of these is the N-terminal portion of motif F and the second is a large insertion in motif F present in the RNA-dependent RNA polymerases of some dsRNA viruses.

Amino Acid Motifs↗

Gene discovery and expression profile analysis through sequencing of expressed sequence tags from different developmental stages of the chytridiomycete Blastocladiella emersonii.

Blastocladiella emersonii is an aquatic fungus of the chytridiomycete class which diverged early from the fungal lineage and is notable for the morphogenetic processes which occur during its life cycle. Its particular taxonomic position makes this fungus an interesting system to be considered when investigating phylogenetic relationships and studying the biology of lower fungi. To contribute to the understanding of the complexity of the B. emersonii genome, we present here a survey of expressed sequence tags (ESTs) from various stages of the fungal development. Nearly 20,000 cDNA clones from 10 different libraries were partially sequenced from their 5' end, yielding 16,984 high-quality ESTs. These ESTs were assembled into 4,873 putative transcripts, of which 48% presented no matches with existing sequences in public databases. As a result of Gene Ontology (GO) project annotation, 1,680 ESTs (35%) were classified into biological processes of the GO structure, with transcription and RNA processing, protein biosynthesis, and transport as prevalent processes. We also report full-length sequences, useful for construction of molecular phylogenies, and several ESTs that showed high similarity with known proteins, some of which were not previously described in fungi. Furthermore, we analyzed the expression profile (digital Northern analysis) of each transcript throughout the life cycle of the fungus using Bayesian statistics. The in silico approach was validated by Northern blot analysis with good agreement between the two methodologies.

Amino Acid Sequence↗

Nonribosomal peptide synthetase genes in the genome of Fusarium graminearum, causative agent of wheat head blight.

Fungal nonribosomal peptide synthetases (NRPSs) are responsible for the biosynthesis of numerous metabolites which serve as virulence factors in several plant-pathogen interactions. The aim of our work was to investigate the diversity of these genes in a Fusarium graminearum sequence database using bioinformatic techniques. Our search identified 15 NRPS sequences, among which two were found to be closely related to peptide synthetases of various fungi taking part in ferrichrome biosynthesis. Another peptide synthetase gene was similar to that identified in Aspergillus oryzae which is possibly responsible for the biosynthesis of fusarinine, an extracellular iron-chelating siderophore. To our knowledge, this is the first report on the identification of a putative NRPS gene possibly responsible for the biosynthesis of fusarinine-type siderophores. The other NRPSs were found to be related to peptide synthetases taking part in the biosynthesis of various peptides in other fungi. Transcription factors carrying ankyrin repeats were observed in the vicinity of four of the identified peptide synthetase genes. Additionally, NRPS related genes similar to putative long-chain fatty acid CoA ligases, acyl CoA ligases, ABC transport proteins, a highly conserved putative transmembrane protein of Aspergillus nidulans, and alpha-aminoadipate reductases have also been identified. Further studies are in progress to clarify the role of some of the identified NRPS genes in plant pathogenesis.

Amino Acid Sequence↗

The universe of exons revisited.

We study the distribution of exons in eukaryotic genes to determine whether one can detect the reuse of exon sequences and to use the frequency of such reuse to estimate how many ancestral exon sequences there might have been. We use two databases of exons. One contained 56,276 internal exons from putatively unrelated genes (less than 20% sequence identity) and the second contained 8917 internal exons from regions of these genes that are homologous and colinear with prokaryotic genes; these are ancient conserved regions (ACRs). At the 95% significance level we find 3500 exon-sequence matches in the large database and 500 matches in the ACR database. These matches correspond to groups of similar sequences. The size-rank relationship for these groups follows a power law, the size falling off as the inverse square root of the rank. This form of the power law distribution leads us to make an estimate for the size of a possible universe of ancestral exons. Using the data corresponding to the ACR regions, that universe is estimated to be about 15,000-30,000 in size.

Amino Acid Sequence↗

PDBsum more: new summaries and analyses of the known 3D structures of proteins and nucleic acids.

PDBsum is a database of mainly pictorial summaries of the 3D structures of proteins and nucleic acids in the Protein Data Bank. Its pages aim to provide an at-a-glance view of the contents of every 3D structure, plus detailed structural analyses of each protein chain, DNA-RNA chain and any bound ligands and metals. In the past year, the database has been significantly improved, in terms of both appearance and new content. Moreover, it has moved to its new address at http://www.ebi.ac.uk/thornton-srv/databases/pdbsum.

Databases, Protein↗

Rod and cone opsin families differ in spectral tuning domains but not signal transducing domains as judged by saturated evolutionary trace analysis.

The visual receptor of rods and cones is a covalent complex of the apoprotein, opsin, and the light-sensitive chromophore, 11-cis-retinal. This pigment must fulfill many functions including photoactivation, spectral tuning, signal transmission, inactivation, and chromophore regeneration. Rod and cone photoreceptors employ distinct families of opsins. Although it is well known that these opsin families provide unique ranges in spectral sensitivity, it is unclear whether the families have additional functional differences. In this study, we use evolutionary trace (ET) analysis of 188 vertebrate opsin sequences to identify functionally important sites in each opsin family. We demonstrate the following results. (1) The available vertebrate opsin sequences produce a definitive description of all five vertebrate opsin families. This is the first demonstration of sequence saturation prior to ET analysis, which we term saturated ET (SET). (2) The cone opsin classes have class-specific sites compared to the rod opsin class. These sites reside in the transmembrane region and tune the spectral sensitivity of each opsin class to its characteristic wavelength range. (3) The cytoplasmic loops, primarily responsible for signal transmission and inactivation, are essentially invariant in rod versus cone opsins. This indicates that the electrophysiological differences between rod and cone photoreceptors cannot be ascribed to differences in the protein interaction regions of the opsins. SET shows that chromophore binding and regeneration are the only aspects of opsin structure likely to have functionally significant differences between rods and cones, whereas excitatory and adaptational properties of the opsin families appear to be functionally invariant.

Amino Acid Sequence↗