PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The IMB Jena Image Library of biological macromolecules.

The IMB Jena Image Library of Biological Macro-molecules (http://www. imb-jena.de/IMAGE.html ) is aimed at a better dissemination of information on three-dimensional biopolymer structures with an emphasis on visualization and analysis. It provides access to all structure entries deposited at the Protein Data Bank (PDB) and Nucleic Acid Database (NDB). By combining automatic and manual processing it is possible to keep pace with the rapidly growing number of known biopolymer structures and to provide, for selected entries, information not available from automatic procedures. Each entry page contains basic information on the structure, various visualization and analysis tools as well as links to other databases. The visualization techniques adopted include static mono/stereo raster or vector graphics representations, virtual reality modeling (VRML), RasMol/Chime scripts and Java applets. A helix and bending analysis tool provides consistent information on about 750 DNA and RNA duplex structures. Access to metal-containing PDB entries is possible via the Periodic Table of Elements. Finally, general information on amino acids, cis -peptide bonds, structural elements in proteins, base pairs, nucleic acid model conformations and experimental methods for biopolymer structure determination is provided.

Biopolymers↗

Increment of efficiency in the identification of noble genes by colony hybridization assay. (Screening of low redundant clones from human fetal cDNA library).

For the rapid identification of noble genes in a specific tissue by computer analysis from the cDNA sequences determined by single-pass cDNA sequencing, clone redundancy was one of the major obstacles. To facilitate the efficiency in identification of noble genes, it was necessary to reduce the number of clones to be sequenced by eliminating the redundant clones for a rapid analysis. In order to increase the probability of isolating noble sequences from the cDNA clones of human fetal liver tissue origin, colony hybridization assay was adopted and redundant clones were efficiently removed. Four cDNA clones highly redundant in the human fetal liver cDNA libraries including alpha-globin, gamma-globin, serum albumin and H19 RNA sequences were selected as the probes. Two hundreds and sixty two cDNA clones were randomly selected and tested with the probes for hybridization properties. The identity of each cDNA clone giving positive or negative signals in the hybridization assay was determined by DNA homology search with the nucleic acid databases. Among the 76 clones giving positive signals, 57 clones (75%) were found to be identical to the probe sequences and could be eliminated by colony hybridization assay before neucleotide sequencing.

Cloning, Molecular↗

Structure of d(TGCGCA)2 and a comparison to other DNA hexamers.

The X-ray crystal structure of d(TGCGCA)2 has been determined at 120 K to a resolution of 1.3 A. Hexamer duplexes, in the Z-DNA conformation, pack in an arrangement similar to the 'pure spermine form' [Egli et al. (1991). Biochemistry, 30, 11388-11402] but with significantly different cell dimensions. The phosphate backbone exists in two equally populated discrete conformations at one nucleotide step, around phosphate 11. The structure contains two ordered cobalt hexammine molecules which have roles in stabilization of both the Z-DNA conformation of the duplex and in crystal packing. A comparison of d(TGCGCA)2 with other Z-DNA hexamer structures available in the Nucleic Acid Database illustrates the elusive nature of crystal packing. A review of the interactions with the metal cations Na+, Mg2+ and Co3+ reveals a relatively small proportion of phosphate binding and that close contacts between metal ions are common. A prediction of the water structure is compared with the observed pattern in the reported structure.

Binding Sites↗

In vitro binding of cattle PstI SINE with a 33-kDa nuclear protein.

A PstI family of SINEs (short interspersed elements) has been identified in some of the members of the family Bovidae, for example, cattle, buffalo and goat. In vitro DNA-protein interactions were studied to provide a better understanding of the function of these SINEs in the genome. Use of one such cattle PstI interspersed repeat sequence, as a probe in gel retardation assays, has lead to the identification of a repeat DNA-binding factor PIRBP (PstI interspersed repeat binding protein) from cattle liver nuclear extract. Southwestern analysis with liver nuclear extracts from cattle, goat, and buffalo revealed the presence of a PIRBP-like nuclear factor in all three species belonging to the family Bovidae. Deletion analysis localized the PIRBP binding site to an 80-bp (337-417 bp) region within the cattle PstI sequence. UV crosslinking and Southwestern analyses clearly indicated that PIRBP is a singular, small polypeptide of 33-kDa molecular mass. Homology search of the nucleic acids database revealed that the cattle PstI sequence was associated with many different genes of the family Bovidae, either in the 5' flanking region, 5' locus activating region, 3' UTR or in intervening sequences. The binding of the cattle PstI SINE by PIRBP and its association with the regulatory regions of the genes suggests that it plays an important role in the bovine genome.

Animals↗

Global mapping of nucleic acid conformational space: dinucleoside monophosphate conformations and transition pathways among conformational classes.

A global conformational space of 6253 dinucleoside monophosphate (DMP) units consisting of RNA and DNA (free and protein/drug-bound) was 'mapped' using high resolution crystal structures cataloged in the Nucleic Acid Database (NDB). The torsion angles of each DMP were clustered in a reduced three-dimensional space using a classical multi-dimensional scaling method. The mapping of the conformational space reveals nine primary clusters which distinguish among the common A-, B- and Z-forms and their various substates, plus five secondary clusters for kinked or bent structures. Conformational relationships and possible transitional pathways among the substates are also examined using the conformational states of DNA and RNA bound with proteins or drugs as potential pathway intermediates.

Algorithms↗

CORGEN: a FORTRAN-77 generator of standard and non-standard DNA helices from the sequence.

An analytical procedure CORGEN generates a variety of DNA double-stranded structures from user-supplied sequence using a nucleic acid database incorporated into a standard FORTRAN-77 program. Alternatively, the cylindrical polar coordinates of DNA components may be supplied from the external table. An algorithm that performs intercalation sites in DNA is described. This procedure can be used to generate complexes of antibiotics with DNA. Non-standard DNA structures can be built by alternating the global helical twist and global helical rise in the regular DNA helix. The procedures described can be used for computer generation of a variety of non-standard DNA structures which can be subjected to molecular mechanics and dynamics simulations.

Base Sequence↗

Analysis of mRNA expression and protein abundance data: an approach for the comparison of the enrichment of features in the cellular population of proteins and transcripts.

MOTIVATION: Protein abundance is related to mRNA expression through many different cellular processes. Up to now, there have been conflicting results on how correlated the levels of these two quantities are. Given that expression and abundance data are significantly more complex and noisy than the underlying genomic sequence information, it is reasonable to simplify and average them in terms of broad proteomic categories and features (e.g. functions or secondary structures), for understanding their relationship. Furthermore, it will be essential to integrate, within a common framework, the results of many varied experiments by different investigators. This will allow one to survey the characteristics of highly expressed genes and proteins. RESULTS: To this end, we outline a formalism for merging and scaling many different gene expression and protein abundance data sets into a comprehensive reference set, and we develop an approach for analyzing this in terms of broad categories, such as composition, function, structure and localization. As the various experiments are not always done using the same set of genes, sampling bias becomes a central issue, and our formalism is designed to explicitly show this and correct for it. We apply our formalism to the currently available gene expression and protein abundance data for yeast. Overall, we found substantial agreement between gene expression and protein abundance, in terms of the enrichment of structural and functional categories. This agreement, which was considerably greater than the simple correlation between these quantities for individual genes, reflects the way broad categories collect many individual measurements into simple, robust averages. In particular, we found that in comparison to the population of genes in the yeast genome, the cellular populations of transcripts and proteins (weighted by their respective abundances, the transcriptome and what we dub the translatome) were both enriched in: (i) the small amino acids Val, Gly, and Ala; (ii) low molecular weight proteins; (iii) helices and sheets relative to coils; (iv) cytoplasmic proteins relative to nuclear ones; and (v) proteins involved in 'protein synthesis,' 'cell structure,' and 'energy production.' SUPPLEMENTARY INFORMATION: http://genecensus.org/expression/translatome

Algorithms↗

The role of CH/pi interactions implicated in the sequence-dependent deformability of DNA.

The crystal structure of AT-rich deoxynucleotides was retrieved from the Nucleic Acid Database and analyzed with the use of our program CHPI. It has been found that the thymidine 5-methyl group favorably interacts with an adenine ring in the same strand. Since an AT sequence is accompanied with another AT in the complementary strand, the interaction is duplicated, thus forming a twin CH/pi interaction. An AT step becomes stickier than other sequences by the above network. A number of CH/pi contacts have also been found in the crystal structure of A-tract. A successive N/T-methyl stacking in the same strand may contribute in making these steps robust and straight. The role of methyl groups in modified DNA has been discussed on a similar basis.

AT Rich Sequence↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

Prediction of solubility on recombinant expression of Plasmodium falciparum erythrocyte membrane protein 1 domains in Escherichia coli.

BACKGROUND: Cellular interactions elicited by Plasmodium falciparum erythrocyte membrane protein antigen 1 (PfEMP1) are brought about by multiple DBL (Duffy binding like), CIDR (cysteine-rich interdomain region) and C2 domain types. Elucidation of the functional and structural characteristics of these domains is contingent on the abundant availability of recombinant protein in a soluble form. A priori prediction of PfEMP1 domains of the 3D7 genome strain, most likely to be expressed in the soluble form in Escherichia coli was computed and proven experimentally. METHODS: A computational analysis correlating sequence-dependent features to likelihood for expression in soluble form was computed and predictions were validated by the colony filtration blot method for rapid identification of soluble protein expression in E. coli. RESULTS: Solubility predictions for all constituent PfEMP1 domains in the decreasing order of their probability to be expressed in a soluble form (% mean solubility) are as follows: ATS (56.7%) > CIDR1alpha (46.8%) > CIDR2beta (42.9%) > DBL2-4gamma (31.7%) > DBL2beta + C2 (30.6%) > DBL1alpha (24.9%) > DBL2-7epsilon (23.1%) > DBL2-5delta (14.8%). The length of the domains does not correlate to their probability for successful expression in the soluble form. Immunoblot analysis probing for soluble protein confirmed the differential in solubility predictions. CONCLUSION: The acidic terminal segment (ATS) and CIDR alpha/beta domain types are suitable for recombinant expression in E. coli while all DBL subtypes (alpha, beta, gamma, delta, epsilon) are a poor choice for obtaining soluble protein on recombinant expression in E. coli. This study has relevance for researchers pursuing functional and structural studies on PfEMP1 domains.

Animals↗

E-MSD: improving data deposition and structure quality.

The Macromolecular Structure Database (MSD) (http://www.ebi.ac.uk/msd/) [H. Boutselakis, D. Dimitropoulos, J. Fillon, A. Golovin, K. Henrick, A. Hussain, J. Ionides, M. John, P. A. Keller, E. Krissinel et al. (2003) E-MSD: the European Bioinformatics Institute Macromolecular Structure Database. Nucleic Acids Res., 31, 458-462.] group is one of the three partners in the worldwide Protein DataBank (wwPDB), the consortium entrusted with the collation, maintenance and distribution of the global repository of macromolecular structure data [H. Berman, K. Henrick and H. Nakamura (2003) Announcing the worldwide Protein Data Bank. Nature Struct. Biol., 10, 980.]. Since its inception, the MSD group has worked with partners around the world to improve the quality of PDB data, through a clean up programme that addresses inconsistencies and inaccuracies in the legacy archive. The improvements in data quality in the legacy archive have been achieved largely through the creation of a unified data archive, in the form of a relational database that stores all of the data in the wwPDB. The three partners are working towards improving the tools and methods for the deposition of new data by the community at large. The implementation of the MSD database, together with the parallel development of improved tools and methodologies for data harvesting, validation and archival, has lead to significant improvements in the quality of data that enters the archive. Through this and related projects in the NMR and EM realms the MSD continues to improve the quality of publicly available structural data.

Computational Biology↗

Aphid biology: expressed genes from alate Toxoptera citricida, the brown citrus aphid.

The brown citrus aphid, Toxoptera citricida (Kirkaldy), is considered the primary vector of citrus tristeza virus, a severe pathogen which causes losses to citrus industries worldwide. The alate (winged) form of this aphid can readily fly long distances with the wind, thus spreading citrus tristeza virus in citrus growing regions. To better understand the biology of the brown citrus aphid and the emergence of genes expressed during wing development, we undertook a large-scale 5' end sequencing project of cDNA clones from alate aphids. Similar large-scale expressed sequence tag (EST) sequencing projects from other insects have provided a vehicle for answering biological questions relating to development and physiology. Although there is a growing database in GenBank of ESTs from insects, most are from Drosophila melanogaster and Anopheles gambiae, with relatively few specifically derived from aphids. However, important morphogenetic processes are exclusively associated with piercing-sucking insect development and sap feeding insect metabolism. In this paper, we describe the first public data set of ESTs from the brown citrus aphid, T. citricida. The cDNA library was derived from alate adults due to their significance in spreading viruses (e.g., citrus tristeza virus). Over 5180 cDNA clones were sequenced, resulting in 4263 high-quality ESTs. Contig alignment of these ESTs resulted in 2124 total assembled sequences, including both contiguous sequences and singlets. Approximately 33% of the ESTs currently have no significant match in either the non-redundant protein or nucleic acid databases. Sequences returning matches with an E-value of < or = -10 using BLASTX, BLASTN, or TBLASTX were annotated based on their putative molecular function and biological process using the Gene Ontology classification system. These data will aid research efforts in the identification of important genes within insects, specifically aphids and other sap feeding insects within the Order Hemiptera.

Animals↗

Thymine-methyl/pi interaction implicated in the sequence-dependent deformability of DNA.

The crystal structures of deoxy-oligonucleotides were retrieved from the Nucleic Acid Database and analyzed with the use of our program CHPI. The structure of 5'-ApTpApT-3' has been shown to be stabilized by the 5-methyl group in the thymidine moiety that favorably interacts with the adenine pi-ring preceding it. H2' of the deoxyribose in adenine also interacts with the thymine ring next to it. Since a 5'-ApT-3' sequence is accompanied by another 5'-ApT-3' in the complementary strand, the interaction is duplicated, thus forming a 'twin A/T-Me interaction'. Coordinates of oligonucleotides with A-T rich sequences were retrieved and analyzed. In almost every case, the thymidine 5-methyl group favorably interacts with an adenine ring in the same strand. The structure of duplexes incorporating A-tracts was also analyzed. The 5-methyl group in the thymidine moiety has been found to interact favorably with the base pi-ring before it. Since an A-tract is lined with an oligo-T sequence in the complementary strand, a successive N/T-Me stacking may contribute in making the A-tracts robust and straight. The possible involvement of the N/T-Me and the twin A/T-Me motif in the deformability of DNA has been suggested. The role of methyl groups in modified DNA has been discussed on a similar basis.

Crystallography, X-Ray↗

VSD: a database for schizophrenia candidate genes focusing on variations.

Schizophrenia is a common mental disease characterized by delusions, hallucinations, and formal thought disorder. It has been demonstrated with genetic evidence that the disease is a polygenic disorder. Pharmacological, neurochemical, and clinical studies have suggested a number of schizophrenia susceptibility loci. In order to systematically search for genes with small effect in the development of schizophrenia, a database called VSD was established to provide variation data for publicly available candidate genes. Most of the genes encode neurotransmitter receptors, neurotransmitter transporters, and the enzymes involved in their metabolism. Other candidate genes extracted from published literature are also included. The variation information has been collected from publicly available mutation and polymorphism databases such as dbSNP, HGVbase, and OMIM, with single nucleotide polymorphism (SNP) being the most abundant form of collected variations. Reference sequences from NCBI's RefSeq database are used as references when positioning variation at transcript and protein levels. The nonsynonymous SNPs (nsSNPs) that lead to amino acid changes in the functional sites or domains of proteins are distinguished since they are more likely to affect protein function and would be target SNPs for association studies. In addition to variation data, gene descriptions, enzyme information, and other biological information for each gene locus are also included. The latest version of VSD contains 23,648 variations assigned to a total of 186 genes. Five-hundred eighty-eight domains and sites annotated in the SWISS-PROT and InterPro databases are found to contain nsSNPs. VSD may be accessed via the World Wide Web (www.chgb.org.cn/vsd.htm) and will be developed as an up-to-date and comprehensive locus-specific resource for identifying susceptibility genes for schizophrenia.

Databases, Nucleic Acid↗

Water molecules in DNA recognition II: a molecular dynamics view of the structure and hydration of the trp operator.

The structure and hydration of the DNA duplex d-(AGCGTACTAGTACGCT)2 corresponding to the trp operator fragment used in the crystal structure of the half site complex (PDB entry 1TRR) was studied by a 1.4 ns molecular dynamics simulation in water. The simulation, starting from a B-DNA conformation, used a non-bonded cutoff of 1.4 nm with a reaction field correction and resulted in a stable trajectory. The average DNA conformation obtained was closer to the ones found in the crystal structures of the complexes (PDB entries 1TRO and 1TRR) than to the crystal structure of unbound trp operator (Nucleic Acid Database entry BDJ061). The DNA hydration was characterized in terms of hydrogen bond percentages and corresponding residence times. The residence times of water molecules within 0.35 nm of the DNA non-exchangeable protons were calculated for comparison with NMR measurements of intermolecular water-DNA NOEs and nuclear magnetic relaxation dispersion measurements. No significant difference was found between major and minor groove hydration. The DNA donors and acceptors were hydrogen bonded to water molecules for 77(+/-19)% of the time on average. The average residence time of the hydrogen bonded water molecules was 11(+/-11) ps with a maximum of 223 ps. When all water molecules within NOE distance (0.35 nm) of non-exchangeable protons were considered, the average residence times increased to an average of 100(+/-4) ps and a maximum of 608 ps. These results agree with the experimental NMR results of Sunnerhagen et al. which did not show any evidence for water molecules bound with more than 1 ns residence time on the DNA surface. The exchange of hydration water from the DNA occurred in the major groove primarily through direct exchange with the bulk solvent, while access to and from the minor groove frequently proceeded via pathways involving ribose O3' and O4' and phosphate O2P oxygen atoms. The most common water diffusion pathways in the minor groove were perpendicular to the groove direction. In general, water molecules visited only a limited number of sites in the DNA grooves before exiting. The hydrogen bonding sites, where hydrogen bonds could be formed with donor and acceptor groups of the DNA, were filled with water molecules with an average B-factor value of 0.58 mn2. No special values were observed at any of the sites, where water molecules were observed both in the trp repressor/operator co-crystals and in the crystal structure of unbound DNA.

Bacterial Proteins↗

Identification of phosphorylated proteins from thrombin-activated human platelets isolated by two-dimensional gel electrophoresis by electrospray ionization-tandem mass spectrometry (ESI-MS/MS) and liquid chromatography-electrospray ionization-mass spectrometry (LC-ESI-MS).

Two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) is a powerful tool to separate complex protein mixtures including whole cell lysates. In combination with immunoblotting techniques or radioactive labeling techniques it is a fast and convenient way to demonstrate the presence of certain proteins or protein modifications. With the development of extremely sensitive analytical techniques such as matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) or electrospray ionization (ESI)-MS, it has become possible to use 2-D gels not only as an analytical but also as a preparative tool. Starting with a number of spots excised from 2-D gels, a protein can be identified using different strategies involving enzymatic cleavage of the protein in the gel matrix, elution of the resulting peptides and analysis of these peptides by mass spectrometry. The obtained peptide mass fingerprint or fragment ion spectra from peptides can be used to screen protein or nucleic acid databases in order to identify the protein. We have used the techniques described above to identify proteins from human platelets which change their phosphorylation state following activation of platelets by thrombin. Platelets were radioactively labeled with [32P]orthophosphate and stimulated. Several protein spots in the observed range of 10-80 kDa and an isoelectric point of 3-10 showed a significant increase or decrease in phosphorylation. We present the results from the investigation of a spot group representing different isoforms and phosphorylation states of myosin light chain.

Amino Acid Sequence↗

PRIME: a graphical interface for integrating genomic/proteomic databases.

Data mining, finding and integration of information about proteins of interest, is an essential component in modern biological and biomedical research. Even when focusing on a single organism and only on a small number of proteins, there are often dozens fo data sources containing relevant information. We are developing PRIME, a protein information environment, to serve as a virtual central database which integrates distributed heterogeneous information about proteins (linked by common identifier). PRIME has powerful capabilities to visualize all kinds of protein annotation in specialized views. These views can be displayed side by side at the same time and can be synchronized in order to show simultaneously different aspects of identical proteins. These features allow a quick and comprehensive overview of properties of single proteins or protein sets.

Computational Biology↗

The venom of the snake genus Atheris contains a new class of peptides with clusters of histidine and glycine residues.

We investigated venoms from members of the genus Atheris (Serpentes, Viperidae), namely the rough scale bush viper (Atheris squamigera), the green bush viper (A. chlorechis) and the great lakes bush viper (A. nitschei), using mass spectrometry-based strategies, relying on matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry (MALDI-TOF-MS) and electrospray ionisation tandem mass spectrometry (ESI-MS/MS) with de novo peptide sequencing. We discovered a set of novel peptides with masses in the 2-3 kDa range and containing poly-His and poly-Gly segments (pHpG). Complete primary structural elucidation and confirmation of two sequences by Edman degradation indicated the consensus sequence EDDH(9)GVG(10). Bioinformatic investigations in protein sequence databanks did not show relevant homology with known peptides or proteins. However, a more extensive investigation of data in nucleic acid databases revealed some similarities to the precursor sequences of bradykinin potentiating peptides (BPP) and C-type natriuretic peptides (CNP), agents that are known to affect the cardiovascular system by acting on specific metalloproteases and receptors. The novel pHpG peptides found in Atheris venoms might also act on the cardiovascular system by inhibiting particular metalloproteases, which however remain to be identified.

Amino Acid Sequence↗