PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Generalized protein tertiary structure recognition using associative memory Hamiltonians.

In previous papers, a method of protein tertiary structure recognition was described based on the construction of an associative memory Hamiltonian, which encoded the amino acid sequence and the C alpha co-ordinates of a set of database proteins. Using molecular dynamics with simulated annealing, the ability of the Hamiltonian to successfully recall the structure of a protein in the memory database was successfully demonstrated, as long as the total number of database proteins did not exceed a characteristic value, called the capacity of the Hamiltonian, equal to 0.5N to 0.7N, where N is the number of amino acid residues in the protein to be recalled. In this paper, we describe the development of additional methods to increase the capacity of the Hamiltonian, including use of a more complete representation of the protein backbone and the incorporation of contextual information into the Hamiltonian through the use of secondary structure prediction. In addition, we further extend the ability of associative memory models to predict the tertiary structures of proteins not present in the protein data set, by making the Hamiltonian invariant with respect to biological symmetries that represent site mutations and insertions and deletions. The ability of the Hamiltonian to generalize from homologous proteins to an unknown protein in the presence of other unrelated proteins in the data set is demonstrated.

Cytochromes↗

AvGI, an index of genes transcribed in the salivary glands of the ixodid tick Amblyomma variegatum.

Random clones from a cDNA library made from mRNA purified from dissected salivary glands of feeding female Amblyomma variegatum ticks were subjected to single pass sequence analysis. A total of 3992 sequences with an average read length of 580 nucleotides have been used to construct a gene index called AvGI that consists of 2109 non-redundant sequences. A provisional gene identity has been assigned to 39% of the database entries by sequence similarity searches against a non-redundant amino acid database and a protein database that has been assigned gene ontology terms. Homologs of genes encoding basic cellular functions including previously characterised enzyme activities, such as stearoyl CoA saturase and protein phosphatase, of ixodid tick salivary glands were found. Several families of abundant cDNA sequences that may code for protein components of tick cement and A. variegatum proteins which may contribute to anti-haemostatic and anti-inflammatory responses, and, one with potential immunosuppressive activity, were also identified. Interference with the function of such proteins might disrupt the life cycle of A. variegatum and help to control this ectoparasite or to reduce its ability to transmit disease causing organisms. AvGI represents an electronic knowledge base, which can be used to launch investigations of the biology of the salivary glands of this tick species. The database may be accessed via the World Wide Web at http://www.tigr.org/tdb/tgi.shtml.

Amino Acid Sequence↗

Nomenclature and structural biology of allergens.

Purified allergens are named using the systematic nomenclature of the Allergen Nomenclature Sub-Committee of the World Health Organization and International Union of Immunological Societies. The system uses abbreviated Linnean genus and species names and an Arabic number to indicate the chronology of allergen purification. Most major allergens from mites, animal dander, pollens, insects, and foods have been cloned, and more than 40 three-dimensional allergen structures are in the Protein Database. Allergens are derived from proteins with a variety of biologic functions, including proteases, ligand-binding proteins, structural proteins, pathogenesis-related proteins, lipid transfer proteins, profilins, and calcium-binding proteins. Biologic function, such as the proteolytic enzyme allergens of dust mites, might directly influence the development of IgE responses and might initiate inflammatory responses in the lung that are associated with asthma. Intrinsic structural or biologic properties might also influence the extent to which allergens persist in indoor and outdoor environments or retain their allergenicity in the digestive tract. Analyses of the protein family database suggest that the universe of allergens comprises more than 120 distinct protein families. Structural biology and proteomics define recombinant allergen targets for diagnostic and therapeutic purposes and identify motifs, patterns, and structures of immunologic significance.

Air Pollution, Indoor↗

The Database of Interacting Proteins: 2004 update.

The Database of Interacting Proteins (http://dip.doe-mbi.ucla.edu) aims to integrate the diverse body of experimental evidence on protein-protein interactions into a single, easily accessible online database. Because the reliability of experimental evidence varies widely, methods of quality assessment have been developed and utilized to identify the most reliable subset of the interactions. This CORE set can be used as a reference when evaluating the reliability of high-throughput protein-protein interaction data sets, for development of prediction methods, as well as in the studies of the properties of protein interaction networks.

Animals↗

DIP: The Database of Interacting Proteins: 2001 update.

The Database of Interacting Proteins (DIP; http://dip.doe-mbi.ucla. edu) is a database that documents experimentally determined protein-protein interactions. Since January 2000 the number of protein-protein interactions in DIP has nearly tripled to 3472 and the number of proteins to 2659. New interactive tools have been developed to aid in the visualization, navigation and study of networks of protein interactions.

Databases, Factual↗

GCRDb: a G-protein-coupled receptor database.

G-protein-coupled receptors (GCRs) are a very large protein family and are critical components in many different autocrine, paracrine and endocrine signaling systems in animals. Current estimates are that humans have several thousand GCRs encoded by only slightly fewer genes. I have developed GCRDb, a database of sequences and other data relevant to the study of the biology of the receptors. The database implementation, data collection, and query system are described. An integral component of the design of GCRDb is the classification of GCRs in families. The current composition of GCRDb is presented in a table showing the number of entries in each family and group as derived from accepted mutation parsimony analyses.

Amino Acid Sequence↗

An evaluation of the use of databases in protein structure refinement.

The speed of electron-density fitting during X-ray structure solution and refinement, and the quality of the protein model resulting, can both be enhanced by the use of databases of main- and side-chain conformations. Three structures are compared in this report, one refined at high resolution (1.7 A), and two at lower resolutions using either the database method (2.4 A resolution) or more traditional empirical electron-density fitting (1.9 A resolution). An analysis of peptide orientation was used as an aid in finding unusual portions of main-chain structure. The fit of side chains to known rotamer conformations was used to help determine the accuracy of these atomic positions. In addition, the use of an objective measure of the fit of structures to electron-density maps was evaluated, both alone and in combination with side-chain conformational information.

Journal Article↗

Stacking and T-shape competition in aromatic-aromatic amino acid interactions.

The potential of mean force of interacting aromatic amino acids is calculated using molecular dynamics simulations. The free energy surface is determined in order to study stacking and T-shape competition for phenylalanine-phenylalanine (Phe-Phe), phenylalanine-tyrosine (Phe-Tyr), and tyrosine-tyrosine (Tyr-Tyr) complexes in vacuo, water, carbon tetrachloride, and methanol. Stacked structures are favored in all solvents with the exception of the Tyr-Tyr complex in carbon tetrachloride, where T-shaped structures are also important. The effect of anchoring the two alpha-carbons (C(alpha)) at selected distances is investigated. We find that short and large C(alpha)-C(alpha) distances favor stacked and T-shaped structures, respectively. We analyze a set of 2396 protein structures resolved experimentally. Comparison of theoretical free energies for the complexes to the experimental analogue shows that Tyr-Tyr interaction occurs mainly at the protein surface, while Phe-Tyr and Phe-Phe interactions are more frequent in the hydrophobic protein core. This is confirmed by the Voronoi polyhedron analysis on the database protein structures. As found from the free energy calculation, analysis of the protein database has shown that proximal and distal interacting aromatic residues are predominantly stacked and T-shaped, respectively.

Computer Simulation↗

SCOPEC: a database of protein catalytic domains.

MOTIVATION: Domains are the units of protein structure, function and evolution. It is therefore essential to utilize knowledge of domains when studying the evolution of function, or when assigning function to genome sequence data. For this purpose, we have developed a database of catalytic domains, SCOPEC, by combining structural domain information from SCOP, full-length sequence information from Swiss-Prot, and verified functional information from the Enzyme Classification (EC) database. Two major problems need to be overcome to create a database of domain-function relationships; (1) for sequences, EC numbers are typically assigned to whole sequences rather than the functional unit, and (2) The Protein Data Bank (PDB) structures elucidated from a larger multi-domain protein will often have EC annotation although the relevant catalytic domain may lie elsewhere. RESULTS: SCOPEC entries have high quality enzyme assignments; having passed both computational and manual checks. SCOPEC currently contains entries for 75% of all EC annotations in the PDB. Overall, EC number is fairly well conserved within a superfamily, even when the proteins are distantly related. Initial analysis is encouraging; suggesting that there is a 50:50 chance of conserved function in distant homologues first detected by a third iteration PSI-BLAST search. Therefore, we envisage that a knowledge-based approach to function assignment using the domain-EC relationships in SCOPEC will gain a marked improvement over this base line. AVAILABILITY: The SCOPEC database is a valuable resource in the analysis and prediction of protein structure and function. It can be obtained or queried at our website http://www.enzome.com

Catalysis↗

Prospective highlights of functional skin proteomics.

Although a wide variety of protein profiles have been extensively constructed via proteomic analysis, the comprehensive proteomic profiling of the skin, which is considered to be the largest organ of the human body, is still far from complete. Our efforts to establish the functional skin proteome, a protein database describing the protein networks that underlie biological processes, has set in motion the identification and characterization of proteins expressed in the epidermis and dermis of the BALB/c mice. In this review, we will highlight various cutaneous proteins we have characterized and discuss their biological functions associated with skin distress, immunity, and cancer. This type of research into functional skin proteomics will provide a critical step toward understanding disease and developing successful therapeutic strategies.

Animals↗

Probabilistic approach to determining unbiased random-coil carbon-13 chemical shift values from the protein chemical shift database.

We describe a probabilistic model for deriving, from the database of assigned chemical shifts, a set of random coil chemical shift values that are "unbiased" insofar as contributions from detectable secondary structure have been minimized (RCCSu). We have used this approach to derive a set of RCCSu values for 13Calpha and 13Cbeta for 17 of the 20 standard amino acid residue types by taking advantage of the known opposite conformational dependence of these parameters. We present a second probabilistic approach that utilizes the maximum entropy principle to analyze the database of 13Calpha and 13Cbeta chemical shifts considered separately; this approach yielded a second set of random coil chemical shifts (RCCSmax-ent). Both new approaches analyze the chemical shift database without reference to known structure. Prior approaches have used either the chemical shifts of small peptides assumed to model the random coil state (RCCSpeptide) or statistical analysis of chemical shifts associated with structure not in helical or strand conformation (RCCSstruct-stat). We show that the RCCSmax-ent values are strikingly similar to published RCCSpeptide and RCCSstruct-stat values. By contrast, the RCCSu values differ significantly from both published types of random coil chemical shift values. The differences (RCCSpeptide - RCCSu) for individual residue types show a correlation with known intrinsic conformational propensities. These results suggest that random coil chemical shift values from both prior approaches are biased by conformational preferences. RCCSu values appear to be consistent with the current concept of the "random coil" as the state in which the geometry of the polypeptide ensemble samples the allowed region of (phi, psi)-space in the absence of any dominant stabilizing interactions and thus represent an improved basis for the detection of secondary structure. Coupled with the growing database of chemical shifts, this probabilistic approach makes it possible to refine relationships among chemical shifts, their conformational propensities, and their dependence on pH, temperature, or neighboring residue type.

Carbon Isotopes↗

Molecular cloning and expression of a porcine chondrocyte nucleotide pyrophosphohydrolase.

The porcine 127-kDa nucleotide pyrophosphohydrolase (NTPPHase) had been previously purified from the conditioned culture media of porcine articular cartilage. Protein sequencing of an internal 61-kDa proteolytic fragment of NTPPHase (61-kDa NTPPHase) determined the 26 N-terminal amino acids. This sequence was used to amplify a DNA fragment, which was used as a probe to clone the gene encoding the 61-kDa NTPPHase from a porcine chondrocyte cDNA library. DNA sequence analysis showed the cDNA insert to be 2509 bp, corresponding to a predicted open reading frame (ORF) encoding 599 amino acids. The 26 N-terminal amino acids of the 61-kDa NTPPHase were located within the ORF immediately downstream of a putative protease recognition region, RRKRR. This is consistent with this cDNA insert representing an internal proteolytic fragment of the full length 127-kDa NTPPHase. BLAST and FASTA analysis confirmed that the deduced amino acid sequence of 61-kDa NTPPHase was unique and did not possess a high degree of homology to sequence in the non-redundant protein and nucleotide databases. Proteins that possess limited homology (< 17%) with the 61-kDa NTTPPHase include several prokaryotic and eukaryotic ATP pyrophosphate-lyases (adenylate cyclase). Northern blot analysis of porcine chondrocyte RNA showed that the DNA encoding the 61-kDa NTPPHase hybridized to a single 4.0-kb RNA transcript. This DNA probe also hybridized to a single species of human chondrocyte RNA. Expression of a 61-kDa protein was detected by coupled in-vitro transcription/translation. Western blot analysis of this in-vitro transcription/translation reaction detected a 61-kDa protein, using an antibody raised against the peptide sequence that was originally used to clone the 61-kDa NTPPHase. These data indicate the successful in-vitro cloning and expression of the porcine chondrocyte 61-kDa NTPPHase. Future studies that utilize the gene encoding the 61-kDa NTPPHase may allow the characterization of the role of NTPPHase in calcium pyrophosphate dihydrate (CPPD) crystal deposition disease.

Amino Acid Sequence↗

Arabidopsis thaliana defense-related protein ELI3 is an aromatic alcohol:NADP+ oxidoreductase.

We expressed a cDNA encoding the Arabidopsis thaliana defense-related protein ELI3-2 in Escherichia coli to determine its biochemical function. Based on a protein database search, this protein was recently predicted to be a mannitol dehydrogenase [Williamson, J. D., Stoop, J. M. H., Massel, M. O., Conkling, M. A. & Pharr, D. M. (1995) Proc. Natl. Acad. Sci. USA 92, 7148-7152]. Studies on the substrate specificity now revealed that ELI3-2 is an aromatic alcohol: NADP+ oxidoreductase (benzyl alcohol dehydrogenase). The enzyme showed a strong preference for various aromatic aldehydes as opposed to the corresponding alcohols. Highest substrate affinities were observed for 2-methoxybenzaldehyde, 3-methoxybenzaldehyde, salicylaldehyde, and benzaldehyde, in this order, whereas mannitol dehydrogenase activity could not be detected. These and previous results support the notion that ELI3-2 has an important role in resistance-related aromatic acid-derived metabolism.

Journal Article↗

The ankyrin repeat as molecular architecture for protein recognition.

The ankyrin repeat is one of the most frequently observed amino acid motifs in protein databases. This protein-protein interaction module is involved in a diverse set of cellular functions, and consequently, defects in ankyrin repeat proteins have been found in a number of human diseases. Recent biophysical, crystallographic, and NMR studies have been used to measure the stability and define the various topological features of this motif in an effort to understand the structural basis of ankyrin repeat-mediated protein-protein interactions. Characterization of the folding and assembly pathways suggests that ankyrin repeat domains generally undergo a two-state folding transition despite their modular structure. Also, the large number of available sequences has allowed the ankyrin repeat to be used as a template for consensus-based protein design. Such projects have been successful in revealing positions responsible for structure and function in the ankyrin repeat as well as creating a potential universal scaffold for molecular recognition.

Amino Acid Sequence↗