PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

WWW access to the SYSTERS protein sequence cluster set.

SUMMARY: We present a Web server where the SYSTERS cluster set of the non-redundant protein database consisting of sequences from SWISS-PROT and PIR is being made available for querying and browsing. The cluster set can be searched with a new sequence using the SSMAL search tool. Additionally, a multiple alignment is generated for each cluster and annotated with domain information from the Pfam protein family database. AVAILABILITY: The server address is http://www.dkfz-heidelberg.de/tbi/services/cluster/ systersform

Algorithms↗

Identification of glycosylphosphatidylinositol-anchored proteins in Arabidopsis. A proteomic and genomic analysis.

In a recent bioinformatic analysis, we predicted the presence of multiple families of cell surface glycosylphosphatidylinositol (GPI)-anchored proteins (GAPs) in Arabidopsis (G.H.H. Borner, D.J. Sherrier, T.J. Stevens, I.T. Arkin, P. Dupree [2002] Plant Physiol 129: 486-499). A number of publications have since demonstrated the importance of predicted GAPs in diverse physiological processes including root development, cell wall integrity, and adhesion. However, direct experimental evidence for their GPI anchoring is mostly lacking. Here, we present the first, to our knowledge, large-scale proteomic identification of plant GAPs. Triton X-114 phase partitioning and sensitivity to phosphatidylinositol-specific phospholipase C were used to prepare GAP-rich fractions from Arabidopsis callus cells. Two-dimensional fluorescence difference gel electrophoresis and one-dimensional sodium dodecyl sulfate-polyacrylamide gel electrophoresis demonstrated the existence of a large number of phospholipase C-sensitive Arabidopsis proteins. Using liquid chromatography-tandem mass spectrometry, 30 GAPs were identified, including six beta-1,3 glucanases, five phytocyanins, four fasciclin-like arabinogalactan proteins, four receptor-like proteins, two Hedgehog-interacting-like proteins, two putative glycerophosphodiesterases, a lipid transfer-like protein, a COBRA-like protein, SKU5, and SKS1. These results validate our previous bioinformatic analysis of the Arabidopsis protein database. Using the confirmed GAPs from the proteomic analysis to train the search algorithm, as well as improved genomic annotation, an updated in silico screen yielded 64 new candidates, raising the total to 248 predicted GAPs in Arabidopsis.

Arabidopsis↗

The limit of accuracy of protein modeling: influence of crystal packing on protein structure.

The size of the protein database (PDB) makes it now feasible to arrive at statistical conclusions regarding structural effects of crystal packing. These effects are relevant for setting upper practical limits of accuracy on protein modeling. Proteins whose crystals have more than one molecule in the asymmetric unit or whose structures were determined at least twice by X-ray crystallography were paired and their differences analyzed. We demonstrate a clear influence of crystal environment on protein structure, including backbone conformations, hinge-like motions and side-chain conformations. The positions of surface water molecules tend to be variable in different crystal environments while those of ligands are not. Structures determined by independent groups vary more than structures determined by the same authors. The use of different refinement methods is a major source for this effect. Our pair-wise analysis derives a practical limit to the accuracy of protein modeling. For different crystal forms, the limit of accuracy (C(alpha), root-mean-square deviation (RMSD)) is approximately 0.8A for the entire protein, which includes approximately 0.3A due to crystal packing. For organized secondary elements, the upper limit of C(alpha) RMSD is 0.5-0.6A while for loops or protein surface it reaches 1.0A. Twenty percent of exposed side- chains exhibit different chi(1+2) conformations with approximately half of the effect also resulting from crystal packing. A web based tool for analysis and graphic presentation of surface areas of crystal contacts is available (http://ligin.weizmann.ac.il/cryco).

Computational Biology↗

The normal human amniotic fluid supernatant proteome.

Proteomic analysis combining two-dimentional electrophoresis (2DE) and mass spectrometry (MS) has the potential for a wide range of applications in biological and medical sciences, as protein screening in tissues obtained from healthy and diseased conditions can determine drug targets and diagnostic markers. Conventionally, amniotic fluid (AF) samples are routinely used for prenatal diagnosis of a wide range of fetal abnormalities. Proteomics have already been applied in the analysis of tissues from fetuses with Down's syndrome, in order to detect differences in their protein profile as compared to the normal profiles and to determine possible diagnostic tools. A detailed protein 2DE for the normal human AF has not been reported. In the present study, the 2D protein database of the normal human AF supernatant (AFS) was constructed. Ten AFS samples from women carrying normal fetuses were analysed by 2DE. A mean of 412 spots per gel were analyzed and protein identification was carried out by MALDI-MS and MALDI-MS-MS. A 2D protein map comprising of 136 different gene products was constructed. The majority of the identified proteins are regulatory proteins, enzymes, secreted proteins, carriers and immunoglobulins. Twelve hypothetical proteins were also included. The normal AFS proteome map is a valuable tool for the study of aberrant protein expression and the search for proteins as possible markers for the prediction of abnormal fetuses.

Amniotic Fluid↗

Large scale bacterial gene discovery by similarity search.

DNA sequencing efforts frequently uncover genes other than the targeted ones. We have used rapid database scanning methods to search for undescribed eubacterial and archean protein coding frames in regions flanking known genes. By searching all prokaryotic DNA sequences not marked as coding for proteins or stable RNAs against the protein databases, we have identified more than 450 new examples of bacterial proteins, as well as a smaller number of possible revisions to known proteins, at a surprisingly high rate of one new protein or revision for every 24 initial DNA sequences or 8,300 nucleotides examined. Seven proteins are members of families which have not been described in prokaryotic sequences. We also describe 49 re-interpretations of existing sequence data of particular biological significance.

Amino Acid Sequence↗

A tool for analyzing and annotating genomic sequences.

We describe a tool for analyzing and annotating large genomic sequences containing introns. The analysis and annotation tool (AAT) includes two sets of programs, one for comparing the query sequence with a protein database and the other for comparing the query with a cDNA database. Each set contains a fast database search program and a rigorous alignment program. The database search program quickly identifies regions of the query sequence that are similar to a database sequence. Then the alignment program constructs an optimal alignment for each region and the database sequence. The alignment program also reports the coordinates of exons in the query sequence. Pairwise alignments of the query sequence with protein and cDNA database sequences are combined into multiple sequence alignments, which provide a view of all protein and cDNA sequences matching a query region. On a data set of 570 DNA sequences, AAT identified 94% of coding nucleotides correctly and 74% of exons exactly. Results of analyzing a human BAC sequence with the AAT tool are also presented. The AAT tool reduces the labor-intensive work of locating the exons of the query sequence and improves the process of defining intron-exon boundaries by using the wealth of available protein and cDNA data.

Amino Acid Sequence↗

The EROP-Moscow oligopeptide database.

Natural oligopeptides may regulate nearly all vital processes. To date, the chemical structures of nearly 6000 oligopeptides have been identified from >1000 organisms representing all the biological kingdoms. We have compiled the known physical, chemical and biological properties of these oligopeptides--whether synthesized on ribosomes or by non-ribosomal enzymes--and have constructed an internet-accessible database, EROP-Moscow (Endogenous Regulatory OligoPeptides), which resides at http://erop.inbi.ras.ru. This database enables users to perform rapid searches via many key features of the oligopeptides, and to carry out statistical analysis of all the available information. The database lists only those oligopeptides whose chemical structures have been completely determined (directly or by translation from nucleotide sequences). It provides extensive links with the Swiss-Prot-TrEMBL peptide-protein database, as well as with the PubMed biomedical bibliographic database. EROP-Moscow also contains data on many oligopeptides that are absent from other convenient databases, and is designed for extended use in classifying new natural oligopeptides and for production of novel peptide pharmaceuticals.

Anti-Infective Agents↗

GFSWeb: a web tool for genome-based identification of proteins from mass spectrometric samples.

The interpretation of mass spectrometry data for protein identification has become a vital component of proteomics research. However, since most existing software tools rely on protein databases, their success is limited, especially as the pace of annotation efforts fails to keep pace with sequencing. We present a publicly available, web-based version of a software tool that maps peptide mass fingerprint data directly to their genomic origin, allowing for genome-based, annotation-independent protein identification.

Databases, Protein↗

A rigorous method for multigenic families' functional annotation: the peptidyl arginine deiminase (PADs) proteins family example.

BACKGROUND: large scale and reliable proteins' functional annotation is a major challenge in modern biology. Phylogenetic analyses have been shown to be important for such tasks. However, up to now, phylogenetic annotation did not take into account expression data (i.e. ESTs, Microarrays, SAGE, ...). Therefore, integrating such data, like ESTs in phylogenetic annotation could be a major advance in post genomic analyses. We developed an approach enabling the combination of expression data and phylogenetic analysis. To illustrate our method, we used an example protein family, the peptidyl arginine deiminases (PADs), probably implied in Rheumatoid Arthritis. RESULTS: the analysis was performed as follows: we built a phylogeny of PAD proteins from the NCBI's NR protein database. We completed the phylogenetic reconstruction of PADs using an enlarged sequence database containing translations of ESTs contigs. We then extracted all corresponding expression data contained in EST database This analysis allowed us 1/To extend the spectrum of homologs-containing species and to improve the reconstruction of genes' evolutionary history. 2/To deduce an accurate gene expression pattern for each member of this protein family. 3/To show a correlation between paralogous sequences' evolution rate and pattern of tissular expression. CONCLUSION: coupling phylogenetic reconstruction and expression data is a promising way of analysis that could be applied to all multigenic families to investigate the relationship between molecular and transcriptional evolution and to improve functional annotation.

Animals↗

ProtoNet: hierarchical classification of the protein space.

The ProtoNet site provides an automatic hierarchical clustering of the SWISS-PROT protein database. The clustering is based on an all-against-all BLAST similarity search. The similarities' E-score is used to perform a continuous bottom-up clustering process by applying alternative rules for merging clusters. The outcome of this clustering process is a classification of the input proteins into a hierarchy of clusters of varying degrees of granularity. ProtoNet (version 1.3) is accessible in the form of an interactive web site at http://www.protonet.cs.huji.ac.il. ProtoNet provides navigation tools for monitoring the clustering process with a vertical and horizontal view. Each cluster at any level of the hierarchy is assigned with a statistical index, indicating the level of purity based on biological keywords such as those provided by SWISS-PROT and InterPro. ProtoNet can be used for function prediction, for defining superfamilies and subfamilies and for large-scale protein annotation purposes.

Animals↗

Classifying a protein in the CATH database of domain structures.

The CATH database of protein domain structures classifies structures according to their (C)lass, (A)rchitecture, (T)opology or fold and (H)omologous family (http://www.biochem.ucl.ac.uk/bsm/cath). Although the protocol used is mostly automatic, manual inspection is used to check assignments at some critical stages, such as the detection of very distantly related homologues and anologues and the assignment of novel architectures. Described in this article is a recently established facility to search the database with the coordinates of a newly determined structure. The CATH server first locates domain boundaries and then uses automatic sequence and structure comparison methods to assign this new structure to one or more of the domain families within CATH. Diagnostic reports are generated, together with multiple structural alignments for close relatives. The Server can be accessed over the World Wide Web (WWW) and mirror sites are planned to improve access.

Amino Acid Sequence↗

Rapid discovery of putative protein biomarkers of traumatic brain injury by SDS-PAGE-capillary liquid chromatography-tandem mass spectrometry.

We report the rapid discovery of putative protein biomarkers of traumatic brain injury (TBI) by SDS-PAGE-capillary liquid chromatography-tandem mass spectrometry (SDS-PAGE-Capillary LC-MS(2)). Ipsilateral hippocampus (IH) samples were collected from naive rats and rats subjected to controlled cortical impact (a rodent model of TBI). Protein database searching with 15,558 uninterpreted MS(2) spectra, collected in 3 days via data-dependent capillary LC-MS(2) of pooled cyanine dye-labeled samples separated by SDS-PAGE, identified more than 306 unique proteins. Differential proteomic analysis revealed differences in protein sequence coverage for 170 mammalian proteins (57 in naive only, 74 in injured only, and 39 of 64 in both), suggesting these are putative biomarkers of TBI. Confidence in our results was obtained by the presence of several known biomarkers of TBI (including alphaII-spectrin, brain creatine kinase, and neuron-specific enolase) in our data set. These results show that SDS-PAGE prior to in vitro proteolysis and capillary LC-MS(2) is a promising strategy for the rapid discovery of putative protein biomarkers associated with a specific physiological state (i.e., TBI) without a priori knowledge of the molecules involved.

Amino Acid Sequence↗

Stochastic motif extraction using hidden Markov model.

In this paper, we study the application of an HMM (hidden Markov model) to the problem of representing protein sequences by a stochastic motif. A stochastic protein motif represents the small segments of protein sequences that have a certain function or structure. The stochastic motif, represented by an HMM, has conditional probabilities to deal with the stochastic nature of the motif. This HMM directly reflects the characteristics of the motif, such as a protein periodical structure or grouping. In order to obtain the optimal HMM, we developed the "ilerative duplication method" for HMM topology learning. It starts from a small fully-connected network and iterates the network generation and parameter optimization until it achieves sufficient discrimination accuracy. Using this method, we obtained an HMM for a leucine zipper motif. Compared to the accuracy of a symbolic pattern representation with accuracy of 14.8 percent, an HMM achieved 79.3 percent in prediction. Additionally, the method can obtain an HMM for various types of zinc finger motifs, and it might separate the mixed data. We demonstrated that this approach is applicable to the validation of the protein database; a constructed HMM has indicated that one protein sequence annotated as "leucine-zipper like sequence" in the database is quite different from other leucine-zipper sequences in terms of likelihood, and we found this discrimination is plausible.

Algorithms↗

Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach.

The function of a protein that has no sequence homolog of known function is difficult to assign on the basis of sequence similarity. The same problem may arise for homologous proteins of different functions if one is newly discovered and the other is the only known protein of similar sequence. It is desirable to explore methods that are not based on sequence similarity. One approach is to assign functional family of a protein to provide useful hint about its function. Several groups have employed a statistical learning method, support vector machines (SVMs), for predicting protein functional family directly from sequence irrespective of sequence similarity. These studies showed that SVM prediction accuracy is at a level useful for functional family assignment. But its capability for assignment of distantly related proteins and homologous proteins of different functions has not been critically and adequately assessed. Here SVM is tested for functional family assignment of two groups of enzymes. One consists of 50 enzymes that have no homolog of known function from PSI-BLAST search of protein databases. The other contains eight pairs of homologous enzymes of different families. SVM correctly assigns 72% of the enzymes in the first group and 62% of the enzyme pairs in the second group, suggesting that it is potentially useful for facilitating functional study of novel proteins. A web version of our software, SVMProt, is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.

Artificial Intelligence↗

Schematic representation of residue-based protein context-dependent data: an application to transmembrane proteins.

An algorithmic method for drawing residue-based schematic diagrams of proteins on a 2D page is presented and illustrated. The method allows the creation of rendering engines dedicated to a given family of sequences, or fold. The initial implementation provides an engine that can produce a 2D diagram representing secondary structure for any transmembrane protein sequence. We present the details of the strategy for automating the drawing of these diagrams. The most important part of this strategy is the development of an algorithm for laying out residues of a loop that connects to arbitrary points of a 2D plane. As implemented, this algorithm is suitable for real-time modification of the loop layout. This work is of interest for the representation and analysis of data from (1) protein databases, (2) mutagenesis results, or (3) various kinds of protein context-dependent annotations or data.

Algorithms↗

NCI: A server to identify non-canonical interactions in protein structures.

NCI is a server for the identification of non-canonical interactions in protein structures. These interactions, which include N-H...pi, C(alpha)-H...pi, C(alpha)-H...O=C and variants of them, were first observed in small molecules and subsequently in high-resolution protein structures. Such interactions have been subjected to extensive structural analysis to elucidate the different geometric criteria required to identify them. These interactions have also recently been shown to be important for the stability of protein structures. In this work, I describe a server called NCI, which allows the user to either upload protein/peptide coordinates in Protein Data Bank (PDB) format or enter a Structural Classification of Proteins database (SCOP)/PDB identifier for which NCI identifies the different non-canonical interactions, based purely on geometric criteria. Results are presented as an HTML table, as a parseable text file and as a color-coded interaction matrix. In addition, the user can view the RasMol image highlighting the interactions in the protein structure and download the RasMol script. The NCI server is available at: http://www.mrc-lmb.cam.ac.uk/genomes/nci/.

Amino Acids↗

The cis-Pro touch-turn: a rare motif preferred at functional sites.

A new motif of three-dimensional (3D) protein structure is described, called the cis-Pro touch-turn. In this four-residue, three-peptide motif, the central peptide is cis. Residue 2, which precedes the proline, has phi, psi values either in the "prePro region" of the Ramachandran plot near -130 degrees, 75 degrees or in the Lalpha region near +60 degrees, +60 degrees. The Calpha(1)-Calpha(4) distance is 4-5 A and the two flanking peptides lie parallel to one another, making van der Waals contact rather than a hydrogen bond. Apparently, this arrangement is locally unfavorable and therefore rare, usually occurring only if needed for biological function. Of the 12 examples in a 500-protein database, cis-Pro touch-turns are found at the catalytic sites of pectate lyase, Ni-Fe hydrogenase, glucoamylase, xylanase, and opine dehydrogenase and at the primary binding sites of ribonuclease H, type I DNA polymerase, ribotoxin, and phage gene 3 protein. In each of these protein families, the touch-turns serve different roles; their functional importance is supported by conservation and mutagenesis data. In analyzing the conservation patterns of these 3D motifs, new methods for in-depth quality evaluation of the structural bioinformatic data are employed to distinguish between significant exceptions and errors

Amino Acid Motifs↗