PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Histone Sequence Database: new histone fold family members.

Searches of the major public protein databases with core and linker chicken and human histone sequences have resulted in the compilation of an annotated set of histone protein sequences. In addition, new database searches with two distinct motif search algorithms have identified several members of the histone fold family, including human DRAP1 and yeast CSE4. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, links to the Entrez integrated information retrieval system, structures for histone and histone fold proteins, and the ability to visualize structural data through Cn3D. The database currently contains >1000 protein sequences, which are searchable by protein type, accession number, organism name, or any other free text appearing in the definition line of the entry. All sequences and alignments in this database are available through the World Wide Web at http://www.nhgri.nih. gov/DIR/GTB/HISTONES or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Amino Acid Sequence↗

Proteome analysis of a human heptocellular carcinoma cell line, HCC-M: an update.

Recently, we reported the proteome analysis of a human hepatocellular carcinoma cell line, HCC-M (Electrophoresis 2000, 21, 1787-1813), using two-dimensional gel electrophoresis (2-DE) and matrix assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS). From a total of 408 unique spots excised from the 2-DE gel, 301 spots yielded good MALDI spectra. Out of these, 272 spots had matches returned from the database search leading to the identification of these proteins. Here, we report the results on the identification of the remaining 29 spots using nanoelectrospray ionization-tandem mass spectrometry (nESI-MS/MS). First, "peptide tag sequencing" was performed to obtain partial amino acid sequences of the peptides to search the SWISS-PROTand NCBI nonredundant protein databases. Spots that were still not able to find any matches from the databases were subjected to de novo peptide sequencing. The tryptic peptide sequences were used to search for homologues in the protein and nucleotide databases with the NCBI Basic Local Alignment Search Tool (BLAST), which was essential for the characterization of novel or post-translationally modified proteins. Using this approach, all the 29 spots were unambiguously identified. Among them, phosphotyrosyl phosphatase activator (PTPA), RNA-binding protein regulatory subunit, replication protein A 32 kDa subunit (RP-A) and N-acetylneuraminic acid phosphate synthase were reported to be cancer-related proteins.

Carcinoma, Hepatocellular↗

Identification of protein phosphorylation sites by combination of elastase digestion, immobilized metal affinity chromatography, and quadrupole-time of flight tandem mass spectrometry.

Using the combination of in-gel elastase digestion, immobilized metal affinity chromatography and high resolution electrospray tandem mass spectrometry, the phosphorylation sites of two phosphoproteins were determined. Complete coverage of all phosphorylation sites (Ser10, Ser139, Thr197, Ser338) of the model phosphoprotein protein kinase A C(alpha)-subunit could be achieved by this strategy in the low picomole range. In addition, three previously unknown phosphorylation sites of the human transcription initiation factor TIF-IA (Ser44, Ser170, Ser172) were determined in this way. Both phosphoproteins could be identified in a protein database on the basis of their elastase generated phosphopeptides alone. The data of seven phosphopeptides were used for identification of protein kinase A, and those of two phosphopeptides for TIF-IA, respectively. The accurate mass data of the electrospray mass spectra recorded at high resolution are extremely useful for sequencing of the elastase generated phosphopeptides and for protein identification by database searching.

Chromatography, Affinity↗

Thousands of proteins likely to have long disordered regions.

Neural network predictors of protein disorder using primary sequence information were developed and applied to the Swiss Protein Database. More than 15,000 proteins were predicted to contain disordered regions of at least 40 consecutive amino acids, with more than 1,000 having especially high scores indicating disorder. These results support proposals that consideration of structure-activity relationships in proteins need to be broadened to include unfolded or disordered protein.

Amino Acid Sequence↗

Selection of peptides with affinity for the N-terminal domain of GATA-1: Identification of a potential interacting protein.

As most transcription factors, GATA-1 activities are mediated by interactions with multiple proteins. Those identified so far associate with the zinc-finger domain and/or surrounding sequences. In contrast, no proteins interacting with the N-terminal domain have been identified although several evidences suggest its involvement in the control of hematopoiesis. In an attempt to identify proteins that interact with the N-terminal transactivation domain of GATA-1, a random phage peptide library was screened with recombinant GATA-1 protein and the sequence of a selected peptide was used for database protein sequence retrieval. We selected a set of peptides sharing the core sequence phi-B((2-3))-nu((2-4)) (where phi, B, and nu represent hydrophobic, basic, and neutral residues, respectively). Using the sequence of the most represented peptide (pep5) as query, we retrieved the HIV accessory protein Nef. We show that Nef binds GATA-1 and GATA-3 in vitro in virtue of its sequence homology with pep5.

Amino Acid Sequence↗

BETAWRAP: successful prediction of parallel beta -helices from primary sequence reveals an association with many microbial pathogens.

The amino acid sequence rules that specify beta-sheet structure in proteins remain obscure. A subclass of beta-sheet proteins, parallel beta-helices, represent a processive folding of the chain into an elongated topologically simpler fold than globular beta-sheets. In this paper, we present a computational approach that predicts the right-handed parallel beta-helix supersecondary structural motif in primary amino acid sequences by using beta-strand interactions learned from non-beta-helix structures. A program called BETAWRAP (http://theory.lcs.mit.edu/betawrap) implements this method and recognizes each of the seven known parallel beta-helix families, when trained on the known parallel beta-helices from outside that family. BETAWRAP identifies 2,448 sequences among 595,890 screened from the National Center for Biotechnology Information (NCBI; http://www.ncbi.nlm.nih.gov/) nonredundant protein database as likely parallel beta-helices. It identifies surprisingly many bacterial and fungal protein sequences that play a role in human infectious disease; these include toxins, virulence factors, adhesins, and surface proteins of Chlamydia, Helicobacteria, Bordetella, Leishmania, Borrelia, Rickettsia, Neisseria, and Bacillus anthracis. Also unexpected was the rarity of the parallel beta-helix fold and its predicted sequences among higher eukaryotes. The computational method introduced here can be called a three-dimensional dynamic profile method because it generates interstrand pairwise correlations from a processive sequence wrap. Such methods may be applicable to recognizing other beta structures for which strand topology and profiles of residue accessibility are well conserved.

Bacteria↗

CDART: protein homology by domain architecture.

The Conserved Domain Architecture Retrieval Tool (CDART) performs similarity searches of the NCBI Entrez Protein Database based on domain architecture, defined as the sequential order of conserved domains in proteins. The algorithm finds protein similarities across significant evolutionary distances using sensitive protein domain profiles rather than by direct sequence similarity. Proteins similar to a query protein are grouped and scored by architecture. Relying on domain profiles allows CDART to be fast, and, because it relies on annotated functional domains, informative. Domain profiles are derived from several collections of domain definitions that include functional annotation. Searches can be further refined by taxonomy and by selecting domains of interest. CDART is available at http://www.ncbi.nlm.nih.gov/Structure/lexington/lexington.cgi.

BRCA1 Protein↗

Plant protein annotation in the UniProt Knowledgebase.

The Swiss-Prot, TrEMBL, Protein Information Resource (PIR), and DNA Data Bank of Japan (DDBJ) protein database activities have united to form the Universal Protein Resource (UniProt) Consortium. UniProt presents three database layers: the UniProt Archive, the UniProt Knowledgebase (UniProtKB), and the UniProt Reference Clusters. The UniProtKB consists of two sections: UniProtKB/Swiss-Prot (fully manually curated entries) and UniProtKB/TrEMBL (automated annotation, classification and extensive cross-references). New releases are published fortnightly. A specific Plant Proteome Annotation Program (http://www.expasy.org/sprot/ppap/) was initiated to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Through UniProt, our aim is to provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information that will allow the plant community to fully explore and utilize the wealth of information available for both plant and non-plant model organisms.

Amino Acid Sequence↗

An approach to searching protein sequences for superfamily relationships or chance similarities relevant to the molecular mimicry hypothesis: application to the basic proteins of myelin.

A rapid method for similarity searches (FASTP program) was used to identify similarities between a protein database and the human basic proteins from myelin [P2 protein and 17.2K, 18.5K, and 21.5K variants of myelin basic protein (MBP)]. From similarity scores, we concluded that none of the presently known proteins are in a family containing the MBPs. No new members were found for the lipid-binding family of which P2 is a member. Sequence similarities deemed relevant to the molecular mimicry hypothesis for virus-induced autoimmunity were identified in FASTP data with the aid of microcomputer programs. Several MBP/viral protein similarities were found that have not been reported previously. Of note because of their association with demyelinating conditions were proteins from visna and vaccinia. Similarity with visna was specific to the 21.5K and 20.2K MBPs. The most interesting new possibility for mimicry involving the P2 protein was between the influenza A NS2 protein and a sequence region of P2 thought to be neuritogenic in animals and mitogenic for lymphocytes from some patients with Guillain-Barré syndrome (GBS). This may have relevance for some cases of GBS associated with the 1976 U.S.A. swine flu vaccination program. Because FASTP reports only the best similarities between proteins, searches with FASTP may not have detected all the examples of mimicry present in the database. Searches might also be more effective if similarities could be scored on immunological rather than structural relatedness.

Animals↗

Probity: a protein identification algorithm with accurate assignment of the statistical significance of the results.

An algorithm for protein identification based on mass spectrometric proteolytic peptide mapping and genome database searching is presented. The algorithm ranks database proteins based on direct calculation of the probability of random matching and assigns the statistical significance to each result. We investigate the performance of the algorithm by simulation and show that the algorithm responds to random data in the desired manner and that the statistical significance computed indicates the risk that a particular identification result is false.

Algorithms↗

PINT: Protein-protein Interactions Thermodynamic Database.

The first release of Protein-protein Interactions Thermodynamic Database (PINT) contains >1500 data of several thermodynamic parameters along with sequence and structural information, experimental conditions and literature information. Each entry contains numerical data for the free energy change, dissociation constant, association constant, enthalpy change, heat capacity change and so on of the interacting proteins upon binding, which are important for understanding the mechanism of protein-protein interactions. PINT also includes the name and source of the proteins involved in binding, their Protein Information Resource, SWISS-PROT and Protein Data Bank (PDB) codes, secondary structure and solvent accessibility of residues at mutant positions, measuring methods, experimental conditions, such as buffers, ions and additives, and literature information. A WWW interface facilitates users to search data based on various conditions, feasibility to select the terms for output and different sorting options. Further, PINT is cross-linked with other related databases, PIR, SWISS-PROT, PDB and NCBI PUBMED literature database. The database is freely available at http://www.bioinfodatabase.com/pint/index.html.

Databases, Protein↗

The specificity of UDP-GalNAc:polypeptide N-acetylgalactosaminyltransferase as inferred from a database of in vivo substrates and from the in vitro glycosylation of proteins and peptides.

The acceptor substrate specificity of UDP-GalNAc:polypeptide N-acetylgalactosaminyltransferase (GalNAc-transferase) was inferred from the amino acid sequences surrounding 196 O-glycosylation sites extracted from the National Biomedical Research Foundation Protein Database. When analyzed according to the cumulative enzyme specificity model (Poorman, R.A., Tomasselli, A.G., Heinrikson, R.L., and Kézdy, F.J. (1991) J. Biol. Chem. 266, 14554-14561) these data were found to be consistent with an enzymatic active site which interacts with an 8-amino-acid long segment of the substrate, spanning 3 amino acid residues preceding and 4 amino acid residues following the reactive serine or threonine. The model postulates independent interactions of the 8 amino acid moieties with their respective binding sites, designated as subsites P3 through P0 and P1' to P4'. High selectivity is expressed at all subsites toward serine, threonine, and proline. The inferred specificity was confirmed by in vitro bovine colostrum GalNAc-transferase-catalyzed glycosylation of unglycosylated proteins containing predicted sites for O-glycosylation and synthetic peptides designed to be GalNAc acceptors. In synthetic peptides the bovine colostrum GalNAc-transferase glycosylates threonine about 35 times faster than serine. Our results suggest that the specificity of the enzyme is not dependent on any particular secondary structure of the substrate but, rather, it is determined by the amino acids in the acceptor peptide segment as well as by the accessibility of this segment. It also appears likely that bovine colostrum GalNAc-transferase is able to catalyze in vivo the glycosylation of both threonine and serine residues.

Amino Acid Sequence↗

Profiling the malaria genome: a gene survey of three species of malaria parasite with comparison to other apicomplexan species.

We have undertaken the first comparative pilot gene discovery analysis of approximately 25,000 random genomic and expressed sequence tags (ESTs) from three species of Plasmodium, the infectious agent that causes malaria. A total of 5482 genome survey sequences (GSSs) and 5582 ESTs were generated from mung bean nuclease (MBN) and cDNA libraries, respectively, of the ANKA line of the rodent malaria parasite Plasmodium berghei, and 10,874 GSSs generated from MBN libraries of the Salvador I and Belem lines of Plasmodium vivax, the most geographically wide-spread human malaria pathogen. These tags, together with 2438 Plasmodium falciparum sequences present in GenBank, were used to perform first-pass assembly and transcript reconstruction, and non-redundant consensus sequence datasets created. The datasets were compared against public protein databases and more than 1000 putative new Plasmodium proteins identified based on sequence similarity. Homologs of previously characterized Plasmodium genes were also identified, increasing the number of P. vivax and P. berghei sequences in public databases at least 10-fold. Comparative studies with other species of Apicomplexa identified interesting homologs of possible therapeutic or diagnostic value. A gene prediction program, Phat, was used to predict probable open reading frames for proteins in all three datasets. Predicted and non-redundant BLAST-matched proteins were submitted to InterPro, an integrated database of protein domains, signatures and families, for functional classification. Thus a partial predicted proteome was created for each species. This first comparative analysis of Plasmodium protein coding sequences represents a valuable resource for further studies on the biology of this important pathogen.

Animals↗

The srhSR gene pair from Staphylococcus aureus: genomic and proteomic approaches to the identification and characterization of gene function.

Systematic analysis of the entire two-component signal transduction system (TCSTS) gene complement of Staphylococcus aureus revealed the presence of a putative TCSTS (designated SrhSR) which shares considerable homology with the ResDE His-Asp phospho-relay pair of Bacillus subtilis. Disruption of the srhSR gene pair resulted in a dramatic reduction in growth of the srhSR mutant, when cultured under anaerobic conditions, and a 3-log attenuation in growth when analyzed in the murine pyelonephritis model. To further understand the role of SrhSR, differential display two-dimensional gel electrophoresis was used to analyze the cell-free extracts derived from the srhSR mutant and the corresponding wild type. Proteins shown to be differentially regulated were identified by mass spectrometry in combination with protein database searching. An srhSR deletion led to changes in the expression of proteins involved in energy metabolism and other metabolic processes including arginine catabolism, xanthine catabolism, and cell morphology. The impaired growth of the mutant under anaerobic conditions and the dramatic changes in proteins involved in energy metabolism shed light on the mechanisms used by S. aureus to grow anaerobically and indicate that the staphylococcal SrhSR system plays an important role in the regulation of energy transduction in response to changes in oxygen availability. The combination of proteomics, bio-informatics, and microbial genetics employed here represents a powerful set of techniques which can be applied to the study of bacterial gene function.

Amino Acid Sequence↗

SPINS: standardized protein NMR storage. A data dictionary and object-oriented relational database for archiving protein NMR spectra.

Modern protein NMR spectroscopy laboratories have a rapidly growing need for an easily queried local archival system of raw experimental NMR datasets. SPINS (Standardized ProteIn Nmr Storage) is an object-oriented relational database that provides facilities for high-volume NMR data archival, organization of analyses, and dissemination of results to the public domain by automatic preparation of the header files required for submission of data to the BioMagResBank (BMRB). The current version of SPINS coordinates the process from data collection to BMRB deposition of raw NMR data by standardizing and integrating the storage and retrieval of these data in a local laboratory file system. Additional facilities include a data mining query tool, graphical database administration tools, and a NMRStar v2. 1.1 file generator. SPINS also includes a user-friendly internet-based graphical user interface, which is optionally integrated with Varian VNMR NMR data collection software. This paper provides an overview of the data model underlying the SPINS database system, a description of its implementation in Oracle, and an outline of future plans for the SPINS project.

Archives↗

GOblet: a platform for Gene Ontology annotation of anonymous sequence data.

GOblet is a comprehensive web server application providing the annotation of anonymous sequence data with Gene Ontology (GO) terms. It uses a variety of different protein databases (human, murines, invertebrates, plants, sp-trembl) and their respective GO mappings. The user selects the appropriate database and alignment threshold and thereafter submits single or multiple nucleotide or protein sequences. Results are shown in different ways, e.g. as survey statistics for the main GO categories for all sequences or as detailed results for each single sequence that has been submitted. In its newest version, GOblet allows the batch submission of sequences and provides an improved display of results with the aid of Java applets. All output data, together with the Java applet, are packed to a downloadable archive for local installation and analysis. GOblet can be accessed freely at http://goblet.molgen.mpg.de.

Animals↗

Detecting protein sequence conservation via metric embeddings.

MOTIVATION: Comparing two protein databases is a fundamental task in biosequence annotation. Given two databases, one must find all pairs of proteins that align with high score under a biologically meaningful substitution score matrix, such as a BLOSUM matrix (Henikoff and Henikoff, 1992). Distance-based approaches to this problem map each peptide in the database to a point in a metric space, such that peptides aligning with higher scores are mapped to closer points. Many techniques exist to discover close pairs of points in a metric space efficiently, but the challenge in applying this work to proteomic comparison is to find a distance mapping that accurately encodes all the distinctions among residue pairs made by a proteomic score matrix. Buhler (2002) proposed one such mapping but found that it led to a relatively inefficient algorithm for protein-protein comparison. RESULTS: This work proposes a new distance mapping for peptides under the BLOSUM matrices that permits more efficient similarity search. We first propose a new distance function on peptides derived from a given score matrix. We then show how to map peptides to bit vectors such that the distance between any two peptides is closely approximated by the Hamming distance (i.e. number of mismatches) between their corresponding bit vectors. We combine these two results with the LSH-ALL-PAIRS-SIM algorithm of Buhler (2002) to produce an improved distance-based algorithm for proteomic comparison. An initial implementation of the improved algorithm exhibits sensitivity within 5% of that of the original LSH-ALL-PAIRS-SIM, while running up to eight times faster.

Algorithms↗

Frequencies of hydrophobic and hydrophilic runs and alternations in proteins of known structure.

Patterns of alternation of hydrophobic and polar residues are a profound aspect of amino acid sequences, but a feature not easily interpreted for soluble proteins. Here we report statistics of hydrophobicity patterns in proteins of known structure in a current protein database as compared with results from earlier, more limited structure sets. Previous studies indicated that long hydrophobic runs, common in membrane proteins, are underrepresented in soluble proteins. Long runs of hydrophobic residues remain significantly underrepresented in soluble proteins, with none longer than 16 residues observed. These long runs most commonly occur as buried alpha helices, with extended hydrophobic strands less common. Avoiding aggregation of partially folded intermediates during intracellular folding remains a viable explanation for the rarity of long hydrophobic runs in soluble proteins. Comparison between database editions reveals robustness of statistics on aqueous proteins despite an approximately twofold increase in nonredundant sequences. The expanded database does now allow us to explain several deviations of hydrophobicity statistics from models of random sequence in terms of requirements of specific secondary structure elements. Comparison to prior membrane-bound protein sequences, however, shows significant qualitative changes, with the average hydrophobicity and frequency of long runs of hydrophobic residues noticeably increasing between the database editions. These results suggest that the aqueous proteins of solved structure may represent an essentially complete sample of the universe of aqueous sequences, while the membrane proteins of known structure are not yet representative of the universe of membrane-associated proteins, even by relatively simple measures of hydrophobic patterns.

Computational Biology↗