The following protein sequences were reprinted from the protein sequence database of the Protein Identification Resource (PIR).
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The entire protein sequence database has been exhaustively matched. Definitive mutation matrices and models for scoring gaps were obtained from the matching and used to organize the sequence database as sets of evolutionarily connected components. The methods developed are general and can be used to manage sequence data generated by major genome sequencing projects. The alignments made possible by the exhaustive matching are the starting point for successful de novo prediction of the folded structures of proteins, for reconstructing sequences of ancient proteins and metabolisms in ancient organisms, and for obtaining new perspectives in structural biochemistry.
Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.
Explore the source record for details and available documents.
Analysis of cellular protein patterns by computer-aided 2-dimensional gel electrophoresis together with recent advances in protein sequence analysis have made possible the establishment of comprehensive 2-dimensional gel protein databases that may link protein and DNA information and that offer a global approach to the study of the cell. Using the integrated approach offered by 2-dimensional gel protein databases it is now possible to reveal phenotype specific protein (or proteins), to microsequence them, to search for homology with previously identified proteins, to clone the cDNAs, to assign partial protein sequence to genes for which the full DNA sequence and the chromosome location is known, and to study the regulatory properties and function of groups of proteins that are coordinately expressed in a given biological process. Human 2-dimensional gel protein databases are becoming increasingly important in view of the concerted effort to map and sequence the entire genome.
Recently, we have developed a sequence-structure database of protein information, NRL_3D, that is extracted from the Protein Data Bank (PDB) of the Brookhaven National Laboratory. NRL_3D provides a vehicle for the retrieval of the three-dimensional coordinates of protein fragments as identified by sequence properties. These data are formulated to allow access by standard sequence analysis programs such as those provided by the Protein Identification Resource (PIR). Because the PDB is updated four times per year, semimanual construction of NRL_3D in coordination with these updates becomes a time-consuming and inefficient task. Hence, we have developed a computer program (PRENRL_3D) in the "C" computer language that automatically extracts NRL_3D from the PDB. Although the program was developed in a VAX/VMS environment, care was taken to ensure its portability to other computer systems. Customized versions of the NRL_3D database can be created from the PDB entry files using various options available in PRENRL_3D, such as selection of entries determined at high resolution and with low R-value. The program has been developed modularly and it contains a number of generalized procedures for manipulating various information in the PDB.
For the identification of newly sequenced proteins it is necessary to have a large stock of known proteins for comparison. In this paper we present an automatically generated protein sequence database. The translation program introduced allows a periodical translation of every new release of the EMBL database. Possible errors of the translation are discussed as well as the reliability of the nucleotide sequence data, which turns out to be quite good. A comparison of our translated database with some established ones is given.
A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.
A relational database of protein structure has been developed to enable rapid and flexible enquiries about the occurrence of many aspects of protein architecture. The coordinates of 294 proteins from the Brookhaven Data Bank have been processed by standard computer programs to generate many additional terms that quantify aspects of protein structure. These terms include solvent accessibility, main-chain and side-chain dihedral angles, and secondary structure. In a relational database, the information is stored in tables with columns holding the different terms and rows holding the different entries for the terms. The different relational base tables store the information about the protein coordinate set, the different chains in the protein, the amino acid residues and ligands, the atomic coordinates, the salt bridges, the hydrogen bonds, the disulphide bridges and the close tertiary contacts. The database was established under ORACLE management system. Enquiries are constructed in ORACLE using SQL (structured query language) which is simple to use and alleviates the need for extensive computer programs. A single table can be searched for entries that meet various criteria, e.g. all protein solved to better than a given resolution. The power of the database occurs when several tables, or the entries in a single table, are cross-correlated. For example the dihedral angles of proline in the fourth position in an alpha-helix in high resolution structures can be rapidly obtained. The structural database provides a powerful tool to obtain empirical rules about protein conformation. This database of protein structures is part of a joint project between Birkbeck College and Leeds University to establish an integrated data resource of protein sequences and structures (ISIS) that encodes the complex patterns of residues and coordinates that define protein conformation. The entire data resource (ISIS) will provide a system to guide all areas of protein modelling including structure prediction, site-directed mutagenesis and de novo protein design. The availability of ISIS is described in the paper.