PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Two-dimensional database of mouse liver proteins. An update.

We updated the two-dimensional protein database for mouse liver. Microsomal and cytosolic fractions of the liver proteins from male mice were separated by two-dimensional electrophoresis. The proteins were identified by Matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) on the basis of peptide mass fingerprinting, following in-gel digestion with trypsin and matching with the theoretical peptide masses of all known proteins from all species. Approximately 5800 spots, excised from 14 two-dimensional gels, were analyzed which resulted in the identification of about 2500 proteins that were the products of 328 different genes. The database includes 112 newly identified gene products. The fractionation prior to two-dimensional electrophoresis was essential for the detection of the new proteins, 55% of which were found in the microsomal and 35% in the cytosolic fraction. The more frequently identified proteins in the various gels were heat shock proteins, house-keeping enzymes, such as ATP synthase chains, disulfide isomerase, and structural proteins, such as tropomyosin. About 45% of the identified proteins were detected 1-3 times, 45% 4-9 times, and the rest 10 or more times. Most proteins were represented by many spots. In average, about 18-20 spots were detected per gene product.

Animals↗

The PIR-International Protein Sequence Database.

From its origin the Protein Information Resource (http://www-nbrf. georgetown.edu/pir/) has supported research on evolution and computational biology by designing and compiling a comprehensive, quality controlled, and well-organized protein sequence database. The database has been produced and updated on a regular schedule since 1984. Since 1988 it has been maintained collaboratively by the PIR-International, an association of data collection centers engaged in international cooperation for the development of this research resource during a period of explosive acquisition of new data. As of June 1997, essentially all sequence entries have been classified into families, allowing the efficient application of methods to propagate and standardize annotation among related sequences. The databases are available through the Internet by the World-Wide Web and FTP, or on CD-ROM and magnetic media.

Amino Acid Sequence↗

The PIR-International Protein Sequence Database.

From its origin the Protein Sequence Database has been designed to support research and has focused on comprehensive coverage, quality control and organization of the data in accordance with biological principles. Since 1988 the database has been maintained collaboratively within the framework of PIR-International, an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. The database is widely distributed and is available on the World Wide Web, via ftp, email server, on CD-ROM and magnetic media. It is widely redistributed and incorporated into many other protein sequence data compilations, including SWISS-PROT and the Entrez system of the NCBI.

Amino Acid Sequence↗

OWL--a non-redundant composite protein sequence database.

A comprehensive, non-redundant composite protein sequence database is described. The database, OWL, is an amalgam of data from six publicly-available primary sources, and is generated using strict redundancy criteria. The database is updated monthly and its size has increased almost eight-fold in the last six years: the current version contains > 76,000 entries. For added flexibility, OWL is distributed with a tailor-made query language, together with a number of programs for database exploration, information retrieval and sequence analysis, which together form an integrated database and software resource for protein sequences.

Amino Acid Sequence↗

Human cellular protein patterns and their link to genome DNA sequence data: usefulness of two-dimensional gel electrophoresis and microsequencing.

Analysis of cellular protein patterns by computer-aided 2-dimensional gel electrophoresis together with recent advances in protein sequence analysis have made possible the establishment of comprehensive 2-dimensional gel protein databases that may link protein and DNA information and that offer a global approach to the study of the cell. Using the integrated approach offered by 2-dimensional gel protein databases it is now possible to reveal phenotype specific protein (or proteins), to microsequence them, to search for homology with previously identified proteins, to clone the cDNAs, to assign partial protein sequence to genes for which the full DNA sequence and the chromosome location is known, and to study the regulatory properties and function of groups of proteins that are coordinately expressed in a given biological process. Human 2-dimensional gel protein databases are becoming increasingly important in view of the concerted effort to map and sequence the entire genome.

Amino Acid Sequence↗

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence↗

Two-dimensional database of mouse liver proteins: changes in hepatic protein levels following treatment with acetaminophen or its nontoxic regioisomer 3-acetamidophenol.

Overdose of acetaminophen (APAP) causes acute hepatotoxicity in rodents and man. The mechanism underlying APAP-induced liver injury remains unclear, but experimental evidence strongly suggests that activation of APAP and subsequent formation of protein adducts are involved in hepatotoxicity. Using proteomics technologies, we constructed a two-dimensional protein database for mouse liver, comprising 256 different gene products and investigated the proteins affected after APAP-induced hepatotoxicity. Adult male mice received a single dose of APAP (100 or 300 mg/kg) or its nontoxic regioisomer 3-acetamidophenol (AMAP, 300 mg/kg). The extent of liver damage was assessed 8 h after administration by increased liver enzyme release and histopathology. Changes in the protein level were studied by comparison of the intensities of the corresponding spots on two-dimensional (2-D) gels. The expression level of about 35 of the identified proteins was modified due to treatment with APAP or AMAP. The observed changes were usually in the order of 10-50% of the control value and were more marked in the high- than in the low-dose of APAP-treated animals. Most of the changes caused by AMAP occurred in a subset of the proteins modified by APAP. Many of the proteins showing changed expression levels are either known targets for covalent modification by N-acetyl-p-benzoquinoneimine (NAPQI) or involved in the regulation of mechanisms that are believed to drive APAP-induced hepatotoxicity.

Acetaminophen↗

PRENRL_3D: a computer program for an automatic creation of NRL_3D, protein sequence-structure database, from the Protein Data Bank.

Recently, we have developed a sequence-structure database of protein information, NRL_3D, that is extracted from the Protein Data Bank (PDB) of the Brookhaven National Laboratory. NRL_3D provides a vehicle for the retrieval of the three-dimensional coordinates of protein fragments as identified by sequence properties. These data are formulated to allow access by standard sequence analysis programs such as those provided by the Protein Identification Resource (PIR). Because the PDB is updated four times per year, semimanual construction of NRL_3D in coordination with these updates becomes a time-consuming and inefficient task. Hence, we have developed a computer program (PRENRL_3D) in the "C" computer language that automatically extracts NRL_3D from the PDB. Although the program was developed in a VAX/VMS environment, care was taken to ensure its portability to other computer systems. Customized versions of the NRL_3D database can be created from the PDB entry files using various options available in PRENRL_3D, such as selection of entries determined at high resolution and with low R-value. The program has been developed modularly and it contains a number of generalized procedures for manipulating various information in the PDB.

Amino Acid Sequence↗

The design of linear peptides that fold as monomeric beta-sheet structures.

Current knowledge about the determinants of beta-sheet formation has been notably improved by the structural and kinetic analysis of model peptides, by mutagenesis experiments in proteins and by the statistical analysis of the protein structure database (Protein Data Bank; PDB). In the past year, several peptides comprising natural and non-natural amino acids have been designed to fold as monomeric three-stranded beta-sheets. In all these cases, the design strategy has involved both the statistical analysis of the protein structure database and empirical information obtained in model beta-hairpin systems and in proteins. Only in one case was rotamer analysis performed to check for the compatibility of the sidechain packing. It is foreseeable that, in future designs, algorithms exploring the sequence and conformational space will be employed. For the design of small proteins (less than 30 amino acids), questions remain about the demonstration of two-state behavior, the formation of a well-defined network of mainchain hydrogen bonds and the quantification of the structured populations.

Crystallography, X-Ray↗

Databases in protein crystallography.

Applications of structural databases in the protein crystallographic structure determination process are reviewed, using mostly examples from work carried out by the authors. Four application areas are discussed: model building, model refinement, model validation and model analysis.

Crystallography↗

Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases.

A rapid method for the identification of known proteins separated by two-dimensional gel electrophoresis is described in which molecular masses of peptide fragments are used to search a protein sequence database. The peptides are generated by in situ reduction, alkylation, and tryptic digestion of proteins electroblotted from two-dimensional gels. Masses are determined at the subpicomole level by matrix-assisted laser desorption/ionization mass spectrometry of the unfractionated digest. A computer program has been developed that searches the protein sequence database for multiple peptides of individual proteins that match the measured masses. To ensure that the most recent database updates are included, a theoretical digest of the entire database is generated each time the program is executed. This method facilitates simultaneous processing of a large number of two-dimensional gel spots. The method was applied to a two-dimensional gel of a crude Escherichia coli extract that was electroblotted onto poly(vinylidene difluoride) membrane. Ten randomly chosen spots were analyzed. With as few as three peptide masses, each protein was uniquely identified from over 91,000 protein sequences. All identifications were verified by concurrent N-terminal sequencing of identical spots from a second blot. One of the spots contained an N-terminally blocked protein that required enzymatic cleavage, peptide separation, and Edman degradation for confirmation of its identity.

Amino Acid Sequence↗

GTOP: a database of protein structures predicted from genome sequences.

Large-scale genome projects generate an unprecedented number of protein sequences, most of them are experimentally uncharacterized. Predicting the 3D structures of sequences provides important clues as to their functions. We constructed the Genomes TO Protein structures and functions (GTOP) database, containing protein fold predictions of a huge number of sequences. Predictions are mainly carried out with the homology search program PSI-BLAST, currently the most popular among high-sensitivity profile search methods. GTOP also includes the results of other analyses, e.g. homology and motif search, detection of transmembrane helices and repetitive sequences. We have completed analyzing the sequences of 41 organisms, with the number of proteins exceeding 120 000 in total. GTOP uses a graphical viewer to present the analytical results of each ORF in one page in a 'color-bar' format. The assigned 3D structures are presented by Chime plug-in or RasMol. The binding sites of ligands are also included, providing functional information. The GTOP server is available at http://spock.genes.nig.ac.jp/~genome/gtop.html.

Amino Acid Motifs↗

Proclass protein family database: new version with motif alignments.

ProClass is a protein family database which organizes non-redundant sequence entries into families defined collectively by the ProSite patterns and PIR superfamilies. The database consists of about 100,000 entries, more than half of which are classified in about 3,000 families. The new version includes links to various protein family/domain and structural class databases and contains gapped motif alignments for all ProSite patterns. The motif sequences are retrieved from both SwissProt and PIR-international databases, including numerous new members detected by our GeneFIND family identification system. The motif collection represents a 50% increase from those catalogued in ProSite. The ProClass database can be used to maximize family information retrieval, help organize protein sequence databases, and support full-scale genomic annotation. The database and its query program are freely available for on-line record retrieval and direct file transfer from our WWW server at http:/(/)diana.uthct.edu/proclass.html+ ++.

Amino Acid Sequence↗

A Saccharomyces cerevisiae Internet protein resource now available.

The QUEST Protein Database Center is now making available two Saccharomyces cerevisiae protein databases via the Internet. The yeast electrophoretic protein database (YEPD) is a database of approximately one hundred protein identifications on two-dimensional gels. The yeast protein database (YPD) is a database of gene names and properties of over 3500 yeast proteins of known sequence. These databases can be accessed via a World-Wide Web (WWW) server (URL http:@siva.cshl.org). YPD is available via public ftp (isis.cshl.org) as well, in a spreadsheet format, and in ASCII format. When accessed via WWW, both of these databases have hypertext links to other biological data, such as the SWISS-PROT protein sequence database and the Saccharomyces Genome Database (SacchDB), and to each other.

Computer Communication Networks↗

Nucleic acid and protein sequence databases.

Nucleic acid and protein sequences contain a wealth of information of interest to molecular biologists. The advent of molecular sequence databases provides a unique opportunity for the computer analysis of all available sequences. Sequence databases serve two main functions: (i) to facilitate comparisons with newly determined sequences, and (ii) to act as a source of data for the generation and testing of hypotheses concerning molecular sequence organisation and evolution. The large amounts of sequence data now becoming available require that algorithms for database searching be fast and efficient and considerable progress is being made in this area.

Algorithms↗

The HSSP database of protein structure-sequence alignments.

HSSP (homology-derived structures of proteins) is a derived database merging structural (2-D and 3-D) and sequence information (1-D). For each protein of known 3D structure from the Protein Data Bank, the database has a file with all sequence homologues, properly aligned to the PDB protein. Homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of sequence aligned sequence families, but it is also a database of implied secondary and tertiary structures.

Amino Acid Sequence↗

The RESID database of protein structure modifications: 2000 update.

The RESID Database contains supplemental information on post-translational modifications for the standardized annotations appearing in the PIR-International Protein Sequence Database. The RESID Database includes: systematic and frequently observed alternate names, Chemical s Service registry numbers, atomic formulas and weights, enzyme activities, indicators for N-terminal, C-terminal or peptide chain cross-link modifications, keywords, literature citations with database cross-references, structural diagrams and molecular models. Since 1995 updates of the RESID Database have appeared as often as weekly, and full releases appear quarterly. The database is freely accessible through the PIR Web site http://pir.georgetown.edu/pirwww/dbinfo/resid.html and by FTP.

Databases, Factual↗

The RESID Database of Protein Modifications: 2003 developments.

The RESID Database is a comprehensive collection of annotations and structures for protein pre-, co- and post-translational modifications including amino-terminal, carboxyl-terminal and peptide chain cross-link modifications. The RESID Database includes: systematic and alternate names, atomic formulas and masses, enzyme activities generating the modifications, keywords, literature citations, Gene Ontology cross-references, Protein Information Resource (PIR) and SWISS-PROT protein sequence database feature table annotations, structure diagrams and molecular models. This database is freely accessible on the Internet through the European Bioinformatics Institute at http://srs.ebi.ac.uk/srs6bin/cgi-bin/wgetz?-page+LibInfo+-lib+RESID, through the National Cancer Institute - Frederick Advanced Biomedical Computing Center at http://www.ncifcrf.gov/RESID, or through the Protein Information Resource at http://pir.georgetown.edu/pirwww/dbinfo/resid.html.

Animals↗