PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Human cellular protein patterns and their link to genome DNA sequence data: usefulness of two-dimensional gel electrophoresis and microsequencing.

Analysis of cellular protein patterns by computer-aided 2-dimensional gel electrophoresis together with recent advances in protein sequence analysis have made possible the establishment of comprehensive 2-dimensional gel protein databases that may link protein and DNA information and that offer a global approach to the study of the cell. Using the integrated approach offered by 2-dimensional gel protein databases it is now possible to reveal phenotype specific protein (or proteins), to microsequence them, to search for homology with previously identified proteins, to clone the cDNAs, to assign partial protein sequence to genes for which the full DNA sequence and the chromosome location is known, and to study the regulatory properties and function of groups of proteins that are coordinately expressed in a given biological process. Human 2-dimensional gel protein databases are becoming increasingly important in view of the concerted effort to map and sequence the entire genome.

Amino Acid Sequence↗

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence↗

Two-dimensional database of mouse liver proteins: changes in hepatic protein levels following treatment with acetaminophen or its nontoxic regioisomer 3-acetamidophenol.

Overdose of acetaminophen (APAP) causes acute hepatotoxicity in rodents and man. The mechanism underlying APAP-induced liver injury remains unclear, but experimental evidence strongly suggests that activation of APAP and subsequent formation of protein adducts are involved in hepatotoxicity. Using proteomics technologies, we constructed a two-dimensional protein database for mouse liver, comprising 256 different gene products and investigated the proteins affected after APAP-induced hepatotoxicity. Adult male mice received a single dose of APAP (100 or 300 mg/kg) or its nontoxic regioisomer 3-acetamidophenol (AMAP, 300 mg/kg). The extent of liver damage was assessed 8 h after administration by increased liver enzyme release and histopathology. Changes in the protein level were studied by comparison of the intensities of the corresponding spots on two-dimensional (2-D) gels. The expression level of about 35 of the identified proteins was modified due to treatment with APAP or AMAP. The observed changes were usually in the order of 10-50% of the control value and were more marked in the high- than in the low-dose of APAP-treated animals. Most of the changes caused by AMAP occurred in a subset of the proteins modified by APAP. Many of the proteins showing changed expression levels are either known targets for covalent modification by N-acetyl-p-benzoquinoneimine (NAPQI) or involved in the regulation of mechanisms that are believed to drive APAP-induced hepatotoxicity.

Acetaminophen↗

HUGE: a database for human KIAA proteins, a 2004 update integrating HUGEppi and ROUGE.

We have been developing a Human Unidentified Gene-Encoded (HUGE) protein database (http://www.kazusa.or.jp/huge) to summarize results from sequence analysis of human novel large (>4 kb) cDNAs identified in the Kazusa cDNA sequencing project. At present, HUGE contains 2031 cDNA entries (KIAA cDNAs), for each of which a gene/protein characteristic table has been prepared. Since we have been shifting our research attention from the identification and cloning of novel cDNAs to the functional analysis of the proteins encoded by these cDNAs (KIAA proteins), we have not substantially increased the number of cDNA entries in HUGE for some time. Instead, we have manually curated 451 KIAA cDNAs in order to prepare a set of genetic resources to facilitate the functional analysis of KIAA proteins. In addition, we have updated the contents of the corresponding gene/protein characteristic tables in HUGE and have constructed two subsidiary databases, HUGEppi (http://www. kazusa.or.jp/huge/ppi) and ROUGE (http://www. kazusa.or.jp/rouge), to make available the results from our study of KIAA protein function. HUGEppi shows detailed information on protein-protein interactions detected between 84 pairs of KIAA proteins by yeast two-hybrid screening. ROUGE summarizes the results of computer-assisted analyses of approximately 1000 mouse homologues of human large cDNAs that we identified.

Animals↗

Protein structure database search and evolutionary classification.

As more protein structures become available and structural genomics efforts provide structural models in a genome-wide strategy, there is a growing need for fast and accurate methods for discovering homologous proteins and evolutionary classifications of newly determined structures. We have developed 3D-BLAST, in part, to address these issues. 3D-BLAST is as fast as BLAST and calculates the statistical significance (E-value) of an alignment to indicate the reliability of the prediction. Using this method, we first identified 23 states of the structural alphabet that represent pattern profiles of the backbone fragments and then used them to represent protein structure databases as structural alphabet sequence databases (SADB). Our method enhanced BLAST as a search method, using a new structural alphabet substitution matrix (SASM) to find the longest common substructures with high-scoring structured segment pairs from an SADB database. Using personal computers with Intel Pentium4 (2.8 GHz) processors, our method searched more than 10 000 protein structures in 1.3 s and achieved a good agreement with search results from detailed structure alignment methods. [3D-BLAST is available at http://3d-blast.life.nctu.edu.tw].

Acetyltransferases↗

The RESID Database of Protein Modifications as a resource and annotation tool.

The RESID Database of Protein Modifications is a comprehensive collection of annotations and structures for protein modifications and cross-links including pre-, co-, and post-translational modifications. The database provides: systematic and alternate names, atomic formulas and masses, enzymatic activities that generate the modifications, keywords, literature citations, Gene Ontology (GO) cross-references, protein sequence database feature table annotations, structure diagrams, and molecular models. This database is freely accessible on the Internet through resources provided by the European Bioinformatics Institute (http://www.ebi.ac.uk/RESID), and by the National Cancer Institute--Frederick Advanced Biomedical Computing Center (http://www.ncifcrf.gov/RESID). Each RESID Database entry presents a chemically unique modification and shows how that modification is currently annotated in the protein sequence databases, Swiss-Prot and the Protein Information Resource (PIR). The RESID Database provides a table of corresponding equivalent feature annotations that is used in the UniProt project, an international effort to combine the resources of the Swiss-Prot, TrEMBL and PIR. As an annotation tool, the RESID Database is used in standardizing and enhancing modification descriptions in the feature tables of Swiss-Prot entries. As an Internet resource, the RESID Database assists researchers in high-throughput proteomics to search monoisotopic masses and mass differences and identify known and predicted protein modifications.

Databases, Factual↗

A comprehensive and non-redundant database of protein domain movements.

MOTIVATION: The current DynDom database of protein domain motions is a user-created database that suffers from selectivity and redundancy. The aim of the analysis presented here was to overcome both these limitations and to produce both a comprehensive and a non-redundant description of domain movements from structures stored in the current protein data bank. RESULTS: A multi-step procedure is applied that starts with grouping proteins in the structural databank into families based on sequence similarity. Multiple sequence alignment, conformational clustering and a dimensional clustering method based on the Gram-Schmidt algorithm are applied to members of each family to remove dynamic redundancy in their domain movements. Representative domain movements are described in terms of domains, hinge axes and hinge-bending residues using the DynDom program. The results show that within an average family of 11.5 members, there are on average only 1.31 different domain movements indicating a high redundancy in the movements these structures represent. This verifies earlier findings that domain movements are usually highly controlled. Despite the removal of this considerable redundancy, the process has resulted in double the number of domain movements stored in the user-created database. The data are organized in a relational database with a web-interface. AVAILABILITY: The database can be browsed and searched at http://www.cmp.uea.ac.uk/dyndom CONTACT: sjh@cmp.uea.ac.uk.

Algorithms↗

PRENRL_3D: a computer program for an automatic creation of NRL_3D, protein sequence-structure database, from the Protein Data Bank.

Recently, we have developed a sequence-structure database of protein information, NRL_3D, that is extracted from the Protein Data Bank (PDB) of the Brookhaven National Laboratory. NRL_3D provides a vehicle for the retrieval of the three-dimensional coordinates of protein fragments as identified by sequence properties. These data are formulated to allow access by standard sequence analysis programs such as those provided by the Protein Identification Resource (PIR). Because the PDB is updated four times per year, semimanual construction of NRL_3D in coordination with these updates becomes a time-consuming and inefficient task. Hence, we have developed a computer program (PRENRL_3D) in the "C" computer language that automatically extracts NRL_3D from the PDB. Although the program was developed in a VAX/VMS environment, care was taken to ensure its portability to other computer systems. Customized versions of the NRL_3D database can be created from the PDB entry files using various options available in PRENRL_3D, such as selection of entries determined at high resolution and with low R-value. The program has been developed modularly and it contains a number of generalized procedures for manipulating various information in the PDB.

Amino Acid Sequence↗

The design of linear peptides that fold as monomeric beta-sheet structures.

Current knowledge about the determinants of beta-sheet formation has been notably improved by the structural and kinetic analysis of model peptides, by mutagenesis experiments in proteins and by the statistical analysis of the protein structure database (Protein Data Bank; PDB). In the past year, several peptides comprising natural and non-natural amino acids have been designed to fold as monomeric three-stranded beta-sheets. In all these cases, the design strategy has involved both the statistical analysis of the protein structure database and empirical information obtained in model beta-hairpin systems and in proteins. Only in one case was rotamer analysis performed to check for the compatibility of the sidechain packing. It is foreseeable that, in future designs, algorithms exploring the sequence and conformational space will be employed. For the design of small proteins (less than 30 amino acids), questions remain about the demonstration of two-state behavior, the formation of a well-defined network of mainchain hydrogen bonds and the quantification of the structured populations.

Crystallography, X-Ray↗

Multigenic families and proteomics: extended protein characterization as a tool for paralog gene identification.

In classical proteomic studies, the searches in protein databases lead mostly to the identification of protein functions by homology due to the non-exhaustiveness of the protein databases. The quality of the identification depends on the studied organism, its complexity and its representation in the protein databases. Nevertheless, this basic function identification is insufficient for certain applications namely for the development of RNA-based gene-silencing strategies, commonly termed RNA interference (RNAi) in animals and post-transcriptional gene silencing (PTGS) in plants, that require an unambiguous identification of the targeted gene sequence. A PTGS strategy was considered in the study of the infection of Oryza sativa by the Rice Yellow Mottle Virus (RYMV). It is suspected that the RYMV recruits host proteins after its entry into plant cells to form a complex facilitating virus multiplication and spreading. The protein partners of this complex were identified by a classical proteomic approach, nano liquid chromatography tandem mass spectrometry. Among the identified proteins, several were retained for a PTGS strategy. Nevertheless most of the protein candidates appear to be members of multigenic families for which all paralog genes are not present in protein databases. Thus the identification of the real expressed paralog gene with classical protein database searches is impossible. Consequently, as the genome contains all genes and thus all paralog genes, a whole genome search strategy was developed to determine the specific expressed paralog gene. With this approach, the identification of peptides matching only a single gene, called discriminant peptides, allows definitive proof of the expression of this identified gene. This strategy has several requirements: (i) a genome completely sequenced and accessible; (ii) high protein sequence coverage. In the present work, through three examples, we report and validate for the first time a genome database search strategy to specifically identify paralog genes belonging to multigenic families expressed under specific conditions.

Chaperonin 60↗

Databases in protein crystallography.

Applications of structural databases in the protein crystallographic structure determination process are reviewed, using mostly examples from work carried out by the authors. Four application areas are discussed: model building, model refinement, model validation and model analysis.

Crystallography↗

A database of unique protein sequence identifiers for proteome studies.

In proteome studies, identification of proteins requires searching protein sequence databases. The public protein sequence databases (e.g., NCBInr, UniProt) each contain millions of entries, and private databases add thousands more. Although much of the sequence information in these databases is redundant, each database uses distinct identifiers for the identical protein sequence and often contains unique annotation information. Users of one database obtain a database-specific sequence identifier that is often difficult to reconcile with the identifiers from a different database. When multiple databases are used for searches or the databases being searched are updated frequently, interpreting the protein identifications and associated annotations can be problematic. We have developed a database of unique protein sequence identifiers called Sequence Globally Unique Identifiers (SEGUID) derived from primary protein sequences. These identifiers serve as a common link between multiple sequence databases and are resilient to annotation changes in either public or private databases throughout the lifetime of a given protein sequence. The SEGUID Database can be downloaded (http://bioinformatics.anl.gov/SEGUID/) or easily generated at any site with access to primary protein sequence databases. Since SEGUIDs are stable, predictions based on the primary sequence information (e.g., pI, Mr) can be calculated just once; we have generated approximately 500 different calculations for more than 2.5 million sequences. SEGUIDs are used to integrate MS and 2-DE data with bioinformatics information and provide the opportunity to search multiple protein sequence databases, thereby providing a higher probability of finding the most valid protein identifications.

Amino Acid Sequence↗

Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases.

A rapid method for the identification of known proteins separated by two-dimensional gel electrophoresis is described in which molecular masses of peptide fragments are used to search a protein sequence database. The peptides are generated by in situ reduction, alkylation, and tryptic digestion of proteins electroblotted from two-dimensional gels. Masses are determined at the subpicomole level by matrix-assisted laser desorption/ionization mass spectrometry of the unfractionated digest. A computer program has been developed that searches the protein sequence database for multiple peptides of individual proteins that match the measured masses. To ensure that the most recent database updates are included, a theoretical digest of the entire database is generated each time the program is executed. This method facilitates simultaneous processing of a large number of two-dimensional gel spots. The method was applied to a two-dimensional gel of a crude Escherichia coli extract that was electroblotted onto poly(vinylidene difluoride) membrane. Ten randomly chosen spots were analyzed. With as few as three peptide masses, each protein was uniquely identified from over 91,000 protein sequences. All identifications were verified by concurrent N-terminal sequencing of identical spots from a second blot. One of the spots contained an N-terminally blocked protein that required enzymatic cleavage, peptide separation, and Edman degradation for confirmation of its identity.

Amino Acid Sequence↗

GTOP: a database of protein structures predicted from genome sequences.

Large-scale genome projects generate an unprecedented number of protein sequences, most of them are experimentally uncharacterized. Predicting the 3D structures of sequences provides important clues as to their functions. We constructed the Genomes TO Protein structures and functions (GTOP) database, containing protein fold predictions of a huge number of sequences. Predictions are mainly carried out with the homology search program PSI-BLAST, currently the most popular among high-sensitivity profile search methods. GTOP also includes the results of other analyses, e.g. homology and motif search, detection of transmembrane helices and repetitive sequences. We have completed analyzing the sequences of 41 organisms, with the number of proteins exceeding 120 000 in total. GTOP uses a graphical viewer to present the analytical results of each ORF in one page in a 'color-bar' format. The assigned 3D structures are presented by Chime plug-in or RasMol. The binding sites of ligands are also included, providing functional information. The GTOP server is available at http://spock.genes.nig.ac.jp/~genome/gtop.html.

Amino Acid Motifs↗

TOPS: an enhanced database of protein structural topology.

The TOPS database holds topological descriptions of protein structures. These compact and highly abstract descriptions reduce the protein fold to a sequence of Secondary Structure Elements (SSEs) and three sets of pairwise relationships between them, hydrogen bonds relating parallel and anti- parallel beta strands, spatial adjacencies relating neighbouring SSEs, and the chiralities of selected supersecondary structures, including connections in betaalphabeta units and between parallel alpha helices. The database is used as a resource for visualizing folding topologies, fast topological pattern searching and structure comparison. Here, significant enhancements to the TOPS database are described. The topological description has been enhanced to include packing relationships between helices, which significantly improves the description of protein folds with little beta strand content. Further, the topological description has been annotated with sequence information. The query interfaces to the database have been improved and the new version can be found at http://www.tops.leeds.ac.uk/.

Animals↗

Proclass protein family database: new version with motif alignments.

ProClass is a protein family database which organizes non-redundant sequence entries into families defined collectively by the ProSite patterns and PIR superfamilies. The database consists of about 100,000 entries, more than half of which are classified in about 3,000 families. The new version includes links to various protein family/domain and structural class databases and contains gapped motif alignments for all ProSite patterns. The motif sequences are retrieved from both SwissProt and PIR-international databases, including numerous new members detected by our GeneFIND family identification system. The motif collection represents a 50% increase from those catalogued in ProSite. The ProClass database can be used to maximize family information retrieval, help organize protein sequence databases, and support full-scale genomic annotation. The database and its query program are freely available for on-line record retrieval and direct file transfer from our WWW server at http:/(/)diana.uthct.edu/proclass.html+ ++.

Amino Acid Sequence↗

The C terminus of the nuclear protein NuMA: phylogenetic distribution and structure.

The C terminus of the nuclear protein NuMA, NuMA-CT, has a well-known function in mitosis via its proximal segment, but it seems also involved in the control of differentiation. To further investigate the structure and function of NuMA, we exploited established computational techniques and tools to collate and characterize proteins with regions similar to the distal portion of NuMA-CT (NuMA-CTDP). The phylogenetic distribution of NuMA-CTDP was examined by PSI-BLAST- and TBLASTN-based analysis of genome and protein sequence databases. Proteins and open reading frames with a NuMA-CTDP-like region were found in a diverse set of vertebrate species including mammals, birds, amphibia, and early teleost fish. The potential structure of NuMA-CTDP was investigated by searching a database of protein sequences of known three-dimensional structure with a hidden Markov model (HMM) estimated using representative (human, frog, chicken, and pufferfish) sequences. The two highest scoring sequences that aligned to the HMM were the extracellular domains of beta3-integrin and Her2, suggesting that NuMA-CTDP may have a primarily beta fold structure. These data indicate that NuMA-CTDP may represent an important functional sequence conserved in vertebrates, where it may act as a receptor to coordinate cellular events.

Amino Acid Sequence↗

Immunoaffinity purification of plasma membrane with secondary antibody superparamagnetic beads for proteomic analysis.

Plasma membrane (PM) has very important roles in cell-cell interaction and signal transduction, and it has been extensively targeted for drug design. A major prerequisite for the analysis of PM proteome is the preparation of PM with high purity. Density gradient centrifugation has been commonly employed to isolate PM, but it often occurred with contamination of internal membrane. Here we describe a method for plasma membrane purification using second antibody superparamagnetic beads that combines subcellular fractionation and immunoisolation strategies. Four methods of immunoaffinity were compared, and the variation of crude plasma membrane (CPM), superparamagnetic beads, and antibodies was studied. The optimized method and the number of CPM, beads, and antibodies suitable for proteome analysis were obtained. The PM of mouse liver was enriched 3-fold in comparison with the density gradient centrifugation method, and contamination from mitochondria was reduced 2-fold. The PM protein bands were extracted and trypsin-digested, and the resulting peptides were resolved and characterized by MALDI-TOF-TOF and ESI-Q-TOF, respectively. Mascot software was used to analyze the data against IPI-mouse protein database. Nonredundant proteins (248) were identified, of which 67% are PM or PM-related proteins. No endoplasmic reticulum (ER) or nuclear proteins were identified according to the GO annotation in the optimized method. Our protocol represents a simple, economic, and reproducible tool for the proteomic characterization of liver plasma membrane.

Animals↗