PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

SMART: a web-based tool for the study of genetically mobile domains.

SMART (a Simple Modular Architecture Research Tool) allows the identification and annotation of genetically mobile domains and the analysis of domain architectures (http://SMART.embl-heidelberg.de ). More than 400 domain families found in signalling, extra-cellular and chromatin-associated proteins are detectable. These domains are extensively annotated with respect to phyletic distributions, functional class, tertiary structures and functionally important residues. Each domain found in a non-redundant protein database as well as search parameters and taxonomic information are stored in a relational database system. User interfaces to this database allow searches for proteins containing specific combinations of domains in defined taxa.

Database Management Systems↗

Refinement of NMR-determined protein structures with database derived distance constraints.

The protein structures determined by NMR (Nuclear Magnetic Resonance Spectroscopy) are not as detailed and accurate as those by X-ray crystallography and are often underdetermined due to the inadequate distance data available from NMR experiments. The uses of NMR-determined structures in such important applications as homology modeling and rational drug design have thus been severely limited. Here we show that with the increasing numbers of high quality protein structures being determined, a computational approach to enhancing the accuracy of the NMR-determined structures becomes possible by deriving additional distance constraints from the distributions of the distances in databases of known protein structures. We show through a survey on 462 NMR structures that, in fact, many inter-atomic distances in these structures deviate considerably from their database distributions and based on the refinement results on 10 selected NMR structures that these structures can actually be improved significantly when a selected set of distances are constrained within their high probability ranges in their database distributions.

Algorithms↗

UniProt: the Universal Protein knowledgebase.

To provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information, the Swiss-Prot, TrEMBL and PIR protein database activities have united to form the Universal Protein Knowledgebase (UniProt) consortium. Our mission is to provide a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and query interfaces. The central database will have two sections, corresponding to the familiar Swiss-Prot (fully manually curated entries) and TrEMBL (enriched with automated classification, annotation and extensive cross-references). For convenient sequence searches, UniProt also provides several non-redundant sequence databases. The UniProt NREF (UniRef) databases provide representative subsets of the knowledgebase suitable for efficient searching. The comprehensive UniProt Archive (UniParc) is updated daily from many public source databases. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). The scientific community is encouraged to submit data for inclusion in UniProt.

Animals↗

Role of beta-turn residues in beta-hairpin formation and stability in designed peptides.

The sequence RGITVNGKTYGR has been reported as part of a de novo design peptide system. This peptide folds as a beta-hairpin structure with three residues per strand and two residue turns. Asn6 side-chain, the residue in position L1 of the beta-turn, appeared to be solvent exposed, interacting only within the turn but not with the rest of the peptide. We have chosen this position as a good candidate to design mutations, based on the protein database statistical abundances, that should mainly affect the turn stability and possibly the pairing between strands. We have found that all NMR parameters, in particular the conformational shift analysis of CalphaH and the coupling constants, 3JHNalpha, correlate very well and show similar conformational features in all the turn mutant peptides. The population estimates are in reasonable agreement among the different methods used. It appears that the peptide with Asn in position L1 is the most structured peptide, followed by the one with Asp6. The next structured peptide is the one with Gly6. The least populated peptides were those with Ala6 and Ser6. We have found a strong correlation between the hairpin population, as determined from the conformational shift of CalphaH and the occurrence of the different residues at position L1 of beta-hairpins with type I' beta-turn, in the protein database. Our analysis demonstrates that this peptide system is sensitive enough to register small energy changes in the hairpin structure; therefore, it constitutes an appropriate model to quantify energy contributions, once the appropriate sheet/coil transition algorithm is developed. Comparison with the other studies indicate that the design of a specific hairpin structure must involve a sequence at the turn region favouring the desired turn type, and a sequence at the strands that avoids alternative interstrand side-chain pairings.

Amino Acid Sequence↗

The use of extended amino acid motifs for focussing on toxic peptides in coeliac disease.

Cereal prolamins of wheat, rye and barley are the major proteins that have been implicated in toxicity in patients with coeliac disease. The gliadins of wheat are the best characterised with the identification of toxic peptides from rye and barley not as well advanced. This study has employed extended motifs, based on the known toxic motifs are derived from the sequence of A-gliadin, to search protein databases for matches with coeliac-toxic cereals. The results obtained have provided pointers to specific regions in rye and barley prolamins, which have received little attention in in vitro and in vivo studies of toxicity in coeliac disease. The results obtained in this study indicate that the size of the extended motif is critical when searching for coeliac-toxic cereals using protein databases. Extended motifs that are common to all three coeliac-toxic cereals and found in active wheat gliadin peptides are QQPYP, PQQPY and QQQPFP.

Amino Acid Sequence↗

The identification of CD4+ T cell epitopes with dedicated synthetic peptide libraries.

For a large number of T cell-mediated immunopathologies, the disease-related antigens are not yet identified. Identification of T cell epitopes is of crucial importance for the development of immune-intervention strategies. We show that CD4+ T cell epitopes can be defined by using a new system for synthesis and screening of synthetic peptide libraries. These libraries are designed to bind to the HLA class II restriction molecule of the CD4+ T cell clone of interest. The screening is based on three selection rounds using partial release of 14-mer peptides from synthesis beads and subsequent sequencing of the remaining peptide attached to the bead. With this approach, two peptides were identified that stimulate the beta cell-reactive CD4+ T cell clone 1c10, which was isolated from a newly diagnosed insulin-dependent diabetes mellitus patient. After performing amino acid-substitution studies and protein database searches, a Haemophilus influenzae TonB-derived peptide was identified that stimulates clone 1c10. The relevance of this finding for the pathogenesis of insulin-dependent diabetes mellitus is currently under investigation. We conclude that this system is capable of determining epitopes for (autoreactive) CD4+ T cell clones with previously unknown peptide specificity. This offers the possibility to define (auto)antigens by searching protein databases and/or to induce tolerance by using the peptide sequences identified. In addition the peptides might be used as leads to develop T cell receptor antagonists or anergy-inducing compounds.

Amino Acid Sequence↗

OPM: orientations of proteins in membranes database.

SUMMARY: The Orientations of Proteins in Membranes (OPM) database provides a collection of transmembrane, monotopic and peripheral proteins from the Protein Data Bank whose spatial arrangements in the lipid bilayer have been calculated theoretically and compared with experimental data. The database allows analysis, sorting and searching of membrane proteins based on their structural classification, species, destination membrane, numbers of transmembrane segments and subunits, numbers of secondary structures and the calculated hydrophobic thickness or tilt angle with respect to the bilayer normal. All coordinate files with the calculated membrane boundaries are available for downloading. AVAILABILITY: http://opm.phar.umich.edu.

Computer Graphics↗

Enhanced expression of hepatic genes in copper-deficient rats detected by the messenger RNA differential display method.

The influence of copper (Cu) status on hepatic gene expression was examined by using the "messenger RNA differential display" technology. This method involves the distribution of mRNA in a two-dimensional array for the rapid identification and cloning of differentially expressed genes. Livers from male Sprague-Dawley rats that had been fed a Cu-deficient (CD) diet (9.4 micromol/kg) or a Cu-adequate (CA) diet (103.9 micromol/kg) for 6 wk were used to supply cytosolic RNA. Cytosolic RNA were reverse-transcribed in the presence of anchor primers and then amplified by polymerase chain reaction with anchor and arbitrary primer sets. The amplified cDNA were then resolved by denaturing polyacrylamide gel electrophoresis. Differences in mRNA expression between the CD and CA rats were identified. DNA fragments were cloned, sequenced and used as probes for Northern blot analysis to confirm that the identified genes were differentially expressed. The analysis of cDNA sequences by computer searches against DNA and protein databases revealed that one cDNA fragment, whose mRNA abundance was enhanced 1.2-fold by copper deficiency, is novel. Four other cDNA fragments were found to have substantial homology with rat ferritin mRNA; rat fetuin mRNA; rat mitochondrial 12S and 16S rRNA, phenylalanine-, valine- and leucine-tRNA genes; rat mitochondrial genes for 16S rRNA, tRNA-leucine and tRNA-valine; and their mRNA abundance was 0.6- to 0.8-fold higher in Cu-deficient rats. Five additional cDNAs detected by this method appeared to represent novel genes because they exhibited no substantial homology to recorded gene and protein sequences deposited in DNA and protein databases. These results demonstrate the usefulness of this technology in the detection of genes which were differentially expressed as a result of the deprivation of a single nutrient, dietary copper, in this research project.

Animals↗

Evaluation of storage phosphor imaging for quantitative analysis of 2-D gels using the Quest II system.

The advent of storage phosphor technology has been of considerable benefit to the imaging of gel-separated radiolabeled proteins due to the rapid and quantitative nature of the data acquisition process. Previously, times over one month were required to obtain fluorographs of the same gel to yield data of sufficient dynamic range for quantitative analysis of high-resolution two-dimensional (2-D) gels. As we are in the process of building a human 2-D gel protein database, and therefore have a high throughput of 2-D gels both to image and quantitate using the Quest II software, we undertook an evaluation of a storage phosphor imager, including an evaluation of signal fade. The results of this evaluation demonstrate the feasibility of using such a system, and we describe the procedures that allow us to use this technique for quantitative analysis of many complex 2-D gel patterns. These procedures include a useful batch printing program that allows printing of many images in a non-interactive mode. Examples will be presented of how autoradiography, using storage phosphor plates and the Quest II system, have enabled us to begin building a human 2-D gel protein database including posttranslational modification information, without the previous time constraints associated with such a project.

Autoradiography↗

The human pituitary proteome: the characterization of differentially expressed proteins in an adenoma compared to a control.

In order to clarify the basic molecular mechanisms that participate in the formation of human pituitary macroadenomas, this study, for the first time, describes the comparative proteomics between a pituitary adenoma tissue and a control tissue. A vertical, two-dimensional polyacrylamide gel electrophoresis system and PDQuest image analysis software were used to provide a high level of between-gel reproducibility and electrophoretic separation to accurately locate each differentially expressed protein. Mass spectrometry (MALDI-TOF and LC-ESI-Q-IT) and protein databases were used to characterize each differentially expressed protein. A total of 137 differential gel spots (37 increased spot volumes, 39 decreased, 19 new and 42 lost) were found when we compared an adenoma proteome to a control proteome. Seventy-one spots (20 increased, 27 decreased, 13 new, 11 lost), representing 39 differentially regulated proteins, were identified. Five differentially regulated proteins (prolactin, cellular retinoic acid-binding protein II, G-protein beta subunit 3, secretagogin and calreticulin) were also validated with results from a comparative transcriptomics study of pituitary adenomas and controls. The functional characteristics of these differentially expressed proteins provide a differential proteomic profile between a pituitary adenoma and a control.

Adenoma↗

Molecular cloning and characterization of bacteriophage P2 genes R and S involved in tail completion.

The sequences of two previously known tail genes, R and S, of the temperate bacteriophage P2 and the sequence of an additional open reading frame (orf-30) located between S and V, were determined. Amber mutations mapping within R and S, Ram3, Ram42, Ram23, Sam75, and Sam89 were sequenced and found to be within their corresponding open reading frames. We constructed overproducing plasmids for R and S and identified these proteins by SDS-PAGE of whole-cell lysates and Coomassie blue staining. The predicted molecular masses of proteins R and S were M(r) 17,400 and 17,300, respectively, although both polypeptides migrated more slowly during gel electrophoresis than would be expected from the sequence data. orf-30 occupies the strand opposite from RS and V and is preceded by several weak potential sigma 70-RNA polymerase promoters, some of which overlap with the V promoter. A construct that had the putative orf-30 promoter region upstream of the lacZ gene produced low levels of beta-galactosidase activity in vivo. Expression from the orf-30 promoter was not stimulated by the phage P4 transcriptional activator protein, delta, which acts at all the known P2 and P4 late promoters. Insertion mutagenesis showed that orf-30 was not an essential gene for P2 growth in Escherichia coli. None of the gene or protein sequences exhibited extensive homology to sequences in the nucleic acid and protein databases. However, the R protein contains a small region homologous to one in the phage T4 tail protein gp15, which is required for T4 tails to bind heads. We propose that R and S are tail completion proteins that are essential for stable head joining.

Amino Acid Sequence↗

Systematic method for the detection of potential lambda Cro-like DNA-binding regions in proteins.

We have developed and tested a systematic method for the location and statistical evaluation of potential DNA-binding regions of the lambda Cro type in protein sequences. Using this approach to examine proteins expected to contain such regions, we have been able to compile a statistically homogeneous master set of 37 lambda Cro-like DNA-binding domains. Examination of a protein database revealed other prokaryotic proteins that are similar to this lambda Cro-like group. There are also many DNA-binding proteins that are not found to be significantly similar to the lambda Cro group, consistent with previous suggestions that different types of protein sequence may be able to achieve a similar mode of binding and that there exist other modes of sequence-specific DNA-binding. A useful feature of the method is that it can be applied without a computer.

Amino Acid Sequence↗

The myxoma virus EcoRI-O fragment encodes the DNA binding core protein and the major envelope protein of extracellular poxvirus.

The nucleotide sequences of the myxoma virus gene homologs encoding the DNA binding core protein (MF17) and the major envelope protein of the extracellular poxvirus particle (MF13) have been localized to the myxoma virus 4 kB EcoRI-O fragment. The EcoRI-O fragment is located approximately 22 kb from the left end of the 163 kb DNA genome and encodes homologs of the F12L, F13L, F15L, F16L, F17R and E1L genes of the Copenhagen strain of vaccinia virus. The inferred amino acid sequences of the myxoma virus EcoRI-O encoded products have been compared to the protein databases to identify related proteins. The myxoma virus open reading frames MF12, MF15, MF16, MF17 and ME1 encode homologs of poxvirus specific proteins while the MF13 envelope protein also shares amino acid similarity with other poxvirus and cellular proteins.

Amino Acid Sequence↗

Comparative proteome analysis of human temporal cortex lobes by two-dimensional electrophoresis and identification of selected common proteins.

Human brain proteins were isolated from left and right temporal cortex lobes at the age of 73, 23, 84 years and separated by two-dimensional gel electrophoresis (2-DE). 2-DE was carried out with an immobilized pH gradient strip in the first dimension and by sodium dodecyl sulfate-polyacrylamide gel electrophoresis in the second dimension. Over 800 polypeptide spots were resolved with a silver-staining protocol by computerized 2-D gel analsis. Seven of the polypeptide spots were evidently distinguishable between human left and right temporal lobes. Four of the polypeptide spots were larger and three were smaller in human right temporal lobe. One of these three protein spots that have descendent expression in human right temporal lobe was identified as carbonyl reductase (NADPH) 1 by MALDI-TOF MS. Thirty-three common spots were identified by ESI-MS/MALDI-TOF MS/Edman sequencing and a protein database search. These identified proteins include some important enzymes and regulating proteins.

Amino Acid Sequence↗

Protein profiling of the medicinal leech salivary gland secretion by proteomic analytical methods.

Protein diversity of the high molecular weight fraction (molecular mass > 500 daltons) of salivary grand secretion of the medicinal leech Hirudo medicinalis has been demonstrated using methods of proteomic analysis. One-dimensional (1D) electrophoresis revealed the presence of more than 60 bands corresponding to molecular masses ranging from 11 to 483 kD. 2D-electrophoresis revealed more than 100 specific protein spots differing in molecular masses and pI values. SELDI-mass spectrometry analysis using the ProteinChip. System based on chromatography surfaces of strong anion or weak cation exchanger detected 45 individual compounds of molecular masses ranged from 1.964 to 66.5 kD. Comparison of SELDI-MS data with protein databases revealed eight known proteins from the medicinal leech. Other masses detected by proteomic analytical methods may be related to both modifications of known proteins and unknown biologically active components of leech saliva secretion.

Animals↗

Proteomics of chloroplast envelope membranes.

Proteomics is a very powerful approach to link the information contained in sequenced genomes, like Arabidopsis, to the functional knowledge provided by studies of plant cell compartments, such as chloroplast envelope membranes. This review summarizes the present state of proteomic analyses of highly purified spinach and Arabidopsis envelope membranes. Methods targeted towards the hydrophobic core of the envelope allow identifying new proteins, and especially new transport systems. Common features were identified among the known and newly identified putative envelope inner membrane transporters and were used to mine the complete Arabidopsis genome to establish a virtual plastid envelope integral protein database. Arabidopsis envelope membrane proteins were extracted using different methods, that is, chloroform/methanol extraction, alkaline or saline treatments, in order to retrieve as many proteins as possible, from the most to the less hydrophobic ones. Mass spectrometry analyses lead to the identification of more than 100 proteins. More than 50% of the identified proteins have functions known or very likely to be associated with the chloroplast envelope. These proteins are (a) involved in ion and metabolite transport, (b) components of the protein import machinery and (c) involved in chloroplast lipid metabolism. Some soluble proteins, like proteases, proteins involved in carbon metabolism or in responses to oxidative stress, were associated with envelope membranes. Almost one third of the newly identified proteins have no known function. The present stage of the work demonstrates that a combination of different proteomics approaches together with bioinformatics and the use of different biological models indeed provide a better understanding of chloroplast envelope biochemical machinery at the molecular level.

Journal Article↗

Biochemical characterization and lysosomal localization of the mannose-6-phosphate protein p76 (hypothetical protein LOC196463).

Most soluble lysosomal proteins carry Man6P (mannose 6-phosphate), a specific carbohydrate marker that enables their binding to cellular MPRs (Man6P receptors) and their subsequent targeting towards the lysosome. This characteristic was exploited to identify novel soluble lysosomal proteins by proteomic analysis of Man6P proteins purified from a human cell line. Among the proteins identified during the course of the latter study [Journet, Chapel, Kieffer, Roux and Garin (2002) Proteomics, 2, 1026-1040], some had not been previously described as lysosomal proteins. We focused on a protein detected at 76 kDa by SDS/PAGE. We named this protein 'p76' and it appeared later in the NCBI protein database as the 'hypothetical protein LOC196463'. In the present paper, we describe the identification of p76 by MS and we analyse several of its biochemical characteristics. The presence of Man6P sugars was confirmed by an MPR overlay experiment, which showed the direct and Man6P-dependent interaction between p76 and the MPR. The presence of six N-glycosylation sites was validated by progressive peptide-N-glycosidase F deglycosylation. Experiments using N- and C-termini directed anti-p76 antibodies provided insights into p76 maturation. Most importantly, we were able to demonstrate the lysosomal localization of this protein, which was initially suggested by its Man6P tags, by both immunofluorescence and sub-cellular fractionation of mouse liver homogenates.

Animals↗