PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

MyHits: a new interactive resource for protein annotation and domain identification.

The MyHits web server (http://myhits.isb-sib.ch) is a new integrated service dedicated to the annotation of protein sequences and to the analysis of their domains and signatures. Guest users can use the system anonymously, with full access to (i) standard bioinformatics programs (e.g. PSI-BLAST, ClustalW, T-Coffee, Jalview); (ii) a large number of protein sequence databases, including standard (Swiss-Prot, TrEMBL) and locally developed databases (splice variants); (iii) databases of protein motifs (Prosite, Interpro); (iv) a precomputed list of matches ('hits') between the sequence and motif databases. All databases are updated on a weekly basis and the hit list is kept up to date incrementally. The MyHits server also includes a new collection of tools to generate graphical representations of pairwise and multiple sequence alignments including their annotated features. Free registration enables users to upload their own sequences and motifs to private databases. These are then made available through the same web interface and the same set of analytical tools. Registered users can manage their own sequences and annotations using only web tools and freeze their data in their private database for publication purposes.

Computer Graphics↗

Evolution of function in protein superfamilies, from a structural perspective.

The recent growth in protein databases has revealed the functional diversity of many protein superfamilies. We have assessed the functional variation of homologous enzyme superfamilies containing two or more enzymes, as defined by the CATH protein structure classification, by way of the Enzyme Commission (EC) scheme. Combining sequence and structure information to identify relatives, the majority of superfamilies display variation in enzyme function, with 25 % of superfamilies in the PDB having members of different enzyme types. We determined the extent of functional similarity at different levels of sequence identity for 486,000 homologous pairs (enzyme/enzyme and enzyme/non-enzyme), with structural and sequence relatives included. For single and multi-domain proteins, variation in EC number is rare above 40 % sequence identity, and above 30 %, the first three digits may be predicted with an accuracy of at least 90 %. For more distantly related proteins sharing less than 30 % sequence identity, functional variation is significant, and below this threshold, structural data are essential for understanding the molecular basis of observed functional differences. To explore the mechanisms for generating functional diversity during evolution, we have studied in detail 31 diverse structural enzyme superfamilies for which structural data are available. A large number of variations and peculiarities are observed, at the atomic level through to gross structural rearrangements. Almost all superfamilies exhibit functional diversity generated by local sequence variation and domain shuffling. Commonly, substrate specificity is diverse across a superfamily, whilst the reaction chemistry is maintained. In many superfamilies, the position of catalytic residues may vary despite playing equivalent functional roles in related proteins. The implications of functional diversity within supefamilies for the structural genomics projects are discussed. More detailed information on these superfamilies is available at http://www.biochem.ucl.ac.uk/bsm/FAM-EC/.

Binding Sites↗

An algorithm for interpretation of low-energy collision-induced dissociation product ion spectra for de novo sequencing of peptides.

An algorithm for interpretation of product ion spectra of peptides generated from ion trap mass spectrometry is developed for de novo amino acid sequencing of peptides for the purpose of protein identification. It is based on a multi-pass analysis of product ion data using a rigorous data extraction and sequence interpretation protocol in the initial pass. The extraction/interpretation algorithm becomes more relaxed in subsequent passes, considering more of the fragment ions, and potentially more sequence candidates. The possible peptide sequences generated by the algorithm are scored according to those sequences which best explain the fragment ion spectrum. These sequences are searched against a protein database using a BLAST search engine to find likely protein candidates. The method is also suitable for locating and determining protein modifications, and can be applied to de novo interpretation of peptide fragment ions in the tandem mass (MS/MS) spectrum produced from a mixture of two peptides having similar nominal mass, but different sequences. Using a known protein, bovine serum albumin, as an example, it is illustrated that this method is rapid and efficient for MS/MS spectral interpretation. This method combined with BLAST programs is then applied to search homologies and to generate information on post-translational modifications of an unknown protein isolated from shark cartilage that does not have a complete genome or proteome database.

Algorithms↗

PDB-REPRDB: a database of representative protein chains from the Protein Data Bank (PDB) in 2003.

PDB-REPRDB is a database of representative protein chains from the Protein Data Bank (PDB). Started at the Real World Computing Partnership (RWCP) in August 1997, it developed to the present system of PDB-REPRDB. In April 2001, the system was moved to the Computational Biology Research Center (CBRC), National Institute of Advanced Industrial Science and Technology (AIST) (http://www.cbrc.jp/); it is available at http://www.cbrc.jp/pdbreprdb/. The current database includes 33 368 protein chains from 16 682 PDB entries (1 September, 2002), from which are excluded (a) DNA and RNA data, (b) theoretically modeled data, (c) short chains (1<40 residues), or (d) data with non-standard amino acid residues at all residues. The number of entries including membrane protein structures in the PDB has increased rapidly with determination of numbers of membrane protein structures because of improved X-ray crystallography, NMR, and electron microscopic experimental techniques. Since many protein structure studies must address globular and membrane proteins separately, this new elimination factor, which excludes membrane protein chains, is introduced in the PDB-REPRDB system. Moreover, the PDB-REPRDB system for membrane protein chains begins at the same URL. The current membrane database includes 551 protein chains, including membrane domains in the SCOP database of release 1.59 (15 May, 2002).

Amino Acids↗

Identification and characterization of the expression of the translation initiation factor 4A (eIF4A) from Drosophila melanogaster.

We have identified the initiation factor 4A (eIF4A) in a two-dimensional protein database of Drosophila wing imaginal discs. eIF4A, a member of the DEAD-box family of RNA helicases, forms the active eIF4F complex that in the presence of eIF4B and eIF4H unwinds the secondary structure of the 5'-UTR of mRNAs during translational initiation. Two-dimensional gel electrophoresis and microsequencing allowed us to purify eIF4A, and generate specific polyclonal antibodies. A combination of immunoblotting and labelling with [(35)S]methionine + [(35)S]cysteine revealed the existence of a single eIF4A isoform encoded by a previously reported gene that maps to chromosome 2L at 26A7-9. Expression of this gene yields two mRNA species, generated by alternative splicing in the 3'-untranslated region. The two mRNAs contain the same open reading frame and produce the identical eIF4A protein. No expression was detected of the eIF4A-related gene CG7483. We detected eIF4A protein expression in the wing imaginal discs of several Drosophila species, and in haltere, leg 1, leg 2, leg 3, and eye-antenna imaginal discs of D. melanogaster. Examination of eIF4A in tumor suppressor mutants showed significantly increased (> 50%) expression in the wing imaginal discs of these larvae. We observed ubiquitous expression of eIF4A mRNA and protein during Drosophila embryogenesis. Yeast two-hybrid analysis demonstrated the in vivo interaction of Drosophila eIF4G with the N-terminal third of eIF4A.

Alternative Splicing↗

Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.

Protein identification via peptide mass fingerprinting (PMF) remains a key component of high-throughput proteomics experiments in post-genomic science. Candidate protein identifications are made using bioinformatic tools from peptide peak lists obtained via mass spectrometry (MS). These algorithms rely on several search parameters, including the number of potential uncut peptide bonds matching the primary specificity of the hydrolytic enzyme used in the experiment. Typically, up to one of these "missed cleavages" are considered by the bioinformatics search tools, usually after digestion of the in silico proteome by trypsin. Using two distinct, nonredundant datasets of peptides identified via PMF and tandem MS, a simple predictive method based on information theory is presented which is able to identify experimentally defined missed cleavages with up to 90% accuracy from amino acid sequence alone. Using this simple protocol, we are able to "mask" candidate protein databases so that confident missed cleavage sites need not be considered for in silico digestion. We show that that this leads to an improvement in database searching, with two different search engines, using the PMF dataset as a test set. In addition, the improved approach is also demonstrated on an independent PMF data set of known proteins that also has corresponding high-quality tandem MS data, validating the protein identifications. This approach has wider applicability for proteomics database searching, and the program for predicting missed cleavages and masking Fasta-formatted protein sequence databases has been made available via http:// ispider.smith.man.ac uk/MissedCleave.

Algorithms↗

The Mouse Functional Genome Database (MfunGD): functional annotation of proteins in the light of their cellular context.

MfunGD (http://mips.gsf.de/genre/proj/mfungd/) provides a resource for annotated mouse proteins and their occurrence in protein networks. Manual annotation concentrates on proteins which are found to interact physically with other proteins. Accordingly, manually curated information from a protein-protein interaction database (MPPI) and a database of mammalian protein complexes is interconnected with MfunGD. Protein function annotation is performed using the Functional Catalogue (FunCat) annotation scheme which is widely used for the analysis of protein networks. The dataset is also supplemented with information about the literature that was used in the annotation process as well as links to the SIMAP Fasta database, the Pedant protein analysis system and cross-references to external resources. Proteins that so far were not manually inspected are annotated automatically by a graphical probabilistic model and/or superparamagnetic clustering. The database is continuously expanding to include the rapidly growing amount of functional information about gene products from mouse. MfunGD is implemented in GenRE, a J2EE-based component-oriented multi-tier architecture following the separation of concern principle.

Animals↗

A conserved double-stranded RNA-binding domain.

We have identified a double-stranded (ds)RNA-binding domain in each of two proteins: the product of the Drosophila gene staufen, which is required for the localization of maternal mRNAs, and a protein of unknown function, Xlrbpa, from Xenopus. The amino acid sequences of the binding domains are similar to each other and to additional domains in each protein. Database searches identified similar domains in several other proteins known or thought to bind dsRNA, including human dsRNA-activated inhibitor (DAI), human trans-activating region (TAR)-binding protein, and Escherichia coli RNase III. By analyzing in detail one domain in staufen and one in Xlrbpa, we delimited the minimal region that binds dsRNA. On the basis of the binding studies and computer analysis, we have derived a consensus sequence that defines a 65- to 68-amino acid dsRNA-binding domain.

Amino Acid Sequence↗

Automated Gene Ontology annotation for anonymous sequence data.

Gene Ontology (GO) is the most widely accepted attempt to construct a unified and structured vocabulary for the description of genes and their products in any organism. Annotation by GO terms is performed in most of the current genome projects, which besides generality has the advantage of being very convenient for computer based classification methods. However, direct use of GO in small sequencing projects is not easy, especially for species not commonly represented in public databases. We present a software package (GOblet), which performs annotation based on GO terms for anonymous cDNA or protein sequences. It uses the species independent GO structure and vocabulary together with a series of protein databases collected from various sites, to perform a detailed GO annotation by sequence similarity searches. The sensitivity and the reference protein sets can be selected by the user. GOblet runs automatically and is available as a public service on our web server. The paper also addresses the reliability of automated GO annotations by using a reference set of more than 6000 human proteins. The GOblet server is accessible at http://goblet.molgen.mpg.de.

Animals↗

The mouse SWISS-2D PAGE database: a tool for proteomics study of diabetes and obesity.

A number of two-dimensional electrophoresis (2-DE) reference maps from mouse samples have been established and could be accessed through the internet. An up-to-date list can be found in WORLD-2D PAGE (http://www.expasy.ch/ch2d/2d- index.html), an index of 2-DE databases and services. None of them were established from mouse white and brown adipose tissues, pancreatic islets, liver nuclei and skeletal muscle. This publication describes the mouse SWISS-2D PAGE database. Proteins present in samples of mouse (C57BI/6J) liver, liver nuclei, muscle, white and brown adipose tissue and pancreatic islets are assembled and described in an accessible uniform format. SWISS-2D PAGE can be accessed through the World Wide Web (WWW) network on the ExPASy molecular biology server (http://www.expasy.ch/ ch2d/).

Animals↗

IR-MALDI-mass analysis of electroblotted proteins directly from the membrane: comparison of different membranes, application to on-membrane digestion, and protein identification by database searching.

A systematic membrane study investigating different neutral, cationic derivatized, and hydrophilic PVDF membranes for their suitability to carry out on-membrane tryptic digestions and to obtain infrared-matrix-assisted laser desorption/ionization (IR-MALDI) mass information on the proteolytic fragments directly from the membrane was performed. Clearly, the Immobilon CD membrane (Millipore) showed the most reproducible results over a protein mass range from 12 to 66 kDa. Typical protein load to SDS-PAGE was in the 1-2 micrograms range. The protein amount used for enzymatic treatment was estimated to be in the low picomole range. Now both the intact protein mass and the masses of the specific proteolytic fragments are available directly from the membrane. Protein databases can be searched via search algorithms on the Internet using the information on the intact protein mass and the masses, e.g., of its tryptic fragments. Investigations were performed to search for neutral, enzyme-compatible IR matrixes which allow the enzymatic treatment (on-membrane digestion) while the membrane is matrix-incubated. Thiourea could be tolerated during enzymatic cleavage in solution in concentrations of 15 g/L and resulted in high-quality spectra of intact protein signals and turned, therefore, out to be the most promising candidate.

Algorithms↗

Nucleotide sequence, organisation and structural analysis of the products of genes in the nirB-cysG region of the Escherichia coli K-12 chromosome.

The DNA sequence and derived amino-acid sequence of a 5618-base region in the 74-min area of the Escherichia coli chromosome has been determined in order to locate the structural gene, nirB, for the NADH-dependent nitrite reductase and a gene, cysG, required for the synthesis of the sirohaem prosthetic group. Three additional open reading frames, nirD, nirE and nirC, were found between nirB and cysG. Potential binding sites on the NirB protein for NADH and FAD, as well as conserved central core and interface domains, were deduced by comparing the derived amino-acid sequence with those of database proteins. A directly repeated sequence, which includes the motif -Cys-Xaa-Xaa-Cys-, is suggested as the binding site for either one [4Fe-4S] or two [2Fe-2S] clusters. The nirD gene potentially encodes a soluble, cytoplasmic protein of unknown function. No significant similarities were found between the derived amino-acid sequence of NirD and either NirB or any other protein in the database. If the nirE open reading frame is translated, it would encode a 33-amino-acid peptide of unknown function which includes 8 phenylalanyl residues. The product of the nirC gene is a highly hydrophobic protein with regions of amino-acid sequence similar to cytochrome oxidase polypeptide 1.

Amino Acid Sequence↗

Complete genome sequence of the alkaliphilic bacterium Bacillus halodurans and genomic sequence comparison with Bacillus subtilis.

The 4 202 353 bp genome of the alkaliphilic bacterium Bacillus halodurans C-125 contains 4066 predicted protein coding sequences (CDSs), 2141 (52.7%) of which have functional assignments, 1182 (29%) of which are conserved CDSs with unknown function and 743 (18. 3%) of which have no match to any protein database. Among the total CDSs, 8.8% match sequences of proteins found only in Bacillus subtilis and 66.7% are widely conserved in comparison with the proteins of various organisms, including B.subtilis. The B. halodurans genome contains 112 transposase genes, indicating that transposases have played an important evolutionary role in horizontal gene transfer and also in internal genetic rearrangement in the genome. Strain C-125 lacks some of the necessary genes for competence, such as comS, srfA and rapC, supporting the fact that competence has not been demonstrated experimentally in C-125. There is no paralog of tupA, encoding teichuronopeptide, which contributes to alkaliphily, in the C-125 genome and an ortholog of tupA cannot be found in the B.subtilis genome. Out of 11 sigma factors which belong to the extracytoplasmic function family, 10 are unique to B. halodurans, suggesting that they may have a role in the special mechanism of adaptation to an alkaline environment.

ATP-Binding Cassette Transporters↗

Identification of multidrug resistant protein 1 of mouse leukemia P388 cells on a PVDF membrane using 6-aminoquinolyl-carbamyl (AQC)-amino acid analysis and World Wide Web (WWW)-accessible tools.

Multidrug resistant protein 1 (MDR1) in a doxorubicin-resistant mouse leukemia cell line (P388/DOX) was identified using its amino acid composition combined with protein database searching (ExPASy and EMBL PROPSEARCH) via the World Wide Web. The proteins were separated by one-dimensional SDS-polyacrylamide gel electrophoresis, blotted onto a polyvinylidene fluoride membrane, and stained with Coomassie brilliant blue. A 160-kDa protein band was acid-hydrolyzed in the vapor phase (6 N HC1) and converted to 6-aminoquinolyl-carbamyl (AQC)-amino acids without extraction of the amino acids from the membrane. The amino acid composition of the protein was determined using the sensitive AQC-amino acid analysis method, improving our previously described method. The improved method involved using a Cosmosil 5C8-MS column instead of a Pegasil C8; replacement of the mobile phase A, constituent, 75 mM ammonium phosphate (pH 7.5), with 30 mM sodium phosphate buffer (pH 7.2); and slight modification of the separation program (9). All manipulations for protein hydrolysis and AQC derivatization were carried out in a hood using clean tools. This minimized contamination of amino acids at the low femtomolar level. A database search was carried out with bovine serum albumin as a calibration protein. MDR1 in P388/DOX was ranked first by both databases with high reliability (score 14 for ExPASy, distance 1.34 for EMBL).

ATP Binding Cassette Transporter, Subfamily B, Mem↗

A two-dimensional gel database of human plasma proteins.

An updated two-dimensional electrophoretic map of human plasma proteins is presented, together with a complete listing of the individual protein spots, their locations, size and isoelectric points relative to internal charge standards. Forty-nine polypeptide species are identified, many consisting of multiple spots differing in glycosylation or sequence (e.g., immunoglobulins). A further series of 35 as yet uncharacterized proteins is indicated.

Blood Proteins↗

Effect of chronic morphine exposure on the synaptic plasma-membrane subproteome of rats: a quantitative protein profiling study based on isotope-coded affinity tags and liquid chromatography/mass spectrometry.

The effect of chronic morphine exposure on the synaptic plasma-membrane subproteome in rats was studied by the isotope-coded affinity tag (ICAT) method coupled with capillary reversed-phase liquid chromatography/electrospray ionization mass spectrometry and tandem mass spectrometry. ICAT-labeled tryptic peptides of synaptic membrane proteins were successfully identified using tandem mass spectrometry in conjunction with protein database searching. Several important synaptic plasma-membrane proteins displayed significant regulation changes as a result of chronic morphine exposure in vivo. In particular, an integral membrane protein Na(+)/K+ ATPase (alpha-subunit) involved in regulation of the cell membrane potential by controlling sodium and potassium ion permeability was downregulated by 39 +/- 2%. This result was in excellent agreement with the reduction in electrogenic Na+, K+ pumping due to about 40% downregulation of Na(+)/K+ ATPase alpha3-isoform in myenteric S-neurons of morphine-exposed guinea-pigs measured by others via immunohistochemistry. The decrease in the abundance of non-erythroid alpha II-spectrin in the synaptic plasma-membrane fraction was also observed, which was hypothetically associated with the breakdown of the protein due to the upregulation of the proteolytic enzyme caspase-3 upon chronic morphine exposure.

Animals↗

Proteomic analysis and comparison of the biopsy and autopsy specimen of human brain temporal lobe.

The proteomic study on human temporal lobe can help us to understand the physiological function of CNS in normal as well as in pathological state. Proteomic tools are potent for the assessment of protein stability post mortem. In this pilot study, the human temporal lobe biopsy specimen with chronic pharmacoresistant temporal lobe epilepsy (TLE) and autopsy specimen in control were separated by 2-DE. Using MALDI-TOF-MS and MS/MS, 375 protein spots were identified which were the products of 267 genes. Six down-regulated and 23 up-regulated protein spots in the autopsy specimen were ascertained after the gel image analysis with the ImageMaster software. A number of proteins that include neurotransmitter metabolic and glycolytic enzymes, cytoprotective proteins and cytoskeleton were found decreased while the precursor of apolipoprotein A-I increased in the TLE brain. We tried several methods to prepare the protein samples and found that DNase and RNase treatment, ultracentrifugation and Amersham clean-up kit purification can improve gel separation quality. This work optimized the sample preparation method and constructed a primary protein database of human temporal lobe and found some proteins with remarkable level change probably involved in the post-mortem process and chronic pharmacoresistant TLE pathogenesis.

Autopsy↗

Cloning of the cDNAs for the small subunits of bovine and human DNA polymerase delta and chromosomal location of the human gene (POLD2).

cDNAs encoding the small subunit of bovine and human DNA polymerase delta have been cloned and sequenced. The predicted polypeptides, 50,885 and 51,289 Daltons, respectively, are 94% identical, similar to the catalytic subunits. The high degree of conservation of the polypeptides suggests an essential function for the small subunit in the heterodimeric core enzyme. Although the catalytic subunit of DNA polymerase delta shares significant homology with those of the herpes virus family of DNA polymerases, the small subunit of mammalian DNA polymerase delta is not homologous to the small subunit of either herpes simplex virus type 1 DNA polymerase (UL42 protein) or the Epstein-Barr virus DNA polymerase (BMRF1 protein). Searches of the protein databases failed to detect significant homology with any protein sequenced thus far. PCR analysis of DNA from a panel of human-hamster hybrid cell lines localized the gene (POLD2) for the small subunit of DNA polymerase delta to human chromosome 7.

Amino Acid Sequence↗