PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

A novel member of the lipocalin superfamily: tammar wallaby late-lactation protein.

The finding that tammar wallaby late-lactation protein is linked to beta-lactoglobulin prompted a search of current GenPeptide and NBRF-PIR protein databases for sequence similarities to late-lactation protein. Similarities were found to von Ebner's gland protein and other members of the lipocalin superfamily of proteins. A conservative replacement of Trp with Tyr suggests that late-lactation protein may represent an unusual member of this protein superfamily.

Amino Acid Sequence↗

Isolation of cDNA clone encoding rat senescence marker protein-30 (SMP30) and its tissue distribution.

We have isolated and characterized two cDNA clones encoding senescence marker protein-30 (SMP30), the amounts of which are known to decrease androgen-independently with aging in the livers of rats. Of these cDNA clones, one consisted of 1588 bp nucleotides and the other of 1195 bp nucleotides generated by alternative polyadenylation. These two cDNA clones shared the same open reading frame, but the larger species had 393 bp nucleotides of 3' untranslated region in addition to the first polyadenylation site of smaller species. Northern hybridization analysis showed that two species of mRNA (1.7 kb and 1.4 kb) located in the liver and kidney were consistent with these short and long forms of cDNA. The open reading frame, 897 bp could encode 299 amino acids. The estimated molecular weight and pI of the deduced polypeptide were 33,387 and 5.1, respectively. Furthermore, immunohistochemical analysis confirmed that SMP30 was preferentially localized in the hepatocytes and renal proximal tubular epithelium. Genomic Southern hybridization analysis demonstrated that SMP30 was widely conserved among higher animals. A computer-assisted homology analysis of nucleic acid and protein databases revealed no remarkable homology with other known proteins. Therefore, SMP30 seems to be a novel protein. In addition, the existence of putative A-U rich mRNA degradation signals and protein degradation signals (PEST sequence) in the structure of SMP30 may suggest important regulatory function of this unique protein manifested by changes in its concentrations.

Aging↗

The gene G13 in the class III region of the human MHC encodes a potential DNA-binding protein.

G13 is a single-copy gene lying approx. 75 kb centromeric of the complement gene cluster in the class III region of the human MHC. The gene spans approx. 17 kb of DNA and has been shown to encode mRNA of approx. 2.7 kb that is present in cell lines representing lymphoid and non-lymphoid tissues, indicating that it is ubiquitously expressed. The complete nucleotide sequence of the 2.7 kb mRNA has been derived from cDNA and genomic clones. The longest open reading frame obtained for G13 codes for a 703 amino acid protein of approx. 77 kDa in molecular mass. Comparison of the putative G13 amino acid sequence with the protein databases revealed significant similarities with DNA-binding proteins of the leucine zipper class, including a human cAMP response element binding protein. G13 contains a bZIP motif, a region rich in basic amino acids adjacent to a coiled-coil leucine zipper domain, common to this class of proteins that is known to be involved in dimerization and DNA binding. Antibodies raised against a fragment encoding the C-terminal half of the putative G13 protein recognized a major polypeptide of approx. 86 kDa and a minor polypeptide of approx. 78 kDa on immunoblotting of U937 cell extracts; this has been confirmed by immunoprecipitation experiments. Even though it contained at least one potential bipartite nuclear localization signal, the G13 protein was present both in the cytoplasm and the nucleus of the fibroblast cells. Thus G13 might be a novel DNA-binding protein that is perhaps translocated to the nucleus in a regulated manner.

Amino Acid Sequence↗

The CATH Domain Structure Database and related resources Gene3D and DHS provide comprehensive domain family information for genome analysis.

The CATH database of protein domain structures (http://www.biochem.ucl.ac.uk/bsm/cath/) currently contains 43,229 domains classified into 1467 superfamilies and 5107 sequence families. Each structural family is expanded with sequence relatives from GenBank and completed genomes, using a variety of efficient sequence search protocols and reliable thresholds. This extended CATH protein family database contains 616,470 domain sequences classified into 23,876 sequence families. This results in the significant expansion of the CATH HMM model library to include models built from the CATH sequence relatives, giving a 10% increase in coverage for detecting remote homologues. An improved Dictionary of Homologous superfamilies (DHS) (http://www.biochem.ucl.ac.uk/bsm/dhs/) containing specific sequence, structural and functional information for each superfamily in CATH considerably assists manual validation of homologues. Information on sequence relatives in CATH superfamilies, GenBank and completed genomes is presented in the CATH associated DHS and Gene3D resources. Domain partnership information can be obtained from Gene3D (http://www.biochem.ucl.ac.uk/bsm/cath/Gene3D/). A new CATH server has been implemented (http://www.biochem.ucl.ac.uk/cgi-bin/cath/CathServer.pl) providing automatic classification of newly determined sequences and structures using a suite of rapid sequence and structure comparison methods. The statistical significance of matches is assessed and links are provided to the putative superfamily or fold group to which the query sequence or structure is assigned.

Databases, Nucleic Acid↗

A human genomic library enriched in transcriptionally active sequences (aDNA library).

Core histone hyperacetylation, in particular of H4, is concentrated in the promoter-upstream regions of active genes and in certain cases is locuswide. Antibodies to hyperacetylated H4 were used to immunoprecipitate dinucleosomal chromatin derived from K562 human erythroleukemic cells by micrococcal nuclease digestion. The extracted DNA was made into a genomic library and was expected to contain sequences from genes active in K562 cells (an active, 'aDNA' library). Clones (180) were randomly selected from the library; 24 of 103 tested (23%) contained highly repeated sequences, as determined by their hybridization to total genomic DNA, and were not analyzed further. An additional 10 clones (6%) were shown to contain no insert DNA. The remaining 146 were sequenced and compared with the nucleic acid databases and in all six frames to the protein databases: Sixeen clones could be assigned to known genes, the majority of which (12) were tissue specific. All but 2 of these 16 corresponded to segments 5' of the coding sequences, as expected if H4 acetylation is concentrated at promoter regions. Thirty-three clones (23%) displayed high sequence identity to cDNAs in the expressed sequence tag database (dbEST). Northern blots and reverse transcription (RT)-PCR were used to determine the proportion of clones representing sequences expressed in K562 cells: Although only 1 of 34 tested clones showed a band in Northern hybridization, RT-PCR demonstrated that at least 12 of 40 tested clones (30%) were present in the mRNA population. Because a further 8 of these 40 clones were identified as gene fragments by database sequence comparisons, it follows that about half of this subset of 40 clones is derived from genes. The aDNA library is thus very gene rich and not skewed toward the most highly expressed sequences, as in mRNA libraries. The aDNA library is also rich in promoters and could be a valuable source of such sequences, particularly those that lack CpG islands or other features that allow their specific selection.

Blotting, Northern↗

Comparative salivary gland transcriptomics of sandfly vectors of visceral leishmaniasis.

BACKGROUND: Immune responses to sandfly saliva have been shown to protect animals against Leishmania infection. Yet very little is known about the molecular characteristics of salivary proteins from different sandflies, particularly from vectors transmitting visceral leishmaniasis, the fatal form of the disease. Further knowledge of the repertoire of these salivary proteins will give us insights into the molecular evolution of these proteins and will help us select relevant antigens for the development of a vector based anti-Leishmania vaccine. RESULTS: Two salivary gland cDNA libraries from female sandflies Phlebotomus argentipes and P. perniciosus were constructed, sequenced and proteomic analysis of the salivary proteins was performed. The majority of the sequenced transcripts from the two cDNA libraries coded for secreted proteins. In this analysis we identified transcripts coding for protein families not previously described in sandflies. A comparative sandfly salivary transcriptome analysis was performed by using these two cDNA libraries and two other sandfly salivary gland cDNA libraries from P. ariasi and Lutzomyia longipalpis, also vectors of visceral leishmaniasis. Full-length secreted proteins from each sandfly library were compared using a stand-alone version of BLAST, creating formatted protein databases of each sandfly library. Related groups of proteins from each sandfly species were combined into defined families of proteins. With this comparison, we identified families of salivary proteins common among all of the sandflies studied, proteins to be genus specific and proteins that appear to be species specific. The common proteins included apyrase, yellow-related protein, antigen-5, PpSP15 and PpSP32-related protein, a 33-kDa protein, D7-related protein, a 39- and a 16.1- kDa protein and an endonuclease-like protein. Some of these families contained multiple members, including PPSP15-like, yellow proteins and D7-related proteins suggesting gene expansion in these proteins. CONCLUSION: This comprehensive analysis allows us the identification of genus- specific proteins, species-specific proteins and, more importantly, proteins common among these different sandflies. These results give us insights into the repertoire of salivary proteins that are potential candidates for a vector-based vaccine.

Amino Acid Sequence↗

[The application of human mutation databases].

Researches on genome mutation are becoming more and more important with the finish of human genome DNA draft. This review is to classify the existing human mutation databases, including mutation database, SNP(single nucleotide polymorphisms) databases, mutation databases about disease, mutation databases about proteins, mutation databases about map and mutation information about specific gene. We also give advice on how to utilize these mutation databases, and discuss problems of existing databases.

Databases, Factual↗

Capillary electrophoresis/electrospray ionization high mass accuracy time-of-flight mass spectrometry for protein identification using peptide mapping.

Capillary electrophoresis/electrospray ionization (CE/ESI) high mass accuracy time-of-flight mass spectrometry was used for the first time to characterize small proteins using peptide mapping. To identify small proteins, the intact proteins were first analyzed to obtain their average molecular weights with errors less than 1 Da. On-line capillary electrophoresis mass spectrometry of the tryptic digests of these small proteins was then performed to obtain the accurate molecular weights of the peptides with accuracies of approximately 10 ppm. Next, this information was used for the identification of the proteins using a protein database. It was found that high mass accuracy is an effective tool in reducing the list of most-likely proteins generated by the database. In addition, on-line collision-induced dissociation of the completely or partially resolved capillary electrophoresis peaks of the protein digests was used to unambiguously identify the sequences of these peptides. Each CE/ESI-MS analysis used only 5 nL of sample containing approximately 120 fmol of each peptide in protein digests. The results indicate that the combination of capillary electrophoresis and high resolution, high mass accuracy time-of-flight mass spectrometry is a viable option for the identification of small proteins using peptide mapping.

Calibration↗

SPIDER: software for protein identification from sequence tags with de novo sequencing error.

For the identification of novel proteins using MS/MS, de novo sequencing software computes one or several possible amino acid sequences (called sequence tags) for each MS/MS spectrum. Those tags are then used to match, accounting amino acid mutations, the sequences in a protein database. If the de novo sequencing gives correct tags, the homologs of the proteins can be identified by this approach and software such as MS-BLAST is available for the matching. However, de novo sequencing very often gives only partially correct tags. The most common error is that a segment of amino acids is replaced by another segment with approximately the same masses. We developed a new efficient algorithm to match sequence tags with errors to database sequences for the purpose of protein and peptide identification. A software package, SPIDER, was developed and made available on Internet for free public use. This paper describes the algorithms and features of the SPIDER software.

Algorithms↗

A protein domain interaction interface database: InterPare.

BACKGROUND: Most proteins function by interacting with other molecules. Their interaction interfaces are highly conserved throughout evolution to avoid undesirable interactions that lead to fatal disorders in cells. Rational drug discovery includes computational methods to identify the interaction sites of lead compounds to the target molecules. Identifying and classifying protein interaction interfaces on a large scale can help researchers discover drug targets more efficiently. DESCRIPTION: We introduce a large-scale protein domain interaction interface database called InterPare http://interpare.net. It contains both inter-chain (between chains) interfaces and intra-chain (within chain) interfaces. InterPare uses three methods to detect interfaces: 1) the geometric distance method for checking the distance between atoms that belong to different domains, 2) Accessible Surface Area (ASA), a method for detecting the buried region of a protein that is detached from a solvent when forming multimers or complexes, and 3) the Voronoi diagram, a computational geometry method that uses a mathematical definition of interface regions. InterPare includes visualization tools to display protein interior, surface, and interaction interfaces. It also provides statistics such as the amino acid propensities of queried protein according to its interior, surface, and interface region. The atom coordinates that belong to interface, surface, and interior regions can be downloaded from the website. CONCLUSION: InterPare is an open and public database server for protein interaction interface information. It contains the large-scale interface data for proteins whose 3D-structures are known. As of November 2004, there were 10,583 (Geometric distance), 10,431 (ASA), and 11,010 (Voronoi diagram) entries in the Protein Data Bank (PDB) containing interfaces, according to the above three methods. In the case of the geometric distance method, there are 31,620 inter-chain domain-domain interaction interfaces and 12,758 intra-chain domain-domain interfaces.

Computers, Molecular↗

HOMSTRAD: recent developments of the Homologous Protein Structure Alignment Database.

HOMSTRAD (http://www-cryst.bioc.cam.ac.uk/ homstrad/) is a collection of protein families, clustered on the basis of sequence and structural similarity. The database is unique in that the protein family sequence alignments have been specially annotated using the program, JOY, to highlight a wide range of structural features. Such data are useful for identifying key structurally conserved residues within the families. Superpositions of the structures within each family are also available and a sensitive structure-aided search engine, FUGUE, can be used to search the database for matches to a query protein sequence. Historically, HOMSTRAD families were generated using several key pieces of software, including COMPARER and MNYFIT, and held in a number of flat files and indexes. A new relational database version of HOMSTRAD, HOMSTRAD BETA (http://www-cryst.bioc.cam. ac.uk/homstradbeta/) is being developed using MySQL. This relational data structure provides more flexibility for future developments, reduces update times and makes data more easily accessible. Consequently it has been possible to add a number of new web features including a custom alignment facility. Altogether, this makes HOMSTRAD and its new BETA version, an excellent resource both for comparative modelling and for identifying distant sequence/structure similarities between proteins.

Amino Acid Sequence↗

Dbp5p/Rat8p is a yeast nuclear pore-associated DEAD-box protein essential for RNA export.

To identify Saccharomyces cerevisiae genes important for nucleocytoplasmic export of messenger RNA, we screened mutant strains to identify those in which poly(A)+ RNA accumulated in nuclei under nonpermissive conditions. We describe the identification of DBP5 as the gene defective in the strain carrying the rat8-1 allele (RAT = ribonucleic acid trafficking). Dbp5p/Rat8p, a previously uncharacterized member of the DEAD-box family of proteins, is closely related to eukaryotic initiation factor 4A(eIF4A) an RNA helicase essential for protein synthesis initiation. Analysis of protein databases suggests most eukaryotic genomes encode a DEAD-box protein that is probably a homolog of yeast Dbp5p/Rat8p. Temperature-sensitive alleles of DBP5/RAT8 were prepared. In rat8 mutant strains, cells displayed rapid, synchronous accumulation of poly(A)+ RNA in nuclei when shifted to the non-permissive temperature. Dbp5p/Rat8p is located within the cytoplasm and concentrated in the perinuclear region. Analysis of the distribution of Dbp5p/Rat8p in yeast strains where nuclear pore complexes are tightly clustered indicated that a fraction of this protein associates with nuclear pore complexes (NPCs). The strong mutant phenotype, association of the protein with NPCs and genetic interaction with factors involved in RNA export provide strong evidence that Dbp5p/Rat8p plays a direct role in RNA export.

Alleles↗

Gleaning non-trivial structural, functional and evolutionary information about proteins by iterative database searches.

Using a number of diverse protein families as test cases, we investigate the ability of the recently developed iterative sequence database search method, PSI-BLAST, to identify subtle relationships between proteins that originally have been deemed detectable only at the level of structure-structure comparison. We show that PSI-BLAST can detect many, though not all, of such relationships, but the success critically depends on the optimal choice of the query sequence used to initiate the search. Generally, there is a correlation between the diversity of the sequences detected in the first pass of database screening and the ability of a given query to detect subtle relationships in subsequent iterations. Accordingly, a thorough analysis of protein superfamilies at the sequence level is necessary in order to maximize the chances of gleaning non-trivial structural and functional inferences, as opposed to a single search, initiated, for example, with the sequence of a protein whose structure is available. This strategy is illustrated by several findings, each of which involves an unexpected structural prediction: (i) a number of previously undetected proteins with the HSP70-actin fold are identified, including a highly conserved and nearly ubiquitous family of metal-dependent proteases (typified by bacterial O-sialoglycoprotease) that represent an adaptation of this fold to a new type of enzymatic activity; (ii) we show that, contrary to the previous conclusions, ATP-dependent and NAD-dependent DNA ligases are confidently predicted to possess the same fold; (iii) the C-terminal domain of 3-phosphoglycerate dehydrogenase, which binds serine and is involved in allosteric regulation of the enzyme activity, is shown to typify a new superfamily of ligand-binding, regulatory domains found primarily in enzymes and regulators of amino acid and purine metabolism; (iv) the immunoglobulin-like DNA-binding domain previously identified in the structures of transcription factors NFkappaB and NFAT is shown to be a member of a distinct superfamily of intracellular and extracellular domains with the immunoglobulin fold; and (v) the Rag-2 subunit of the V-D-J recombinase is shown to contain a kelch-type beta-propeller domain which rules out its evolutionary relationship with bacterial transposases.

Actins↗

Strategic proteome analysis of Candida magnoliae with an unsequenced genome.

Erythritol is a noncariogenic, low calorie sweetener. It is safe for people with diabetes and obese people. Candida magnoliae is an industrially important organism because of its ability to produce erythritol as a major product. The genome of C. magnoliae has not been sequenced yet, limiting the available proteome database. Therefore, systematic approaches were employed to construct the proteome map of C. magnoliae. Proteomic analysis with systematic approaches is based on two-dimensional electrophoresis, matrix-assisted laser desorption ionization time of flight mass spectrometry (MALDI-TOF MS), tandem mass spectrometry (MS/MS) and database interrogation. First, 24 spots were analyzed using peptide mass fingerprinting along with MALDI-TOF MS with high mass accuracy. Only four spots were reliably identified as carbonyl reductase and its isoforms. The reason for low sequence coverage seemed to be that these identification strategies were based on the presence of the protein database obtained from the publicly accessible genome database and the availability of cross-species protein identification. MS/MS (MS/MS ion search and de novo sequencing) in combination with similarity searches allowed successful identification of 39 spots. Several proteins including transaldolase identified by MS/MS ion searches were further confirmed by partial sequences from the expressed sequence tag database. In this study, 51 protein spots were analyzed and then potentially identified. The identified proteins were involved in glycolysis, stress response, other essential metabolisms and cell structures.

Algorithms↗

RECOORD: a recalculated coordinate database of 500+ proteins from the PDB using restraints from the BioMagResBank.

State-of-the-art methods based on CNS and CYANA were used to recalculate the nuclear magnetic resonance (NMR) solution structures of 500+ proteins for which coordinates and NMR restraints are available from the Protein Data Bank. Curated restraints were obtained from the BioMagResBank FRED database. Although the original NMR structures were determined by various methods, they all were recalculated by CNS and CYANA and refined subsequently by restrained molecular dynamics (CNS) in a hydrated environment. We present an extensive analysis of the results, in terms of various quality indicators generated by PROCHECK and WHAT_CHECK. On average, the quality indicators for packing and Ramachandran appearance moved one standard deviation closer to the mean of the reference database. The structural quality of the recalculated structures is discussed in relation to various parameters, including number of restraints per residue, NOE completeness and positional root mean square deviation (RMSD). Correlations between pairs of these quality indicators were generally low; for example, there is a weak correlation between the number of restraints per residue and the Ramachandran appearance according to WHAT_CHECK (r = 0.31). The set of recalculated coordinates constitutes a unified database of protein structures in which potential user- and software-dependent biases have been kept as small as possible. The database can be used by the structural biology community for further development of calculation protocols, validation tools, structure-based statistical approaches and modeling. The RECOORD database of recalculated structures is publicly available from http://www.ebi.ac.uk/msd/recoord.

Databases, Protein↗

Identification of microbial mixtures by LC-selective proteotypic-peptide analysis (SPA).

This paper describes a method--using a combination of LC-MS/MS of selected bacteria-specific peptides and database search--for determining the species of bacteria present in a mixture. We identified the proteotypic peptides that were associated with specific bacteria by searching protein databases for the LC-MS/MS data. The retention time windows for specific peptide markers were used as an extra constraint so that the peptide markers of many bacterial species could be analyzed in a single LC-selective proteotypic-peptide analysis (SPA). We performed LC-MS/MS analyses on the proteolytic digest of cell extracts and monitored only the selected marker peptide ions at given elution time windows. The corresponding bacterial species could be characterized when the selected peptides that eluted at expected elution windows were identified correctly from the database. We managed to identify up to eight bacterial species simultaneously during a single LC-MS/MS analysis, as well as bacteria mixed in various abundances. Two marker ions having similar values of m/z, but obtained from two different bacterial samples, which would otherwise be selected as precursors within mass tolerance and would complicate the MS/MS data, were time-resolved using LC and then used to correctly identify their bacterial sources. The coupling of selective MS/MS monitoring with separation methods, such as LC, provides a highly selective and accurate analytical method for characterizing complex mixtures of bacterial species.

Amino Acid Sequence↗

Characterization of four outer membrane proteins that play a role in utilization of starch by Bacteroides thetaiotaomicron.

Results of earlier work had suggested that utilization of polysaccharides by Bacteroides spp. did not proceed via breakdown by extracellular polysaccharide-degrading enzymes. Rather, it appeared that the polysaccharide was first bound to a putative outer membrane receptor complex and then translocated into the periplasm, where the degradative enzymes were located. In a recent article, we reported the cloning and sequencing of susC, a gene from Bacteroides thetaiotaomicron that encoded a 115-kDa outer membrane protein. SusC protein proved to be essential for utilization not only of starch but also of intermediate-sized maltooligosaccharides (maltose to maltoheptaose). In this paper, we report the sequencing of a 7-kbp region of the B. thetaiotaomicron chromosome that lies immediately downstream of susC. We found four genes in this region (susD, susE, susF, and susG). Transcription of these genes was maltose inducible, and the genes appeared to be part of the same operon as susC. Western blot (immunoblot) analysis using antisera raised against proteins encoded by each of the four genes showed that all four were outer membrane proteins. Protein database searches revealed that SusE had limited similarity to a glucanohydrolase from Clostridium acetobutylicum and SusG had high similarity to amylases from a variety of sources. SusD and SusF had no significant similarity to any proteins in the databases. Results of 14C-starch binding assays suggested that SusD makes a major contribution to binding. SusE and SusF also appear to contribute to binding but not to the same extent as SusD. SusG is essential for growth on starch but appears to contribute little to starch binding. Our results demonstrate that the binding of starch to the B. thetaiotaomicron surface involves at least four outer membrane proteins (SusC, SusD, SusE, and SusF), which may form a surface receptor complex. The role of SusG in binding is still unclear.

Bacterial Outer Membrane Proteins↗

Towards establishing comprehensive databases of cellular proteins from transformed human epithelial amnion cells (AMA) and normal peripheral blood mononuclear cells.

Databases of protein information derived from the analysis of two-dimensional gels have been established from transformed human amnion cells (AMA) and peripheral blood mononuclear cells (PBMCs). A total of 1781 [35S]methionine-labeled AMA proteins (1274 IEF, 537 NEPHGE) and a total of 1311 proteins from PBMC (948 IEF, 363 NEPHGE) were resolved and recorded using computerized (PDQ-SCAN and PDQUEST softwares) two-dimensional gel electrophoresis. AMA and PBMC proteins (total, 454: 301 IEF, 153 NEPHGE) were matched both manually and by the computer. Information entered in the AMA database (in most cases for some major proteins) includes: molecular weight, protein name, HeLa protein catalogue number, mouse protein catalogue number, nuclear proteins, phosphorylated proteins, distribution of proteins in Triton X-100 supernatants and cytoskeletons, proliferation- and transformation-sensitive proteins, cell cycle-specific proteins, mitochondrial proteins, proteins matched in normal human embryonal lung MRC-5 fibroblasts and PBMC cells, heat shock proteins, proteins affected by interferons, cytoskeletal proteins, and the presence of antibody against protein in human sera. Additional information has been entered for the cell cycle-regulated and DNA replication protein cyclin (PCNA). Information entered in the PBMC database includes molecular weight and potential markers for sorted populations of lymphocyte subtypes. For those proteins that have been matched to AMA proteins, information contained in some entries may be transferred from the AMA database.

Amnion↗