PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.

We explore the ability of a simple simulated annealing procedure to assemble native-like structures from fragments of unrelated protein structures with similar local sequences using Bayesian scoring functions. Environment and residue pair specific contributions to the scoring functions appear as the first two terms in a series expansion for the residue probability distributions in the protein database; the decoupling of the distance and environment dependencies of the distributions resolves the major problems with current database-derived scoring functions noted by Thomas and Dill. The simulated annealing procedure rapidly and frequently generates native-like structures for small helical proteins and better than random structures for small beta sheet containing proteins. Most of the simulated structures have native-like solvent accessibility and secondary structure patterns, and thus ensembles of these structures provide a particularly challenging set of decoys for evaluating scoring functions. We investigate the effects of multiple sequence information and different types of conformational constraints on the overall performance of the method, and the ability of a variety of recently developed scoring functions to recognize the native-like conformations in the ensembles of simulated structures.

Bayes Theorem↗

A role for Edman degradation in proteome studies.

Advances in protein database design and the software used to access the sequence data has led to progress in using protein attributes such as amino acid composition and peptide masses to identify proteins separated by two-dimensional electrophoresis. However, Edman degradation remains the principal technique for protein identification and it presents a significant bottleneck in the progress towards rapid protein identification. Simple modifications to the sequencing hardware, which automate the delivery of protein spots into the sequencer, and parallel sequencing of the protein spots represent a significant advance in the use of Edman degradation to rapidly generate the powerful protein attribute, an N-terminal sequence tag.

Amino Acid Sequence↗

Chaos game representation of protein structures.

Chaos game representation (CGR) was proposed recently to visualize nucleotide sequences as one of the first applications of this technique in the field of biochemistry. In this paper we would like to demonstrate that representations similar to CGR can be generalized and applied for visualizing and analyzing protein databases. Examples of applications will be presented for investigating regularities, and motifs in the primary structure of proteins, and for analyzing possible structural attachments on the super-secondary structure level of proteins. A further application will be presented for testing structure prediction methods using CGR.

Computer Simulation↗

BIA/MS of epitope-tagged peptides directly from E. coli lysate: multiplex detection and protein identification at low-femtomole to subfemtomole levels.

The use of biomolecular interaction analysis mass spectrometry to selectively isolate, detect, and characterize epitope-tagged peptides present in total cell lysates is demonstrated. Epitope-tagged tryptic peptides were captured via affinity interactions with either chelated Ni2+ or monoclonal antibodies and detected using surface plasmon resonance biomolecular interaction analysis (SPR-BIA). After SPR-BIA the tagged peptides were either eluted from the biosensor chips for mass spectrometric analysis or analyzed directly from the biosensor chip using matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF). Protein database searches were performed using the masses of the tagged tryptic peptides, resulting in identification of the protein into which the epitope tag was inserted. Detection limits for both SPR-BIA and MALDI-TOF were at the low-femtomole to subfemtomole level. The approach represents a (multiplexed) high-sensitivity chip-based technique capable of identifying epitope-tagged proteins as they are present in complex mixtures.

Amino Acid Sequence↗

Odorant-binding proteins from a primitive termite.

Hitherto, odorant-binding proteins (OBPs) have been identified from insects belonging to more highly evolved insect orders (Lepidoptera, Coleoptera, Diptera, Hymenoptera, and Hemiptera), whereas only chemosensory proteins have been identified from more primitive species, such as orthopteran and phasmid species. Here, we report for the first time the isolation and cloning of odorant-binding proteins from a primitive termite species, the dampwood termite. Zootermopsis nevadensis nevadensis (Isoptera: Termopsidae). A major antennae-specific protein was detected by native PAGE along with four other minor proteins, which were also absent in the extract from control tissues (hindlegs). Multiple cDNA cloning led to the full characterization of the major antennae-specific protein (ZnevOBP1) and to the identification of two other antennae-specific cDNAs, encoding putative odorant-binding proteins (ZnevOBP2 and ZnevOBP3). N-terminal amino acid sequencing of the minor antennal bands and cDNA cloning showed that olfaction in Z. n. nevadensis may involve multiple odorant-binding proteins. Database searches suggest that the OBPs from this primitive termite are homologues of the pheromone-binding proteins from scarab beetles and antennal-binding proteins from moths.

Amino Acid Sequence↗

Proteomic characterization of two snake venoms: Naja naja atra and Agkistrodon halys.

Snake venom is a complex mixture of proteins and peptides, and a number of studies have described the biological properties of several venomous proteins. Nevertheless, a complete proteomic profile of venom from any of the many species of snake is not available. Proteomics now makes it possible to globally identify proteins from a complex mixture. To assess the venom proteomic profiles from Naja naja atra and Agkistrodon halys, snakes common to southern China, we used a combination strategy, which included the following four different approaches: (i) shotgun digestion plus HPLC with ion-trap tandem MS, (ii) one-dimensional SDS/PAGE plus HPLC with tandem MS, (iii) gel filtration plus HPLC with tandem MS and (iv) gel filtration and 2DE (two-dimensional gel electrophoresis) plus MALDI-TOF (matrix-assisted laser desorption ionization-time-of-flight) MS. In the present paper, we report the novel identification of 124 and 74 proteins and peptides in cobra and viper venom respectively. Functional analysis based upon toxin categories reveals that, as expected, cobra venom has a high abundance of cardio- and neurotoxins, whereas viper venom contains a significant amount of haemotoxins and metalloproteinases. Although approx. 80% of gel spots from 2DE displayed high-quality MALDI-TOF-MS spectra, only 50% of these spots were confirmed to be venom proteins, which is more than likely to be a result of incomplete protein databases. Interestingly, these data suggest that post-translational modification may be a significant characteristic of venomous proteins.

Agkistrodon↗

Bacterial Ohr and OsmC paralogues define two protein families with distinct functions and patterns of expression.

Xanthomonas campestris Ohr (a protein involved in organic peroxide protection) and Escherichia coli OsmC (an osmotically inducible protein of unknown function) are related proteins. Database searches and phylogenetic analyses reveal that Ohr and OsmC homologues cluster into two related subfamilies of proteins widely distributed in both Gram-negative and Gram-positive bacteria. To determine if these two subfamilies are functionally distinct, ohr and osmC in Pseudomonas aeruginosa (a bacterium with one representative from each subfamily) were analysed. Only ohr mutants are hypersensitive to organic peroxide, and this phenotype can be restored by complementation with ohr but not osmC. In addition, expression of ohr was highly induced only by organic peroxides, and not by other oxidants or stresses. In contrast, osmC was induced by ethanol and osmotic stress. A similar pattern of regulation was observed for Ohr and OsmC homologues in the Gram-positive bacterium Deinococcus radiodurans, though uninduced expression was much higher and induction lower in this species. These data clearly support the conclusion that Ohr and OsmC define two functionally distinct subfamilies with distinct patterns of regulation.

Amino Acid Sequence↗

Systematic analysis of the combinatorial nature of epitopes recognized by TCR leads to identification of mimicry epitopes for glutamic acid decarboxylase 65-specific TCRs.

Accumulating evidence indicates that recognition by TCRs is far more degenerate than formerly presumed. Cross-recognition of microbial Ags by autoreactive T cells is implicated in the development of autoimmunity, and elucidating the recognition nature of TCRs has great significance for revelation of the disease process. A major drawback of currently used means, including positional scanning synthetic combinatorial peptide libraries, to analyze diversity of epitopes recognized by certain TCRs is that the systematic detection of cross-recognized epitopes considering the combinatorial effect of amino acids within the epitope is difficult. We devised a novel method to resolve this issue and used it to analyze cross-recognition profiles of two glutamic acid decarboxylase 65-autoreactive CD4(+) T cell clones, established from type I diabetes patients. We generated a DNA-based randomized epitope library based on the original glutamic acid decarboxylase epitope using class II-associated invariant chain peptide-substituted invariant chains. The epitope library was composed of seven sublibraries, in which three successive residues within the epitope were randomized simultaneously. Analysis of agonistic epitopes indicates that recognition by both TCRs was significantly affected by combinations of amino acids in the antigenic peptide, although the degree of combinatorial effect differed between the two TCRs. Protein database searching based on the TCR recognition profile proved successful in identifying several microbial and self-protein-derived mimicry epitopes. Some of the identified mimicry epitopes were actually produced from recombinant microbial proteins by APCs to stimulate T cell clones. Our data demonstrate the importance of the combinatorial nature of amino acid residues of epitopes in molecular mimicry.

Amino Acid Motifs↗

[Epitope of somatic mutations in the hepatocyte growth factor receptor naturally processed and presented by HLA-A2 on human hepatocellular carcinoma].

OBJECTIVE: To isolate and identify peptides bound to HLA I on human hepatocellular carcinoma. METHODS: MHC-associated peptides were extracted by mild acid wash of viable hepatocellular carcinomas cells, collected by gel filtration, and fractioned by reversed phase high pressured liquid chromatography (RP-HPLC). Peptides of individual fractions were reconstitution of T cell epitopes and were identified by cytotoxicity T lymophacyte assay. Actively extracted sample analyses were performed by HPLC-MS-MS (tandom mass spectrometry). Protein database in the internet was used as an additional tool for structure analysis and for determination of protein source of the eluted peptides. RESULT: RP-HPLC showed that there were over 20 different fractions of peptides derived from human hepatocellular carcinoma, and only two active peaks were identified by cytotoxicity T lymophocyte assay. The most promising candidate for T cell epitope was nonamers peptide (SLIVHLNEV), derived from met/hepatocyte growth factor receptor, point mutation was 1180F to L. CONCLUSION: Sensitive sequencing by HPLC-MS may provide a powerful method of identifying tumor specific antigenic peptides, and nonamers peptide (SLIVHLNEV) is the hepatocellular carcinoma antigenic peptide.

Amino Acid Sequence↗

Cloning of a cDNA encoding a novel human nuclear phosphoprotein belonging to the WD-40 family.

We have cloned and expressed in vaccinia virus a cDNA encoding an ubiquitous 501-amino-acid (aa) phosphoprotein that corresponds to protein IEF SSP 9502 (79,400 Da, pI 4.5) in the master 2-D-gel keratinocyte protein database [Celis et al., Electrophoresis 14 (1993) 1091-1198]. The deduced aa sequence contains 9 Trp residues, some of which are localized in repeats and that characterise the protein as a member of the WD-40 family, a group of proteins having 40-aa repeats containing Trp and Asp [Duronio et al., Proteins 13 (1992) 41-56; Van der Voorn and Ploegh, FEBS Lett. 307 (1992) 131-134]. The protein contains a nuclear targeting signal (KKKGK), and fractionation of transformed human amnion cells (AMA) in karyoplasts and cytoplasts confirmed that it is predominantly localized in the nucleus. Database searching indicated that IEF SSP 9502 is a putative human homologue of the Saccharomyces cerevisiae periodic Trp protein, PWP1, a polypeptide that may play a regulatory role in cell growth and/or transcription.

Amino Acid Sequence↗

HSC-2DPAGE and the two-dimensional gel electrophoresis database of dog heart proteins.

A two-dimensional gel electrophoresis database of dog (Canis familiaris) proteins is presented. The database contains 1212 protein spots which have been characterised in terms of their pI and Mr. This database has been integrated into the HSC-2DPAGE database which is accessible on the Internet via the World Wide Web with the uniform resource location (URL): (http://www.harefield.nthames.nhs.uk/nhli/ protein/index.html). Identifications for 80 of the protein spots have been obtained by visual cross-matching with the human heart protein database in HSC-2DPAGE (42 spots), N-terminal microsequence analysis (25 spots) and peptide mass fingerprinting (20 spots). This database is being used in studies of alterations in protein expression in models of heart failure and heart disease.

Amino Acid Sequence↗

PTGL--a web-based database application for protein topologies.

Protein Topology Graph Library (PTGL) is a database application for the representation and retrieval of protein topologies. Protein topologies are based on a graph-theoretical protein model at secondary structure level. Different views on protein topology are given by four linear notations for their characterization. Protein topologies can be derived at different description levels considering alpha- and beta-structures. The on-line search tool is based on an object-relational database and provides a query browser for data interrogation by string patterns, keyword queries and sequence similarity. Protein topologies are represented both as schematic diagrams and as three-dimensional images.

Computer Graphics↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

Database of structural motifs in proteins.

SUMMARY: The database of structural motifs in proteins (DSMP) contains data relevant to helices, beta-turns, gamma-turns, beta-hairpins, psi-loops, beta-alpha-beta motifs, beta-sheets, beta-strands and disulphide bridges extracted from all proteins in the Protein Data Bank primarily using the PROMOTIF program and implemented as a web-based network service using the SRS. The data corresponding to the structural motifs includes; sequence, position in polypeptide chain, geometry, type, unique code, keywords and resolution of crystal structure. This data is available for a representative data set of 1028 protein chains and also for all 10 213 proteins in the Protein Data Bank. The three-dimensional coordinates for all structural motifs (except sheet and disulphide bridge) are also available for the representative data set. Using features in SRS, DSMP can be queried to extract information from one or more structural motifs that may be useful for sequence-structure analysis, prediction, modelling or design. AVAILABILITY: http://www. cdfd.org.in/dsmp.html

Amino Acid Motifs↗

Arf proteins bind to mitotic kinesin-like protein 1 (MKLP1) in a GTP-dependent fashion.

Arf proteins comprise a family of 21-kDa GTP-binding proteins with many proposed functions in mammalian cells, including the regulation of several steps of membrane transport, maintenance of organelle integrity, and activation of phospholipase D. We performed a yeast two-hybrid screen of human cDNA libraries using a dominant activating allele, [Q71L], of human Arf3 as bait. Eleven independent isolates contained plasmids encoding the C-terminal tail of mitotic kinesin-like protein-1 (MKLP1). Further deletion mapping allowed the identification of an 88 amino acid Arf3 binding domain in the C-terminus of MKLP1. This domain has no clear homology to other Arf binding proteins or to other proteins in the protein databases. The C-terminal domain of MKLP1 was expressed and purified from bacteria as a GST fusion protein and shown to bind Arf3 in a GTP-dependent fashion. A screen for mutations in Arf3 that specifically lost the ability to bind MKLP1 identified 10 of 14 point mutations in the GTP-sensitive switch I or switch II regions of Arf3. Two-hybrid assays of the C-terminal domain of MKLP1 with each of the human Arf isoforms revealed strong interaction with each. Taken together, these data are all supportive of the conclusion that activated Arf proteins bind to the C-terminal "tail" domain of MKLP1.

ADP-Ribosylation Factors↗

Ability of trypsin in mimicking germ cell factors that affect Sertoli cell secretory function.

A biological factor that inhibits the in vitro secretion of testin by Sertoli cells was purified to apparent homogeneity from conditioned medium of germ cells isolated using trypsin. Partial N-terminal amino acid sequence analysis of the purified germ cell factor revealed a sequence of NH2-IVGGYTXAAN. Comparison of the sequence with the existing protein database revealed that it is homologous to trypsin. Immunoprecipitation experiments using either [35S]-labeled germ or Sertoli cell proteins and a monospecific anti-trypsin antibody failed to demonstrate the synthesis and secretion of trypsin by these testicular cells, suggesting the isolated factor is the residuary trypsin that was used for isolating germ cells from seminiferous tubules. Subsequent experiments revealed that trypsin per se can inhibit the secretion of Sertoli cell testin and clusterin dose-dependently, whose effect can be prohibited by soybean trypsin inhibitor (STI). In view of these findings, a nonenzymatic procedure was deemed necessary to prepare germ cell conditioned medium (GCCM) to assess whether an authentic biological factor(s) is indeed present. Four batches of conditioned medium of germ cells isolated by a mechanical procedure without the use of trypsin were fractionated by sequential Mono Q anion exchange and C8 reversed-phase HPLC. When these fractions were monitored for testin modulatory activity using an in vitro bioassay with primary cultures of Sertoli cells, it was shown that GCCM prepared by this procedure indeed contained testin modulatory bioactivity. Since testin is a novel component of specialized junctions between Sertoli and germ cells, the identification of a germ cell factor(s) that affects its secretion by Sertoli cells suggests a dynamic biochemical relationship between these cell types in the seminiferous epithelium.

Amino Acid Sequence↗

Characterisation of proteins from two-dimensional electrophoresis gels by matrix-assisted laser desorption mass spectrometry and amino acid compositional analysis.

Amino acid compositional analysis and peptide mass fingerprinting by matrix assisted laser desorption mass spectrometry have been used to characterise proteins obtained from two-dimensional electrophoresis (2-DE) separations of human cardiac proteins. A group of twelve protein spots was selected for analysis. The identities of eight of the proteins had been determined by conventional protein characterisation methods, two were unknown proteins and two had putative identities from protein database spot comparison. Amino acid analysis and peptide mass fingerprinting gave corresponding identities for seven of the twelve proteins, which also agreed with our initial identifications. Three proteins which had been identified previously were not confirmed in this study and putative identities were obtained for the two unknown proteins. The advantages, problems and use of amino acid analysis and peptide mass fingerprinting for the analysis of proteins from 2-DE are discussed. The data highlight the need to use orthogonal techniques for the unequivocal identification of proteins from 2-DE gels.

Amino Acids↗

Cloning and characterization of the human retina-specific gene MPP4, a novel member of the p55 subfamily of MAGUK proteins.

To identify novel retina-specific genes systematically, we are performing expression profiling of retina ESTs that have been assembled in the human UniGene clusters. In this study, we report the 2619-bp full-length cDNA cloning and genomic organization of a gene corresponding to an EST cluster that was demonstrated to be exclusively present in retinal tissue. Alignment of the deduced amino acid sequence to sequence from protein databases revealed this gene, termed MPP4, to be a member of the membrane-associated guanylate kinase (MAGUK) protein family. It consists of 637 amino acids and contains the characteristic MAGUK motifs: an N-terminal PDZ domain, a central src homology 3 region (SH3), and a C-terminal guanylate kinase-like (GUK) domain. Due to the presence of only one PDZ motif, MPP4 is part of the p55 subfamily, named after the major palmitoylated erythrocyte membrane protein p55/MPP1. MAGUK proteins serve as molecular scaffolds to coordinate the membrane-associated cytoskeleton, ion channel and receptor clustering, signaling pathways, and the formation of cellular junctions. The abundant expression of MPP4 in the human retina suggests an important but so far unknown function in this tissue. Colocalization of MPP4 and autosomal recessive retinitis pigmentosa 26 (RP26) on chromosome 2q31-q33 makes this transcript an attractive candidate for the disease gene.

Alternative Splicing↗