PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Mapping cells and sub-cellular organelles on 2-D gels: 'new tricks for an old horse'.

Nowadays, investigators in all fields are faced with the identification of unknown, up- or down-regulated, modified proteins that they are trying to identify. Two-dimensional (2-D) gel electrophoresis, with its ability to resolve several thousand proteins, is an extremely powerful technique. The current resolution and reproducibility of 2-D gel technology and the establishment of computer assisted 2-D gel protein databases have paved new ways for the identification of proteins.

Cell Fractionation↗

Systematic identification of immunoreceptor tyrosine-based inhibitory motifs in the human proteome.

Immunoreceptor tyrosine-based inhibitory motifs (ITIMs) are short sequences of the consensus (ILV)-x-x-Y-x-(LV) in the cytoplasmic tail of immune receptors. The phosphorylation of tyrosines in ITIMs is known to be an important signalling mechanism regulating the activation of immune cells. The shortness of the motif makes it difficult to predict ITIMs in large protein databases. Simple pattern searches find ITIMs in approximately 30% of the protein sequences in the RefSeq database. The majority are false positive predictions. We propose a new database search strategy for ITIM-bearing transmembrane receptors based on the use of sequence context, i.e. the predictions of signal peptides, transmembrane helices (TMs) and protein domains. Our new protocol allowed us to narrow down the number of potential human ITIM receptors to 109 proteins (0.7% of RefPep). Of these, 36 have been described as ITIM receptors in the literature before. Many ITIMs are conserved between orthologous human and mouse proteins which represent novel ITIM receptor candidates. Publicly available DNA array expression data revealed that ITIM receptors are not exclusively expressed in blood cells. We hypothesise that ITIM signalling is not restricted to immune cells, but also functions in diverse solid organs of mouse and man.

Computational Biology↗

Cluster analysis of an extensive human breast cancer cell line protein expression map database.

In the current study, the protein expression maps (PEMs) of 26 breast cancer cell lines and three cell lines derived from normal breast or benign disease tissue were visualised by high resolution two-dimensional gel electrophoresis. Analysis of this data was performed with ChiClust and ChiMap, two analytical bioinformatics tools that are described here. These tools are designed to facilitate recognition of specific patterns shared by two or more (a series) PEMs. Both tools use PEMs that were matched by an image analysis program and locally written programs to create a match table that is saved in an object relational database. The ChiClust tool uses clustering and subclustering methods to extract statistically significant protein expression patterns from a large series of PEMs. The ChiMap tool calculates a differential value (either as percentage change or a fold change) and represents these graphically. All such differentials or just those identified using ChiClust can be submitted to ChiMap. These methods are not dependent on any particular commercial image analysis program, and the whole software package gives an integrated procedure for the comparison and analysis of a series of PEMs. The ChiClust tool was used here to order the breast cell lines into groups according to biological characteristics including morphology in vitro and tumour forming ability in vivo. ChiMap was then used to highlight eight major protein feature-changes detected between breast cancer cell lines that either do or do not proliferate in nude mice. Mass spectrometry was used to identify the proteins. The possible role of these proteins in cancer is discussed.

Algorithms↗

Evaluation of 515 expressed sequence tags obtained from guard cells of Brassica campestris.

As an attempt to examine the transcripts expressed in a single cell type and to unveil the physiology of guard cells at the molecular level, we generated 515 expressed sequence tags (ESTs) from a directional cDNA library constructed from guard-cell protoplasts of Brassica campestris L. ssp. pekinensis. A comparative analysis of the guard-cell ESTs against the National Center for Biotechnological Information non-redundant protein database revealed that 133 ESTs (26%) have significant similarity to protein coding sequences in the database. Among them were 35 clones related to genes that have not yet been identified in higher plants. Analysis of RNA gel blots of 14 database-matched clones revealed that five clones harbor the sequences for mRNAs expressed most abundantly in guard cells, one of them detecting an mRNA with highly preferential expression in guard cells. Functional categorization of the putatively identified guard-cell ESTs showed, when compared with maize leaf ESTs, that guard cells expressed a higher proportion of signal transduction components and a lower proportion of structural or photosynthetic genes, as is consistent with the roles of guard cells.

Brassica↗

Down-regulation of the strawberry Bet v 1-homologous allergen in concert with the flavonoid biosynthesis pathway in colorless strawberry mutant.

Proteomic screening of strawberry (Fragaria ananassa) yielded a 58% success rate in protein identification in spite of the fact that no genomic sequence is available for this species. This was achieved by a combination of MALDI-MS/MS de novo sequencing of double-derivatized peptides and indel-tolerant searching against local protein databases built on both EST and full-length nucleotide sequences. The amino acid sequence of a strawberry allergen, homologous to the well-known major birch pollen allergen Bet v 1, was partially determined. This strawberry allergen, named Fra a 1 according to the nomenclature for allergen proteins, showed sequence identity of 54 and 77%, respectively, with corresponding allergens from birch and apple. Differential expression, as evaluated by 2-D DIGE, occurred in 10% of protein spots when red strawberries were compared to a colorless (white) strawberry mutant. White strawberries, known to be tolerated by individuals affected by allergy, were found to be virtually free from the strawberry allergen. Also several enzymes in the pathway for biosynthesis of flavonoids, to which the red color pelargonidin belongs, were down-regulated. This approach to assess differential protein expression without access to genomic sequence information can also be applied to other crop plants and phenotypic traits.

Allergens↗

Allermatch, a webtool for the prediction of potential allergenicity according to current FAO/WHO Codex alimentarius guidelines.

BACKGROUND: Novel proteins entering the food chain, for example by genetic modification of plants, have to be tested for allergenicity. Allermatch http://allermatch.org is a webtool for the efficient and standardized prediction of potential allergenicity of proteins and peptides according to the current recommendations of the FAO/WHO Expert Consultation, as outlined in the Codex alimentarius. DESCRIPTION: A query amino acid sequence is compared with all known allergenic proteins retrieved from the protein databases using a sliding window approach. This identifies stretches of 80 amino acids with more than 35% similarity or small identical stretches of at least six amino acids. The outcome of the analysis is presented in a concise format. The predictive performance of the FAO/WHO criteria is evaluated by screening sets of allergens and non-allergens against the Allermatch databases. Besides correct predictions, both methods are shown to generate false positive and false negative hits and the outcomes should therefore be combined with other methods of allergenicity assessment, as advised by the FAO/WHO. CONCLUSIONS: Allermatch provides an accessible, efficient, and useful webtool for analysis of potential allergenicity of proteins introduced in genetically modified food prior to market release that complies with current FAO/WHO guidelines.

Allergens↗

A comprehensive update of the sequence and structure classification of kinases.

BACKGROUND: A comprehensive update of the classification of all available kinases was carried out. This survey presents a complete global picture of this large functional class of proteins and confirms the soundness of our initial kinase classification scheme. RESULTS: The new survey found the total number of kinase sequences in the protein database has increased more than three-fold (from 17,310 to 59,402), and the number of determined kinase structures increased two-fold (from 359 to 702) in the past three years. However, the framework of the original two-tier classification scheme (in families and fold groups) remains sufficient to describe all available kinases. Overall, the kinase sequences were classified into 25 families of homologous proteins, wherein 22 families (approximately 98.8% of all sequences) for which three-dimensional structures are known fall into 10 fold groups. These fold groups not only include some of the most widely spread proteins folds, such as the Rossmann-like fold, ferredoxin-like fold, TIM-barrel fold, and antiparallel beta-barrel fold, but also all major classes (all alpha, all beta, alpha+beta, alpha/beta) of protein structures. Fold predictions are made for remaining kinase families without a close homolog with solved structure. We also highlight two novel kinase structural folds, riboflavin kinase and dihydroxyacetone kinase, which have recently been characterized. Two protein families previously annotated as kinases are removed from the classification based on new experimental data. CONCLUSION: Structural annotations of all kinase families are now revealed, including fold descriptions for all globular kinases, making this the first large functional class of proteins with a comprehensive structural annotation. Potential uses for this classification include deduction of protein function, structural fold, or enzymatic mechanism of poorly studied or newly discovered kinases based on proteins in the same family.

Algorithms↗

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine↗

PDB-REPRDB: a database of representative protein chains from the Protein Data Bank (PDB).

PDB-REPRDB is a database of representative protein chains from the Protein Data Bank (PDB). The previous version of PDB-REPRDB provided 48 representative sets, whose similarity criteria were predetermined, on the WWW. The current version is designed so that the user may obtain a quick selection of representative chains from PDB. The selection of representative chains can be dynamically configured according to the user's requirement. The WWW interface provides a large degree of freedom in setting parameters, such as cut-off scores of sequence and structural similarity. One can obtain a representative list and classification data of protein chains from the system. The current database includes 20 457 protein chains from PDB entries (August 6, 2000). The system for PDB-REPRDB is available at the Parallel Protein Information Analysis system (PAPIA) WWW server (http://www.rwcp.or.jp/papia/).

Databases, Factual↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (aka EMBL-Bank; http://www.ebi.ac.uk/embl/) incorporates, organises and distributes nucleotide sequences from all available public sources. EMBL-Bank is located and maintained at the European Bioinformatics Institute (EBI) near Cambridge, UK. In an international collaboration with DDBJ (Japan) and GenBank (USA), data are exchanged amongst the collaborating databases on a daily basis. Major contributors to the EMBL database are individual scientists and genome project groups. Webin is the preferred web-based submission system for individual submitters, whilst automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO). Database releases are produced quarterly. Network services allow free access to the most up-to-date data collection via FTP, email and World Wide Web interfaces. EBI's Sequence Retrieval System (SRS), a network browser for databanks in molecular biology, integrates and links the main nucleotide and protein databases plus many other specialized databases. For sequence similarity searching, a variety of tools (e.g. Blitz, Fasta, BLAST) are available which allow external users to compare their own sequences against the latest data in the EMBL Nucleotide Sequence Database and SWISS-PROT. All resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗

Proteome studies of Saccharomyces cerevisiae: identification and characterization of abundant proteins.

Two-dimensional (2-D) gel electrophoresis can now be coupled with protein identification techniques and genome sequence information for direct detection, identification, and characterization of large numbers of proteins from microbial organisms. 2-D electrophoresis, and new protein identification techniques such as amino acid composition, are proteome research techniques in that they allow direct characterization of many proteins at the same time. Another new tool important for yeast proteome research is the Yeast Protein Database (YPD), which provides the sequence-derived protein properties needed for spot identification and tabulations of the currently known properties of the yeast proteins. Studies presented here extend the yeast 2-D protein map to 169 identified spots based upon the recent completion of the yeast genome sequence, and they show that methods of spot identification based on predicted isoelectric point, predicted molecular mass, and determination of partial amino acid composition from radiolabeled gels are powerful enough for the identification of at least 80% of the spots representing abundant proteins. Comparison of proteins predicted by YPD to be detectable on 2-D gels based on calculated molecular mass, isoelectric point and codon bias (a predictor of abundance) with proteins identified in this study suggests that many glycoproteins and integral membrane proteins are missing from the 2-D gel patterns. Using the 2-D gel map and the information available in YDP, 2-D gel experiments were analyzed to characterize the yeast proteins associated with: (i) an environmental change (heat shock), (ii) a temperature-sensitive mutation (the prp2 mRNA splicing mutant), (iii) a mutation affecting post-translational modification (N-terminal acetylation), and (iv) a purified subcellular fraction (the ribosomal proteins). The methods used here should allow future extension of these studies to many more proteins of the yeast proteome.

Databases, Factual↗

In-gel isoelectric focusing of peptides as a tool for improved protein identification.

In the analysis of proteins in complex samples, pre-fractionation is imperative to obtain the necessary depth in the number of reliable protein identifications by mass spectrometry. Here we explore isoelectric focusing of peptides (peptide IEF) as an effective fractionation step that at the same time provides the added possibility to eliminate spurious peptide identifications by filtering for pI. Peptide IEF in IPG strips is fast and sharply confines peptides to their pI. We have evaluated systematically the contribution of pI filtering and accurate mass measurements on the total number of protein identifications in a complex protein mixture (Drosophila nuclear extract). At the same time, by varying Mascot identification cutoff scores, we have monitored the false positive rate among these identifications by searching reverse protein databases. From mass spectrometric analyses at low mass accuracy using an LTQ ion trap, false positive rates can be minimized by filtering of peptides not focusing at their expected pI. Analyses using an LTQ-FT mass spectrometer delivers low false positive rates by itself due to the high mass accuracy. In a direct comparison of peptide IEF with SDS-PAGE as a pre-fractionation step, IEF delivered 25% and 43% more proteins when identified using FT-MS and LTQ-MS, respectively. Cumulatively, 2190 non redundant proteins were identified in the Drosophila nuclear extract at a false positive rate of 0.5%. Of these, 1751 proteins (80%) were identified after peptide IEF and FT-MS alone. Overall, we show that peptide IEF allows to increase the confidence level of protein identifications, and is more sensitive than SDS-PAGE.

Animals↗

PSIST: indexing protein structures using suffix trees.

Approaches for indexing proteins, and for fast and scalable searching for structures similar to a query structure have important applications such as protein structure and function prediction, protein classification and drug discovery. In this paper, we developed a new method for extracting the local feature vectors of protein structures. Each residue is represented by a triangle, and the correlation between a set of residues is described by the distances between Calpha atoms and the angles between the normals of planes in which the triangles lie. The normalized local feature vectors are indexed using a suffix tree. For all query segments, suffix trees can be used effectively to retrieve the maximal matches, which are then chained to obtain alignments with database proteins. Similar proteins are selected by their alignment score against the query. Our results shows classification accuracy up to 97.8% and 99.4% at the superfamily and class level according to the SCOP classification, and shows that on average 7.49 out of 10 proteins from the same superfamily are obtained among the top 10 matches. These results are competitive with the best previous methods.

Algorithms↗

Identification of a Spiroplasma citri hydrophilic protein associated with insect transmissibility.

With the aim of identifying Spiroplasma citri proteins involved in transmission by the leafhopper Circulifer haematoceps, protein maps of four transmissible and four non-transmissible strains were compared. Total cell lysates of strains were analysed by two-dimensional gel electrophoresis using commercially available immobilized pH gradients (IPGs) covering a pH range of 4-7. Approximately 530 protein spots were visualized by silver staining and the resulting protein spot patterns for the eight strains were found to be highly similar. However, comparison using PDQuest 2-D analysis software revealed two trains of protein spots that were present only in the four transmissible strains. Using MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight) mass spectrometry and a nearly complete S. citri protein database, established during the still-ongoing S. citri GII-3-3X genome project, the sequences of both proteins were deduced. One of these proteins was identified in the general databases as adhesion-related protein (P89) involved in the attachment of S. citri to gut cells of the insect vector. The second protein, with an apparent molecular mass of 32 kDa deduced from the electrophoretic mobility, could not be assigned to a known protein and was named P32. The P32-encoding gene (714 bp) was carried by a large plasmid of 35.3 kbp present in transmissible strains and missing in non-transmissible strains. PCR products with primers designed from the p32 gene were obtained only with genomic DNA isolated from transmissible strains. Therefore, P32 has a putative role in the transmission process and it could be considered as a marker for S. citri leafhopper transmissibility. Functional complementation of a non-transmissible strain with the p32 gene did not restore the transmissible phenotype, despite the expression of P32 in the complemented strain. Electron microscopic observations of salivary glands of leafhoppers infected with the complemented strain revealed a close contact between spiroplasmas and the plasmalemma of the insect cells. This further suggests that P32 protein contributes to the association of S. citri with host membranes.

Adhesins, Bacterial↗

Proteome analysis. Novel proteins identified at the peribacteroid membrane from Lotus japonicus root nodules.

The peribacteroid membrane (PBM) forms the structural and functional interface between the legume plant and the rhizobia. The model legume Lotus japonicus was chosen to study the proteins present at the PBM by proteome analysis. PBM was purified from root nodules by an aqueous polymer two-phase system. Extracted proteins were subjected to a global trypsin digest. The peptides were separated by nanoscale liquid chromatography and analyzed by tandem mass spectrometry. Searching the nonredundant protein database and the green plant expressed sequence tag database using the tandem mass spectrometry data identified approximately 94 proteins, a number far exceeding the number of proteins reported for the PBM hitherto. In particular, a number of membrane proteins like transporters for sugars and sulfate; endomembrane-associated proteins such as GTP-binding proteins and vesicle receptors; and proteins involved in signaling, for example, receptor kinases, calmodulin, 14-3-3 proteins, and pathogen response-related proteins, including a so-called HIR protein, were detected. Several ATPases and aquaporins were present, indicating a more complex situation than previously thought. In addition, the unexpected presence of a number of proteins known to be located in other compartments was observed. Two characteristic protein complexes obtained from native gel electrophoresis of total PBM proteins were also analyzed. Together, the results identified specific proteins at the PBM involved in important physiological processes and localized proteins known from nodule-specific expressed sequence tag databases to the PBM.

Cell Membrane↗

Spectral magnitude effects on the analyses of secondary structure from circular dichroism spectroscopic data.

The effects of spectral magnitude on the calculated secondary structures derived from circular dichroism (CD) spectra were examined for a number of the most commonly used algorithms and reference databases. Proteins with different secondary structures, ranging from mostly helical to mostly beta-sheet, but which were not components of existing reference databases, were used as test systems. These proteins had known crystal structures, so it was possible to ascertain the effects of magnitude on both the accuracy of determining the secondary structure and the goodness-of-fit of the calculated structures to the experimental data. It was found that most algorithms are highly sensitive to spectral magnitude, and that the goodness-of-fit parameter may be a useful tool in assessing the correct scaling of the data. This means that parameters that affect magnitude, including calibration of the instrument, the spectral cell pathlength, and the protein concentration, must be accurately determined to obtain correct secondary structural analyses of proteins from CD data using empirical methods.

Albumins↗

Differential proteomic analysis of proteins in wheat spikes induced by Fusarium graminearum.

Scab, caused by Fusarium graminearum, is a serious spike disease in wheat. To identify proteins in resistant wheat cultivar Wangshuibai induced by F. graminearum infection, proteins extracted from spikes 6, 12 and 24 h after inoculation were separated by 2-DE. Thirty protein spots showing 3-fold change in abundance when compared with treatment without inoculation were characterized by MALDI-TOF MS and matched to proteins by querying the mass spectra in protein databases or the Triticeae EST translation database. Based on their volume profiles, these proteins were classified into four categories. The first one fell off rapidly at the initial inoculation and then rose at 12 or 24 hai, the second one decreased considerably after inoculation and remained at low level, the third one rose at the initial inoculation and then declined at 12 or 24 hai, the forth one showed steady increase after inoculation and maintained at a high level. Many of the proteins identified in the first two categories are related to carbon metabolism and photosynthesis. While most of proteins identified in the last two categories are related to stress defense of plants, indicating that proteins associated with the defense reactions were activated or translated shortly after inoculation.

Electrophoresis, Gel, Two-Dimensional↗

The mouse male germ cell-specific gene Tpx-1: molecular structure, mode of expression in spermatogenesis, and sequence similarity to two non-mammalian genes.

Tpx-1 is a testis-specific gene that maps on mouse Chromosome (Chr) 17. The deduced TPX-1 protein shows 55% amino acid sequence similarity to acidic epididymal glycoprotein (AEG), assumed to be involved in sperm maturation. In the present study, we determined the genomic structure of the mouse Tpx-1 gene and the cellular localization of its transcripts. The gene was found to contain ten exons, with an unusually large intron (approximately 17.0 kilobase pairs) between exons 8 and 9. In situ hybridization of testicular sections showed that Tpx-1 is transcribed abundantly by haploid male germ cells. A computer search of protein databases revealed that deduced TPX-1/AEG proteins have significant sequence similarity (approximately 30%) to two non-mammalian proteins: "pathogenesis-related" proteins 1 of tobaccos, and venom sac proteins of white-face hornets, known as Dol m V. Amino acid residues encoded by exon 10 of the Tpx-1 gene and most of those encoded by exon 9 were absent in the non-mammalian proteins. This result suggests that the ancestor of Tpx-1 acquired exons 9 and 10 after its divergence from the ancestors of the plant and insect proteins.

Amino Acid Sequence↗