PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Database searching with DNA and protein sequences: an introduction.

This review of sequence database searching aims to set out current practice in the area, in order to give practical guidelines to the experimental biologist. It describes the basic principles behind the programs and enumerates the range of databases available in the public domain. Of these, the most important are the equivalent DNA databases European Molecular Biology Laboratory (EMBL), GenBank and DNA Databank of Japan (DDBJ), and the protein databases Swiss-Prot and TrEMBL. The commonly used BLAST and FASTA algorithms are described in detail and alternative approaches mentioned briefly. Scoring matrices used to compare amino acid types during protein database searches are compared, with an emphasis on the PAM and BLOSUM series of observed substitution matrices.

Algorithms↗

Homeobox genes in the Australian lungfish, Neoceratodus forsteri.

The aim of the present study was to determine whether the postulated gnathostome duplication from four to eight Hox clusters occurred before or after the split between the actinopterygian and sarcopterygian fish by characterizing Hox genes from the sarcopterygian lungfish, Neoceratodus forsteri. Since lungfish have extremely large genomes, we took the approach of extracting pure high molecular weight (MW) genomic DNA to act as a template for polymerase chain reaction (PCR) of the conserved homeobox domain of the highly conserved Hox genes. The 21 clones thus obtained were sequenced and translated in a BLASTX protein database search to designate Hox gene identity. Fourteen of the clones were from Hox genes, two were Hox pseudogenes, four were Gbx genes, and one most closely resembled the homeobox gene, insulin upstream factor 1. The Hox genes identified were from all four tetrapod clusters A, B, C, and D, confirming their presence in lungfish, and there is no evidence to suggest more than these four functional Hox clusters, as is the case in teleosts. A comparison of Hox group 13 amino acid sequences of lungfish, zebrafish, and mouse provides firm evidence that the expansion of Hox clusters, as seen in zebrafish, occurred after separation of the actinopterygian and sarcopterygian lineages. J. Exp. Zool. (Mol. Dev. Evol.) 285:140-145, 1999.

Amino Acid Sequence↗

Systematic analysis of the total proteins of a mammalian organism: principles, problems and implications for sequencing the human genome.

High-resolution two-dimensional electrophoresis (2-DE) has reached a technological level that allows us to resolve most of the numerous unknown protein species of a mammalian organism if appropriate strategies are used. We will discuss the problems of classification and characterization of proteins and propose a systematic approach to the analysis of the total protein complex. Both a comprehensive as well as a pragmatic approach towards systematic analysis have been considered. A "complex protein database" is suggested and considered with regard to various uses. A systematic analysis of the mouse proteins has been started and some of the preliminary results are summarized here. In particular, genetic properties of the proteins were investigated and are presented in order to demonstrate the significance of a systematic analysis of proteins for research and practical application (e.g. mutagenicity testing). A concept is presented for sequencing the coding DNA of mouse and man, starting with a systematic analysis of mouse proteins and then using two recently developed methods - microsequencing of proteins from spots of 2-DE protein patterns, and utilization of the relatively short N-terminal sequences obtained - to produce the corresponding cDNA's of these proteins.

Animals↗

Proteomic analysis of nuclear proteins from proliferative and differentiated human colonic intestinal epithelial cells.

Self-renewing tissues such as the intestine contain progenitor proliferating cells which subsequently differentiate. Cell proliferation and differentiation involve gene regulation processes which take place in the nucleus. A human intestinal epithelial cell line model (Caco2/TC7) which reproduces these dynamic processes has been used to perform proteomic studies on nuclear proteins. Nuclei from Caco2/TC7 cells at proliferative and differentiated stages were purified by subcellular fractionation. After two-dimensional gel electrophoresis separation and ruthenium staining, 400 protein spots were detected by image analysis. Eighty-five spots corresponding to 60 different proteins were identified by matrix-assisted laser desorption/ionization mass spectrometry in nuclei from proliferative cells. Comparison of nuclear proteomes from proliferative or differentiated cells by differential display resulted in the identification of differentially expressed proteins such as nucleolin, hnRNP A2/B1 and hnRNP A1. By using Western blot analysis, we found that the expression and number of specific isoforms of these nuclear proteins decreased in differentiated cells. Immunocytochemistry experiments also showed that in proliferative cells nucleolin was distributed in nucleoli-like bodies. In contrast, hnRNPs A2/B1 and A1 were dispersed throughout the nucleus. This study of the nuclear proteome from intestinal epithelial cells represents the first step towards the establishment of a protein database which will be a valuable resource in future studies on the differential expression of nuclear proteins in response to physiological, pharmacological and pathological modulations.

Caco-2 Cells↗

Molecular mechanisms involved in the association of HLA-DR4 and rheumatoid arthritis.

Susceptibility to developing rheumatoid arthritis (RA) maps to a highly conserved amino acid motif located in the third hypervariable region of different HLA-DRB1 chains. This motif, namely QKRAA, QRRAA, or RRRAA, helps the development of RA by an unknown mechanism. The QKRAA motif predisposes to more severe disease than the QRRAA or RRRAA motifs. The QKRAA motif carries particular properties: it is a strong B- and T-cell epitope, it shapes the T cell repertoire, it is overrepresented in protein databases, and it is a binding motif for bacterial and human 70-kDa heat-shock proteins. In this article, we propose different models to explain how the QKRAA motif might contribute to RA.

Amino Acid Sequence↗

RNA-binding protein-related sequence in a malaria antigen, clustered-asparagine-rich protein.

Members of the RNA-binding protein superfamily contain RNA binding domains of about 90 amino acids with a highly conserved motif 'GFGF'. Using the conserved motif with some variations G-(F/Y)-(G/A)-(F/Y)-(V/I)-X-(F/Y) as a probe, we screened protein sequences carrying identical amino acids in an NBRF-protein database. It has been found that the C-terminal portion of clustered asparagine-rich protein (CARP), a malaria antigen from Plasmodium falciparum, shows an unexpected sequence similarity with the RNA-binding protein superfamily for the C-terminal half of the RNA-binding domain. Dot matrix comparisons and alignment of these sequences as well as a statistical test have revealed highly significant sequence similarities. From these analyses, we conclude that the malaria antigen CARP belongs to a large family of the RNA-binding proteins. An evolutionary implication of the sequence similarity was also discussed.

Amino Acid Sequence↗

Proteomic analysis of the synaptic plasma membrane fraction isolated from rat forebrain.

Mass spectrometry (MS) in conjunction with liquid chromatography and gel separation techniques has been utilized to identify synaptic plasma membrane (SPM) proteins isolated from rat forebrain and digested with the protease trypsin. Initial results employing two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) separation of the SPM protein mixture have shown that several membrane proteins were under-represented due to solubilization problems in the dimension of isoelectric-point focusing. Given the complexity of the SPM, multiple stages of separation were necessary prior to mass spectrometric detection in order to facilitate protein identification. This particular study involved several approaches using one-dimensional (1D) sodium dodecyl sulfate (SDS)-PAGE, strong cation-exchange (SCX) chromatography and capillary reversed-phase high performance liquid chromatography (HPLC) techniques. In addition to these gel and HPLC separation stages, complementary information was obtained by using both matrix-assisted laser desorption/ionization (MALDI) and electrospray ionization (ESI) mass spectrometry. Data-dependent acquisition employing capillary HPLC-nanoESI/MS allowed for the detection of low-abundance tryptic peptides in the digested SPM fraction and identification of the corresponding proteins when product-ion information of a single or multiple peptides was used in protein database searching. The potential value of this subproteome methodology was exemplified by the identification of several proteins relevant to synaptic physiology which included various transporters, receptors, ion channels, and enzymes.

Animals↗

Computational analysis of human disease-associated genes and their protein products.

The complete genome sequences for human, Drosophila melanogaster and Arabidopsis thaliana have been reported recently. With the availability of complete sequences for many bacteria and archaea, and five eukaryotes, comparative genomics and sequence analysis are enabling us to identify counterparts of many human disease genes in model organisms, which in turn should accelerate the pace of research and drug development to combat human diseases. Continuous improvement of specialized protein databases, together with sensitive computational tools, have enhanced the power and reliability of computational prediction of protein function.

Animals↗

Fine specificity of autoantibodies to La/SSB: epitope mapping, and characterization.

The B cell epitope mapping of La/SSB was performed using 20mer synthetic peptides overlapping by eight amino acids covering the whole sequence of the protein. IgG, purified from sera of five patients with systemic lupus erythematosus (SLE) and four sera from patients with primary Sjögren's syndrome (pSS) were tested against the overlapping synthetic peptides. Peptides highly reactive with purified IgG were those spanning the regions 145-164, 289-308, 301-320 and 349-368 of the La protein. Determination of the minimum required length of the antigenic determinants disclosed the following epitopes: 147HKAFKGSI154, 291NGNLQLRNKEVT302, 301VTWEVLEGEVEKEALKKI318 and 349GSGKGKVQFQGKKTKF364. Predicted features and molecular similarities of the defined epitopes were investigated using protein databases. The La epitope 147HKAFKGSI154 presented 83.3% similarity with the 139HKGFKGVD146 region of human myelin basic protein (MBP) and 72% similarity with the fragment YKNFKGTI of human DNA topoisomerase II. Peptides corresponding to these sequences cross-reacted with anti-La/SSB antibodies. Sixty-three sera with anti-La/SSB antibodies from patients with pSS or SLE, 35 sera without anti-La/SSB antibodies from patients with SS or SLE and 41 sera from age/sex-matched healthy blood donors were tested against biotinylated synthetic epitope analogues in order to determine their sensitivity and specificity for the detection of anti-La/SSB antibodies. Anti-La/SSB were detected with various frequencies ranging from 20% to epitope 147HKAFKGSI154 to 100% to epitope 349GSGKGKVQGKKTKF364. The overall sensitivity and specificity using all assays with the synthetic peptides were found to be 93.6% and 85.6%, respectively. In conclusion, antibodies to La/SSB constitute a heterogeneous population, directed against different linear B cell epitopes of the molecule. The epitope 147HKAFKGSI154 presents molecular similarity with fragments of two other autoantigens, i.e. human MBP and DNA topoisomerase II. Finally, synthetic epitope analogues exhibit high sensitivity and specificity for the detection of anti-La/SSB antibodies.

Adult↗

Gene discovery in the apicomplexa as revealed by EST sequencing and assembly of a comparative gene database.

Large-scale EST sequencing projects for several important parasites within the phylum Apicomplexa were undertaken for the purpose of gene discovery. Included were several parasites of medical importance (Plasmodium falciparum, Toxoplasma gondii) and others of veterinary importance (Eimeria tenella, Sarcocystis neurona, and Neospora caninum). A total of 55192 ESTs, deposited into dbEST/GenBank, were included in the analyses. The resulting sequences have been clustered into nonredundant gene assemblies and deposited into a relational database that supports a variety of sequence and text searches. This database has been used to compare the gene assemblies using BLAST similarity comparisons to the public protein databases to identify putative genes. Of these new entries, approximately 15%-20% represent putative homologs with a conservative cutoff of p < 10(-9), thus identifying many conserved genes that are likely to share common functions with other well-studied organisms. Gene assemblies were also used to identify strain polymorphisms, examine stage-specific expression, and identify gene families. An interesting class of genes that are confined to members of this phylum and not shared by plants, animals, or fungi, was identified. These genes likely mediate the novel biological features of members of the Apicomplexa and hence offer great potential for biological investigation and as possible therapeutic targets.

Animals↗

Cloning and sequencing of the Dermatophagoides pteronyssinus group III allergen, Der p III.

House dust mites are widely recognized as major factors involved in the triggering of allergic diseases such as asthma. It is now apparent that the group III allergens of the Dermatophagoides mite species may play a significant role in a number of house dust mite allergic cases. Natural Der p III was isolated by gel filtration of salt precipitated Dermatophagoides pteronyssinus extract and as reported previously ran as a doublet of Mr 28 and 30 K on sodium dodecyl sulphate-polyacrylamide gel electrophoresis (SDS-PAGE). Natural Der fIII was isolated by affinity purification with the 5A12 monoclonal antibody. Amino acid sequence data was generated for both these proteins which was used to construct DNA probes to screen a Dermatophagoides pteronyssinus cDNA library by hybridization and resulted in the isolation of a recombinant Der p III cDNA clone, P3WS1. The 1059 bp cDNA fragment included a 786 bp open reading frame which encodes a pre-pro region of 29 amino acids and a mature protein of 232 amino acids with a calculated Mr 24,985. A search of the BLAST protein database has confirmed that the Der pIII P3WS1 clone is approximately 50% homologous with other trypsin proteins. We have confirmed with both our natural protein sequence and the P3WS1 amino acid sequence data that the group III allergens are trypsin-like proteins.

Allergens↗

Identification and characterization of four novel peptide motifs that recognize distinct regions of the transcription factor CP2.

Although ubiquitously expressed, the transcriptional factor CP2 also exhibits some tissue- or stage-specific activation toward certain genes such as globin in red blood cells and interleukin-4 in T helper cells. Because this specificity may be achieved by interaction with other proteins, we screened a peptide display library and identified four consensus motifs in numerous CP2-binding peptides: HXPR, PHL, ASR and PXHXH. Protein-database searching revealed that RE-1 silencing factor (REST), Yin-Yang1 (YY1) and five other proteins have one or two of these CP2-binding motifs. Glutathione S-transferase pull-down and coimmunoprecipitation assays showed that two HXPR motif-containing proteins REST and YY1 indeed were able to bind CP2. Importantly, this binding to CP2 was almost abolished when a double amino acid substitution was made on the HXPR sequence of REST and YY1 proteins. The suppressing effect of YY1 on CP2's transcriptional activity was lost by this point mutation on the HXPR sequence of YY1 and reduced by an HXPR-containing peptide, further supporting the interaction between CP2 and YY1 via the HXPR sequence. Mapping the sites on CP2 for interaction with the four distinct CP2-binding motifs revealed at least three different regions on CP2. This suggests that CP2 recognizes several distinct binding motifs by virtue of employing different regions, thus being able to interact with and regulate many cellular partners.

Amino Acid Sequence↗

Protein expression in a transformed trabecular meshwork cell line: proteome analysis.

PURPOSE: Characterization of the human trabecular meshwork (TM) proteome is hindered by the small mass of intact tissue and the slow growth of cultured cell strains. We have previously characterized a transformed TM cell strain (GTM3) that demonstrates many of the same protein expression and cell signaling systems of nontransformed cell strains. The aim of this study was to initiate a proteomic survey of GTM3 cells as the initial step toward characterization of the complete human TM proteome. METHODS: GTM3 cells were cultured to confluence, harvested and solubilized in urea/Nonidet. The protein extract (600 mug) was focused in immobilized isoelectric focusing (IEF) strips, separated by 10% SDS PAGE, and visualized with colloidal Coomassie Blue. Spots of interest were excised, destained, and the contained proteins subjected to in-gel reduction, derivatization, and tryptic digestion. Tryptic peptides were extracted and analyzed by electrospray LC/MS/MS. Protein identification was made using the TurboSequest search algorithm and a recent version of the nonredundant human protein database downloaded from the National Center for Biotechnology Information (NCBI). RESULTS: Eighty-seven (87) primary proteins and 93 variants of these proteins were identified. A website was created (TM proteome) that combines data such as graphic spot location within the gel, peptide sequence, apparent and calculated pI, apparent and calculated mass, percentage of coverage, and protein informatic website links. CONCLUSIONS: Proteomic analysis of a transformed human TM cell line has been initiated combining preparative two-dimensional PAGE separation, LC/MS/MS analysis of major proteins, and bioinformatic cataloging of the data. Further investigation of data from the transformed cell strain will be used in a comparative fashion for spot identification of analytical proteomic gels of human TM tissue and cultured normal cells. These initial data will form the base from which the characterization of protein expression in the normal and glaucomatous TM can be accomplished.

Cell Line, Transformed↗

The human cornea proteome: bioinformatic analyses indicate import of plasma proteins into the cornea.

Increased biochemical knowledge of normal and diseased corneas is essential for the understanding of corneal homeostasis and pathophysiology. In a recent study, we characterized the proteome of the normal human cornea and identified 141 distinct proteins. This dataset represents the most comprehensive protein study of the cornea to date and provides a useful reference for further studies of normal and diseased human corneas. The list of identified proteins is available at the Cornea Protein Database. In the present paper, we review the utilized procedures for extraction and fractionation of corneal proteins and discuss the potential roles of the identified proteins in relation to homeostasis, diseases, and wound-healing of the cornea. In addition, we compare the list of identified proteins with high quality gene expression libraries (cDNA libraries) and Serial Analysis of Gene Expression (SAGE) data. Of the 141 proteins, 86 (61%) were recognized in cDNA libraries from the corneas of dogs and rabbits, or humans with keratoconus, and 98 (69.5%) were recognized in SAGE data of mouse and human corneas. However, the percentages of identified genes in each of the protein functional groups differed markedly. Thus, exceptionally few of the traditional blood/plasma proteins and immune defense proteins that were identified in the human cornea were recognized in the gene expression libraries of the cornea. This observation strongly indicates that these abundant corneal proteins are not expressed in the cornea but originate from the surrounding pericorneal tissue.

Animals↗

Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) for large-scale protein analysis: the yeast proteome.

Highly complex protein mixtures can be directly analyzed after proteolysis by liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS). In this paper, we have utilized the combination of strong cation exchange (SCX) and reversed-phase (RP) chromatography to achieve two-dimensional separation prior to MS/MS. One milligram of whole yeast protein was proteolyzed and separated by SCX chromatography (2.1 mm i.d.) with fraction collection every minute during an 80-min elution. Eighty fractions were reduced in volume and then re-injected via an autosampler in an automated fashion using a vented-column (100 microm i.d.) approach for RP-LC-MS/MS analysis. More than 162,000 MS/MS spectra were collected with 26,815 matched to yeast peptides (7,537 unique peptides). A total of 1,504 yeast proteins were unambiguously identified in this single analysis. We present a comparison of this experiment with a previously published yeast proteome analysis by Yates and colleagues (Washburn, M. P.; Wolters, D.; Yates, J. R., III. Nat. Biotechnol. 2001, 19, 242-7). In addition, we report an in-depth analysis of the false-positive rates associated with peptide identification using the Sequest algorithm and a reversed yeast protein database. New criteria are proposed to decrease false-positives to less than 1% and to greatly reduce the need for manual interpretation while permitting more proteins to be identified.

Cations↗

Trypanosome apoptotic factor mediates apoptosis in human brain vascular endothelial cells.

Human African trypanosomiasis (HAT, sleeping sickness) is a devastating disease caused by infection with Trypanosoma brucei ssp. These hemoflagellates invade the central nervous system (CNS) and induce meningo-encephalitis, neuronal demyelination, blood-brain-barrier (BBB) dysfunction, peri-vascular infiltration, astrocytosis and apoptosis. The molecular basis of these manifestations is unclear. We previously reported T. brucei-induced apoptosis in cerebella and brain-stem nuclei in mice at peak parasitemia. Here, we identify and characterize a trypanosome apoptotic factor (TAF) expressed by T. brucei that mediates apoptosis in mouse-brain and human-brain vascular endothelial cells (HBVEC). Molecular, biochemical and apoptotic assays, coupled with surface enhanced laser desorption ionization (SELDI), and protein database analyses were utilized to show that TAF is a soluble, non-serum, parasite-derived, heat-labile protein that causes DNA laddering and apoptosis in HBVEC. Protein-chip assay analysis of the SELDI spectrum of infected mouse serum and procyclic culture supernatants revealed a single major peak at 8652.7 Da. Further database analysis indicated that the TAF may be a procyclin or procyclin derivative. A synthetic 27 mer peptide (ProEP2-1), corresponding to a region common to EP procyclins (EP2-1), induced apoptosis in HBVEC and in cerebella of mice similar to that induced in T. brucei-infected mice. Western blot analysis utilizing an anti-procyclin monoclonal antibody (mAb) revealed that TAF is present in infected but not uninfected brain tissue lysates. Furthermore, this mAb blocked T. brucei- and ProEP2-1-induced apoptosis in HBVEC in vitro. We conclude that T. brucei TAF or its derivative(s) play a major role in the apoptosis associated with HAT pathology.

Animals↗

Identification of tryptic peptides from large databases using multiplexed tandem mass spectrometry: simulations and experimental results.

Multiplexed tandem mass spectrometry (MS/MS) has recently been demonstrated as a means to increase the throughput of peptide identification in liquid chromatography (LC) MS/MS experiments. In this approach, a set of parent species is dissociated simultaneously and measured in a single spectrum (in the same manner that a single parent ion is conventionally studied), providing a gain in sensitivity and throughput proportional to the number of species that can be simultaneously addressed. In the present work, simulations performed using the Caenorhabditis elegans predicted proteins database show that multiplexed MS/MS data allow the identification of tryptic peptides from mixtures of up to ten peptides from a single dataset with only three "y" or "b" fragments per peptide and a mass accuracy of 2.5 to 5 ppm. At this level of database and data complexity, 98% of the 500 peptides considered in the simulation were correctly identified. This compares favorably with the rates obtained for classical MS/MS at more modest mass measurement accuracy. LC multiplexed Fourier transform-ion cyclotron resonance MS/MS data obtained from a 66 kDa protein (bovine serum albumin) tryptic digest sample are presented to illustrate the approach, and confirm that peptides can be effectively identified from the C. elegans database to which the protein sequence had been appended.

Algorithms↗

Molecular cloning of developmentally specific genes by representational difference analysis during the fruiting body formation in the basidiomycete Lentinula edodes.

To understand molecular mechanisms of the fruiting body development in basidiomycetes, we attempted to isolate developmentally regulated genes expressed specifically during the fruiting body formation of Lentinula edodes (Shiitake-mushroom). cDNA representational difference analysis (cDNA-RDA) between vegetatively growing mycelium and two developmental substages, primordium and mature fruiting body, resulted in an isolation of 105 individual genes (51 in primordium and 54 in mature fruiting body, respectively). A search of homology with the protein databases and two basidiomycetous genomes in Phanerochaete chrysosporium and Coprinopsis cinerea revealed that the obtained genes encoded various proteins similar to those involved in general metabolism, cell structure, signal transduction, and responses to stress; in addition, there were apparently several metabolic pathways and signal transduction cascades that could be involved in the fruiting body development. The expression products of several genes revealed no significant homologies to those in the databases, implying that those genes are unique in L. edodes and the encoding products may possess possible functions in the course of fruiting body development. RT-PCR analyses revealed that 20 candidates of the obtained genes were specifically or abundantly transcribed in the course of the fruiting body formation, suggesting that the obtained genes in this work play roles in fruiting body development in L. edodes.

Agaricales↗