PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Toward a human blood serum proteome: analysis by multidimensional separation coupled with mass spectrometry.

Blood serum is a complex body fluid that contains various proteins ranging in concentration over at least 9 orders of magnitude. Using a combination of mass spectrometry technologies with improvements in sample preparation, we have performed a proteomic analysis with submilliliter quantities of serum and increased the measurable concentration range for proteins in blood serum beyond previous reports. We have detected 490 proteins in serum by on-line reversed-phase microcapillary liquid chromatography coupled with ion trap mass spectrometry. To perform this analysis, immunoglobulins were removed from serum using protein A/G, and the remaining proteins were digested with trypsin. Resulting peptides were separated by strong cation exchange chromatography into distinct fractions prior to analysis. This separation resulted in a 3-5-fold increase in the number of proteins detected in an individual serum sample. With this increase in the number of proteins identified we have detected some lower abundance serum proteins (ng/ml range) including human growth hormone, interleukin-12, and prostate-specific antigen. We also used SEQUEST to compare different protein databases with and without filtering. This comparison is plotted to allow for a quick visual assessment of different databases as a subjective measure of analytical quality. With this study, we have performed the most extensive analysis of serum proteins to date and laid the foundation for future refinements in the identification of novel protein biomarkers of disease.

Blood Proteins↗

The development of a high-density canine microarray.

DNA microarrays can give global transcriptional views of cellular responses to disease, development, nutrition, and other biological states. They can be used to elucidate biological networks, develop diagnostics, and identify genetic targets and molecular mechanisms. The technology is widely used and can be a valuable complement to more "disease-centric" focused arrays. For these reasons, Nestlé designed a custom canine Affymetrix microarray representing transcripts from multiple tissues for use in areas where a more focused microarray had not already been developed. Sufficient numbers of sequences representing messenger RNAs (mRNAs) or expressed sequence tags (ESTs) is integral for the design of a global microarray chip. This chip was designed using public domain sequences (GenBank) and sequences from a proprietary canine EST database. In order to enrich the chip with annotated transcripts, both of these sequence sets were BLASTed against the nonredundant protein database. The sequences on the microarray were isolated from more than 48 different tissues. The final compliment of sequences had sequences unique to GenBank (3160), unique to the proprietary EST database (17,620), and present in both sources (1996). In comparison with human sequences (RefSeq), 74% of the canine sequences matched a human sequence.

Animals↗

Prediction of proprotein convertase cleavage sites.

Many secretory proteins and peptides are synthesized as inactive precursors that in addition to signal peptide cleavage undergo post-translational processing to become biologically active polypeptides. Precursors are usually cleaved at sites composed of single or paired basic amino acid residues by members of the subtilisin/kexin-like proprotein convertase (PC) family. In mammals, seven members have been identified, with furin being the one first discovered and best characterized. Recently, the involvement of furin in diseases ranging from Alzheimer's disease and cancer to anthrax and Ebola fever has created additional focus on proprotein processing. We have developed a method for prediction of cleavage sites for PCs based on artificial neural networks. Two different types of neural networks have been constructed: a furin-specific network based on experimental results derived from the literature, and a general PC-specific network trained on data from the Swiss-Prot protein database. The method predicts cleavage sites in independent sequences with a sensitivity of 95% for the furin neural network and 62% for the general PC network. The ProP method is made publicly available at http://www.cbs.dtu.dk/services/ProP.

Animals↗

Purification and identification of a tributyltin-binding protein from serum of Japanese flounder, Paralichthys olivaceus.

Tributyltin (TBT) is an industrial chemical used as an antifoulant in marine environments. Previously, we reported that TBT accumulates in the serum or plasma of some fishes and is bound to a high molecular weight compound in the serum of the Japanese flounder, Paralichthys olivaceus. In this study, we succeeded in purifying the TBT-binding protein (TBT-bp) from the serum of Japanese flounder by using gel filtration chromatography, anion exchange chromatography, and polyacrylamide gel electrophoresis, with a 2.6% yield and a 77-fold purification. The molecular mass of TBT-bp was approximately 46.5 kDa on sodium dodecyl sulfate-polyacrylamide gel electrophoresis, and its isoelectric point was approximately 3.0 on isoelectric focusing-polyacrylamide gel electrophoresis. The TBT-bp contained 42% N-glycan. The cDNA nucleotide sequence of TBT-bp was determined by reverse transcription-polymerase chain reaction of Japanese flounder liver, and we deduced a sequence of 191 amino acids of mature TBT-bp. No sequence identical to the TBT-bp amino acid sequence was found within the SWISS-PROT (http://www.nig. ac.jp/) protein database; however, a lipocalin-like sequence pattern was observed. We concluded that the TBT-bp was a novel protein that has not yet been reported, although some DNA sequences from expressed sequence tags (ESTs) of Japanese flounder liver had a high identity. A high expression level of TBT-bp gene was found in the liver, but the gene was slightly detectable in the kidney and brain.

Amino Acid Sequence↗

[Proteomic analysis of proteins related to retinoid acid resistance].

OBJECTIVE: To evaluate the applicability of proteomic methods for shedding light on the mechanisms of All-trans retinoic acid resistance. METHODS: The expression of cellular proteins in the retinoid acid sensitive cell line NB4 and retinoid acid resistant cell line MR2 was analyzed using the two-dimensional polyacrylamide gel electrophoresis (2D-PAGE). Differentially expressed proteins were analyzed by mass spectrometry for peptide mass finger and identified by SWISS-PROT protein database. RESULTS: Approximately 600 spots appeared in an sodium dod(SDS-GEL). The match of the expression of cellular proteins between NB4 and MR2 was better. Three significantly differentially expressed proteins were screened and identified to DJ-1 and HSP70, as well as KIAA1289 of which the function is not yet known. CONCLUSION: The utilization of 2D-PAGE, coupled with mass spectrometry is effective in screening the drug resistance-associated proteins and could provide molecular targets for the elucidation of resistance mechanisms.

Antineoplastic Agents↗

Altered mouse leukemia L1210 thymidylate synthase, associated with cell resistance to 5-fluoro-dUrd, is not mutated but rather reflects posttranslational modification.

Thymidylate synthase purified from 5-fluoro-dUrd-resistant mouse leukemia L1210 cells (TSr) was less sensitive to slow-binding inhibition by 5-fluoro-dUMP than the enzyme from the parental cells (TSp), both enzyme forms differing also in sensitivity to several other dump analogues, apparent molecular weights of monomer and dimer, and temperature dependence of the catalyzed reaction. Direct sequencing of products obtained from RT-PCR, performed on total RNA isolated from the parental and 5-fluoro-dUrd-resistant cells, proved both nucleotide sequences to be identical to the mouse thymidylate synthase coding sequence published earlier (NCBI protein database access no. NP_067263). This suggests that the altered properties of TSr are caused by a factor different than protein mutation, presumably posttranslational modification. As a possibility of rat thymidylate synthase phosphorylation has been recently demonstrated (Samsonoff et al. (1997) J Biol Chem 272: 13281), the mouse enzyme amino-acid sequence was analysed, revealing several potential phosphorylation sites. In order to test possible influence of the protein phosphorylation state on enzymatic properties, endogenous TSp and TSr were purified in the presence of inhibitors of phosphatases. Although both enzyme forms were phosphorylated, as shown by electrophoretical separation followed by phosphoprotein detection, the extent of phosphorylation was apparently similar. However, the same two purified enzyme preparations, compared to the corresponding preparations purified in the absence of phosphatase inhibitors, showed certain properties, including sensitivity to the slow-binding inhibition by FdUMP, altered. Thus properties dependence on phosphorylation was indicated.

Animals↗

Reference points for comparisons of two-dimensional maps of proteins from different human cell types defined in a pH scale where isoelectric points correlate with polypeptide compositions.

A highly reproducible, commercial and nonlinear, wide-range immobilized pH gradient (IPG) was used to generate two-dimensional (2-D) gel maps of [35S]methionine-labeled proteins from noncultured, unfractionated normal human epidermal keratinocytes. Forty one proteins, common to most human cell types and recorded in the human keratinocyte 2-D gel protein database were identified in the 2-D gel maps and their isoelectric points (pI) were determined using narrow-range IPGs. The latter established a pH scale that allowed comparisons between 2-D gel maps generated either with other IPGs in the first dimension or with different human protein samples. Of the 41 proteins identified, a subset of 18 was defined as suitable to evaluate the correlation between calculated and experimental pI values for polypeptides with known composition. The variance calculated for the discrepancies between calculated and experimental pI values for these proteins was 0.001 pH units. Comparison of the values by the t-test for dependent samples (paired test) gave a p-level of 0.49, indicating that there is no significant difference between the calculated and experimental pI values. The precision of the calculated values depended on the buffer capacity of the proteins, and on average, it improved with increased buffer capacity. As shown here, the widely available information on protein sequences cannot, a priori, be assumed to be sufficient for calculating pI values because post-translational modifications, in particular N-terminal blockage, pose a major problem. Of the 36 proteins analyzed in this study, 18-20 were found to be N-terminally blocked and of these only 6 were indicated as such in databases. The probability of N-terminal blockage depended on the nature of the N-terminal group. Twenty six of the proteins had either M, S or A as N-terminal amino acids and of these 17-19 were blocked. Only 1 in 10 proteins containing other N-terminal groups were blocked.

Amino Acid Sequence↗

Peptide-mass fingerprinting and the ideal covering set for protein characterisation.

The rules that govern the dynamics of protein characterisation by peptide-mass fingerprinting (PMF) were investigated through multiple interrogations of a nonredundant protein database. This was achieved by analysing the efficiency of identifying each entry in the entire database via perfect in silico digestion with a series of 20 pseudo-endoproteinases cutting at the carboxy terminal of each amino acid residue, and the multiple cutters: trypsin, chymotrypsin and Glu-C. The distribution of peptide fragment masses generated by endoproteinase digestion was examined with a view to designing better approaches to protein characterisation by PMF. On average, and for both common and rare cutters, the combination of approximately two fragments was sufficient to identify most database entries. However, the rare cutters left more entries unidentified in the database. Total coverage of the entire database could not be achieved with one enzymatic cutter alone, nor when all 23 cutters were used together. Peptide fragments of > 5000 Da had little effect on the outcome of PMF to correctly characterise database entries, while those with low mass (near to 350 Da in the case of trypsin) were found to be of most utility. The most frequently occurring fragments were also found in this lower mass region. The maximum size of uncut database entries (those not containing a specific amino acid residue) ranged from 52,908 Da to 258,314 Da, while the failure rate for a single cutter in identifying database entries varied from 10,865 (8.4%) to 23,290 (18.1%). PMF is likely to be a mainstay of any high-throughput protein screening strategy for large-scale proteome analysis. A better understanding of the merits and limitations of this technique will allow researchers to optimise their protein characterisation procedures.

Amino Acid Sequence↗

PPD v1.0--an integrated, web-accessible database of experimentally determined protein pKa values.

The Protein pK(a) Database (PPD) v1.0 provides a compendium of protein residue-specific ionization equilibria (pK(a) values), as collated from the primary literature, in the form of a web-accessible postgreSQL relational database. Ionizable residues play key roles in the molecular mechanisms that underlie many biological phenomena, including protein folding and enzyme catalysis. The PPD serves as a general protein pK(a) archive and as a source of data that allows for the development and improvement of pK(a) prediction systems. The database is accessed through an HTML interface, which offers two fast, efficient search methods: an amino acid-based query and a Basic Local Alignment Search Tool search. Entries also give details of experimental techniques and links to other key databases, such as National Center for Biotechnology Information and the Protein Data Bank, providing the user with considerable background information. The database can be found at the following URL: http://www.jenner.ac.uk/PPD.

Amino Acids↗

Local control of peptide conformation: stabilization of cis proline peptide bonds by aromatic proline interactions.

In the native state of proteins there is a marked tendency for an aromatic amino acid to precede a cis proline. There are also significant differences between the three aromatic amino acids with Tyr exhibiting a noticeably higher propensity than Phe or Trp to precede a cis proline residue. In order to study the role that local interactions play in these conformation preferences, a set of tetrapeptides of the general sequence acetyl-Gly-X-Pro-Gly-carboxamide (GXPG), where X = Tyr, Phe, Trp, Ala, or cyclohexyl alanine, were synthesized and studied by nmr. Analysis of the nmr data shows that none of the peptides adopt a specific backbone structure. Ring current shifts, the equilibrium constant, the Van't Hoff enthalpy, and the measured rate of cis-trans isomerization all indicate that the cis proline conformer is stabilized by favorable interactions between the aromatic ring and the proline residue. Analysis of the side chain conformation of the aromatic residue and analysis of the chemical shifts of the pyrrolidine ring protons shows that the aromatic side chain adopts a preferred conformation in the cis form. The distribution of rotamers and the effect of an aromatic residue on the cis-trans equilibrium indicate that the preferred conformation is populated to approximately 62% for the Phe containing peptide, 67% for the Tyr containing peptide, and between 75 and 80% for the Trp containing peptide. The interaction is unaffected by the addition of 8M urea. These local interactions favor an aromatic residue immediately preceding a cis proline, but they cannot explain the relative propensities for Phe-Pro, Tyr-Pro, and Trp-Pro cis peptide bonds observed in the native state of proteins. In the model peptides the percentage of the cis proline conformer is 21% GYPG while it is 17% for GFPG. This difference is considerably smaller than the almost three to one preponderance observed for cis Tyr-Pro peptide bonds vs cis Phe-Pro peptide bonds in the protein database.

Deuterium↗

Two-dimensional electrophoresis analysis of human serum proteins during the acute-phase response.

The serum of patients with meningitis, due to infection by Haemophilus influenzae type b, was analyzed. Several known acute-phase proteins were separated by two-dimensional electrophoresis and estimated quantitatively. In addition, hitherto undescribed reactants were recognized. Gels were calibrated and relevant spots related to master spot numbers in the human serum protein database.

Acute-Phase Proteins↗

TAPASIN, DAXX, RGL2, HKE2 and four new genes (BING 1, 3 to 5) form a dense cluster at the centromeric end of the MHC.

TAPASIN, a gene recently shown to be required for antigen presentation through MHC class I molecules, is located 180 kbp centromeric of HLA-DP in a region linked to several diseases, and associated with altered developmental phenotypes in the mouse. We present the genomic analysis of a 70 kbp gene-dense segment flanking the TAPASIN locus, including sequence, structure and preliminary characterisation of seven additional genes. BING1 is a Zn finger gene containing a POZ motif. BING3 is similar to myosin regulatory light chain. BING4 shows homologies only to hypothetical yeast and Caenorhabditis elegans proteins. BING5 is found within an intron of BING4 on the complementary strand, and encodes a molecule with no homologies to database proteins. Another three genes were identified whose full sequence was not previously known; namely, RGL2, DAXX (BING2) and HKE2. RGL2 encodes an effector of Ras, homologous to the mouse RalGDS protein, Rlf. DAXX encodes an effector of Fas that stimulates apoptosis through the Jun kinase (JNK) pathway. The location of DAXX is of interest given the linkage of autoimmune disease to the MHC and to apoptosis.

Adaptor Proteins, Signal Transducing↗

Gene discovery in Plasmodium vivax through sequencing of ESTs from mixed blood stages.

Despite the significance of Plasmodium vivax as the most widespread human malaria parasite and a major public health problem, gene expression in this parasite is poorly understood. To accelerate gene discovery and facilitate the annotation phase of the P. vivax genome project, we have undertaken a transcriptome approach to study gene expression in the mixed blood stages of a P. vivax field isolate. Using a cDNA library constructed from purified blood stages, we have obtained single-pass sequences for approximately 21,500 expressed sequence tags (ESTs), the largest number of transcript tags obtained so far for this species. Cluster analysis revealed that the library is highly redundant, resulting in 5407 clusters. Clustered ESTs were searched against public protein databases for functional annotation, and more than one-third showed a significant match, the majority of these to Plasmodium falciparum proteins. The most abundant clusters were to genes encoding ribosomal proteins and proteins involved in metabolism, consistent with the predominance of trophozoites in the field isolate sample. In spite of the scarcity of other parasite stages in the field isolate, we could identify genes that are expressed in rings, schizonts and gametocytes. This study should facilitate our understanding of the gene expression in P. vivax asexual stages and provide valuable data for gene prediction and annotation of the P. vivax genome sequence.

Animals↗

The PEPR GeneChip data warehouse, and implementation of a dynamic time series query tool (SGQT) with graphical interface.

Publicly accessible DNA databases (genome browsers) are rapidly accelerating post-genomic research (see http://www.genome.ucsc.edu/), with integrated genomic DNA, gene structure, EST/ splicing and cross-species ortholog data. DNA databases have relatively low dimensionality; the genome is a linear code that anchors all associated data. In contrast, RNA expression and protein databases need to be able to handle very high dimensional data, with time, tissue, cell type and genes, as interrelated variables. The high dimensionality of microarray expression profile data, and the lack of a standard experimental platform have complicated the development of web-accessible databases and analytical tools. We have designed and implemented a public resource of expression profile data containing 1024 human, mouse and rat Affymetrix GeneChip expression profiles, generated in the same laboratory, and subject to the same quality and procedural controls (Public Expression Profiling Resource; PEPR). Our Oracle-based PEPR data warehouse includes a novel time series query analysis tool (SGQT), enabling dynamic generation of graphs and spreadsheets showing the action of any transcript of interest over time. In this report, we demonstrate the utility of this tool using a 27 time point, in vivo muscle regeneration series. This data warehouse and associated analysis tools provides access to multidimensional microarray data through web-based interfaces, both for download of all types of raw data for independent analysis, and also for straightforward gene-based queries. Planned implementations of PEPR will include web-based remote entry of projects adhering to quality control and standard operating procedure (QC/SOP) criteria, and automated output of alternative probe set algorithms for each project (see http://microarray.cnmcresearch.org/pgadatatable.asp).

Algorithms↗

A search method for homologs of small proteins. Ubiquitin-like proteins in prokaryotic cells?

The question of protein homology versus analogy arises when proteins share a common function or a common structural fold without any statistically significant amino acid sequence similarity. Even though two or more proteins do not have similar sequences but share a common fold and the same or closely related function, they are assumed to be homologs, descendant from a common ancestor. The problem of homolog identification is compounded in the case of proteins of 100 or less amino acids. This is due to a limited number of basic single domain folds and to a likelihood of identifying by chance sequence similarity. The latter arises from two conditions: first, any search of the currently very large protein database is likely to identify short regions of chance match; secondly, a direct sequence comparison among a small set of short proteins sharing a similar fold can detect many similar patterns of hydrophobicity even if proteins do not descend from a common ancestor. In an effort to identify distant homologs of the many ubiquitin proteins, we have developed a combined structure and sequence similarity approach that attempts to overcome the above limitations of homolog identification. This approach results in the identification of 90 probable ubiquitin-related proteins, including examples from the two prokaryotic domains of life, Archaea and Bacteria.

Amino Acid Sequence↗

Sequence analyses of Thogoto viral RNA segment 3: evidence for a distant relationship between an arbovirus and members of the Orthomyxoviridae.

The genome of Thogoto (THO) virus, an unclassified tick-borne virus, comprises six segments of single-stranded RNA. The complete sequence of the third largest RNA segment has been determined from overlapping cDNA clones and by primer extension studies. Segment 3 RNA consists of 1865 nucleotides (approx. 6.2 x 10(5) Mr). It has a large open reading frame (ORF1;597 amino acids, 68.6K) in its virus-complementary sequence, confirming that the RNA has a negative-sense coding strategy. A transcription termination (polyadenylation) site located after the end of ORF1 has been identified. A second ORF (ORF2;98 amino acids in length), overlapping ORF1, is also present in the virus-complementary sequence although whether it is translated is not known. The 3' and 5' sequences of the segment 3 RNA are complementary and similar to those of the tick-borne Dhori (DHO) and the mammalian and avian influenza viruses. Protein database searches have identified regions of homology between the sequence of the THO ORF1 gene product and regions of the PA protein of influenza virus strain A/NT/60/68 (approx. 20% aligned homology) and the corresponding protein of influenza B/Sing/222/79 virus (approx. 15% aligned homology). Although the THO protein sequence is not as closely related to those of the influenza viruses as they are to each other (40% aligned homology), the indicated sequence data provide further evidence of relationships between the tick-borne THO and DHO viruses and the vertebrate orthomyxoviruses.

Amino Acid Sequence↗

BLAST 2 Sequences, a new tool for comparing protein and nucleotide sequences.

'BLAST 2 Sequences', a new BLAST-based tool for aligning two protein or nucleotide sequences, is described. While the standard BLAST program is widely used to search for homologous sequences in nucleotide and protein databases, one often needs to compare only two sequences that are already known to be homologous, coming from related species or, e.g. different isolates of the same virus. In such cases searching the entire database would be unnecessarily time-consuming. 'BLAST 2 Sequences' utilizes the BLAST algorithm for pairwise DNA-DNA or protein-protein sequence comparison. A World Wide Web version of the program can be used interactively at the NCBI WWW site (http://www.ncbi.nlm.nih.gov/gorf/bl2.++ +html). The resulting alignments are presented in both graphical and text form. The variants of the program for PC (Windows), Mac and several UNIX-based platforms can be downloaded from the NCBI FTP site (ftp://ncbi.nlm.nih.gov).

Algorithms↗

Major proteins in normal human lymphocyte subpopulations separated by fluorescence-activated cell sorting and analyzed by two-dimensional gel electrophoresis.

We have compared the overall patterns of protein synthesis of normal human lymphocyte subpopulations taken from five volunteers using high resolution two-dimensional gel electrophoresis. The lymphocytes were isolated using density gradient centrifugation, labeled with subtype-specific MoAbs, and separated to a high degree of homogeneity by FACS into CD4+ helper T cells, CD8+ suppressor T cells, CD20+ B cells, and N901 (NHK-1)+ NK cells. The four lymphocyte subpopulations were labeled with [35S]methionine for 14 hr, solubilized in lysis buffer, and analyzed by two-dimensional gel electrophoresis (IEF). Of about 1000 proteins resolved in each case, most were found to be common to all subpopulations. However, eight putative markers for B1+ (proteins 5525, Mr = 63,700; 5621, Mr = 63,700; 8311, Mr = 36,900; 2202, Mr = 36,300; 6121, Mr = 30,300; 106, Mr = 29,300; 5009, Mr = 23,000; 8012, Mr = 11,600) and one for N901+ (protein 8129, Mr = 30,400) were identified. In contrast, no major protein markers were found that could differentiate T4+ and T8+ cells from each other or from B cells and NK cells. With the exception of two B1+ markers (proteins 5525 and 5621), lower but variable levels of the other markers were observed in all cell types. All the putative protein markers have been identified in the protein database of human peripheral blood mononuclear cells (PBMCs) (see accompanying article by Celis et al.). Comparison of the overall patterns of protein synthesis of the unsorted PBMCs with those of the four subpopulations showed that the synthesis of some major PBMC proteins decreased substantially in the sorted subsets. These proteins are most likely not of monocyte origin, as these cells constituted only about 15% of the total PBMCs. Also, the inhibition does not seem to be due to the addition of the single MoAbs or to cell cycle differences. Taken together, the data provide a background for further studies of protein profiles in normal (resting or activated) and malignant hematopoietic cells.

Antigens, Differentiation↗