PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

A new method to improve sensitivity and resolution in matrix-assisted laser desorption/ionization time of flight mass spectrometry.

High detection sensitivity and resolution are two critical parameters for recording good peptide mass fingerprints (PMF) of low abundance proteins. This paper reports a mass spectrometry (MS) sample preparation technique that could improve sensitivity and resolution. By coating the MS steel target with a thin layer of pentadecafluorooctamido propyltrimethoxysilane, which was both polar and nonpolar solvent repellent, the transferred sample droplets on its surface were significantly smaller. As a result, the analyte of the peptide mixture became more concentrated and homogeneous, which helped to improve the sensitivity. The advantages of a modified MS target were documented by mass spectra improvement of attomole level standard peptides and silver-stained proteins from polyacrylamide gels. The mass signal of angiotensin II at 100 attomole was difficult to record on the conventional support, whereas it was easily detected on the modified one. The PMF of cytochrome C was also better recorded on the modified support, in terms of both signal-to-noise ratio and the number of detected peptides. When silver-stained proteins from two-dimensional electrophoresis gels were analyzed, in most cases more satisfactory peptide mass spectra were obtained from the modified support. Searching protein databases with more mass data from the improved PMFs, several unknown proteins were successfully identified.

Angiotensin II↗

Analysis of the human lumbar cerebrospinal fluid proteome.

Idiopathic low back pain has no known cause, and the molecular basis is unknown. Neuropeptidergic systems have been previously studied, and proteomics methods have been applied in this present study. Proteomics combines high-resolution two-dimensional (2-D) gel electrophoresis, high-sensitivity mass spectrometry, and continuously expanding protein databases. Proteomics offers a comprehensive, bird's-eye view to analyze, at a systems level, all of the proteins in cerebrospinal fluid (CSF) that might contribute to idiopathic low back pain. CSF contains a high salt concentration and low protein concentration. In order to obtain a high-quality 2-D pattern, several sample preparation methods were tested to remove salts - protein precipitation with either acetone or trichloroacetic acid/acetone, or sample treatment with a Bio-Spin column. More spots were visualized on the 2-D gel of human CSF, and a relatively high protein recovery was obtained when a Bio-Spin column was used to process a human CSF sample. Sixty-one protein spots, obtained from 2-D gels with a pH range of either 3-10 or 4-7, were identified by matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS) and MALDI-post-source decay (PSD)-MS. These 61 protein spots represent 22 proteins; six of those proteins were not annotated in any previously published 2-D maps. Those six proteins are PRO2619, pigment epithelium-derived factor, albumin homolog, kallikrein-6 precursor, DJ717I23.1, and AMBP protein precursor. These protein-mapping data will contribute to the database that will be used in the future to compare the proteomes obtained from the CSF of controls and low back pain patients, to characterize differentially expressed proteins, and to elucidate the biological markers for idiopathic low back pain.

Amino Acid Sequence↗

Defining the fold space of membrane proteins: the CAMPS database.

Recent progress in structure determination techniques has led to a significant growth in the number of known membrane protein structures, and the first structural genomics projects focusing on membrane proteins have been initiated, warranting an investigation of appropriate bioinformatics strategies for optimal structural target selection for these molecules. What determines a membrane protein fold? How many membrane structures need to be solved to provide sufficient structural coverage of the membrane protein sequence space? We present the CAMPS database (Computational Analysis of the Membrane Protein Space) containing almost 45,000 proteins with three or more predicted transmembrane helices (TMH) from 120 bacterial species. This large set of membrane proteins was subjected to single-linkage clustering using only sequence alignments covering at least 40% of the TMH present in a given family. This process yielded 266 sequence clusters with at least 15 members, roughly corresponding to membrane structural folds, sufficiently structurally homogeneous in terms of the variation of TMH number between individual sequences. These clusters were further subdivided into functionally homogeneous subclusters according to the COG (Clusters of Orthologous Groups) system as well as more stringently defined families sharing at least 30% identity. The CAMPS sequence clusters are thus designed to reflect three main levels of interest for structural genomics: fold, function, and modeling distance. We present a library of Hidden Markov Models (HMM) derived from sequence alignments of TMH at these three levels of sequence similarity. Given that 24 out of 266 clusters corresponding to membrane folds already have associated known structures, we estimate that 242 additional new structures, one for each remaining cluster, would provide structural coverage at the fold level of roughly 70% of prokaryotic membrane proteins belonging to the currently most populated families.

Bacterial Proteins↗

The multiassembly problem: reconstructing multiple transcript isoforms from EST fragment mixtures.

Recent evidence of abundant transcript variation (e.g., alternative splicing, alternative initiation, alternative polyadenylation) in complex genomes indicates that cataloging the complete set of transcripts from an organism is an important project. One challenge is the fact that most high-throughput experimental methods for characterizing transcripts (such as EST sequencing) give highly detailed information about short fragments of transcripts or protein products, instead of a complete characterization of a full-length form. We analyze this "multiassembly problem"-reconstructing the most likely set of full-length isoform sequences from a mixture of EST fragment data-and present a graph-based algorithm for solving it. In a variety of tests, we demonstrate that this algorithm deals appropriately with coupling of distinct alternative splicing events, increasing fragmentation of the input data and different types of transcript variation (such as alternative splicing, initiation, polyadenylation, and intron retention). To test the method's performance on pure fragment (EST) data, we removed all mRNA sequences, and found it produced no errors in 40 cases tested. Using this algorithm, we have constructed an Alternatively Spliced Proteins database (ASP) from analysis of human expressed and genomic sequences, consisting of 13,384 protein isoforms of 4422 genes, yielding an average of 3.0 protein isoforms per gene.

Algorithms↗

What is the minimum number of letters required to fold a protein?

Experimental studies have shown that the full sequence complexity of naturally occurring proteins is not required to generate rapidly folding and functional proteins, i.e. proteins can be designed with fewer than 20 letters. This raises the question of what is the minimum number of amino acid types required to encode complex protein folds? Here, we investigate this issue from three aspects. First, we study the minimum sequence complexity that can reserve the necessary structural information for detection of distantly related homologues. Second, we compare the ability of designing foldable model sequences over a wide range of reduced amino acid alphabets, which find the minimum number of letters that have the similar design ability as 20. Finally, we survey the lower bound of alphabet size of globular proteins in a non-redundant protein database. These different approaches give a remarkably consistent view, that the minimum number of letters required to fold a protein is around ten.

Amino Acids↗

Searching the Porphyromonas gingivalis genome with peptide fragmentation mass spectra.

An approach is described for genomic database searching based on experimentally observed proteolytic fragments, e.g., isolated from 1D or 2D gels or analyzed directly, that can be applied to unfinished prokaryotic genomic data in the absence of annotations or previously assigned open reading frames (ORFs). This variation on the database search is in contrast to the more familiar use of peptide mass spectral fragmentation data to search fully annotated inferred protein databases, e.g., OWL or SWISS-PROT. We compared the SEQUEST search results from a six reading frame translation of the Porphyromonas gingivalis genome DNA sequence with those from computationally derived ORFs created using publicly available genomics software tools. The ORF approach eliminated many of the artifacts present in output from the six reading frame search. The method was applied to uninterpreted tandem mass spectrometric data derived from proteins secreted by the periodontal pathogen Porphyromonas gingivalis in response to the gingival epithelial cell environment, a model system for the study of host-pathogen interactions relevant to human periodontal disease.

Bacterial Proteins↗

Molecular modeling of phosphorylation sites in proteins using a database of local structure segments.

A new bioinformatics tool for molecular modeling of the local structure around phosphorylation sites in proteins has been developed. Our method is based on a library of short sequence and structure motifs. The basic structural elements to be predicted are local structure segments (LSSs). This enables us to avoid the problem of non-exact local description of structures, caused by either diversity in the structural context, or uncertainties in prediction methods. We have developed a library of LSSs and a profile--profile-matching algorithm that predicts local structures of proteins from their sequence information. Our fragment library prediction method is publicly available on a server (FRAGlib), at http://ffas.ljcrf.edu/Servers/frag.html . The algorithm has been applied successfully to the characterization of local structure around phosphorylation sites in proteins. Our computational predictions of sequence and structure preferences around phosphorylated residues have been confirmed by phosphorylation experiments for PKA and PKC kinases. The quality of predictions has been evaluated with several independent statistical tests. We have observed a significant improvement in the accuracy of predictions by incorporating structural information into the description of the neighborhood of the phosphorylated site. Our results strongly suggest that sequence information ought to be supplemented with additional structural context information (predicted with our segment similarity method) for more successful predictions of phosphorylation sites in proteins.

Amino Acid Sequence↗

Characterization of cysteine residues and disulfide bonds in proteins by liquid chromatography/electrospray ionization tandem mass spectrometry.

Cysteine residues and disulfide bonds are important for protein structure and function. We have developed a simple and sensitive method for determining the presence of free cysteine (Cys) residues and disulfide bonded Cys residues in proteins (<100 pmol) by liquid chromatography/electrospray ionization tandem mass spectrometry (LC/ESI-MS/MS) in combination with protein database searching using the program Sequest. Free Cys residues in a protein were labeled with PEO-maleimide biotin immediately followed by denaturation with 8 M urea. Subsequently, the protein was digested with trypsin or chymotrypsin and the resulting products were analyzed by capillary LC/ESI-MS/MS for peptides containing modified Cys and/or disulfide bonded Cys residues. Although the MS method for identifying disulfide bonds has been routinely employed, methods to prevent thiol-disulfide exchange have not been well documented. Our protocol was found to minimize the occurrence of the thiol-disulfide exchange reaction. The method was validated using well-characterized proteins such as aldolase, ovalbumin, and beta-lactoglobulin A. We also applied this method to characterize Cys residues and disulfide bonds of beta 1,4-galactosyltransferase (five Cys), and human blood group A and B glycosyltransferases (four Cys). Our results demonstrate that beta 1,4-galactosyltransferase contains one free Cys residue and two disulfide bonds, which is in contrast to work previously reported using chemical methods for the characterization of free Cys residues, but is consistent with recently published results from x-ray crystallography. In contrast to the results obtained for beta 1,4-galactosyltransferase, none of the Cys residues in A and B glycosyltransferases were found to be involved in disulfide bonds.

Amino Acid Sequence↗

Analysis of expressed sequence tags from the anamorphic basidiomycetous yeast, Pseudozyma antarctica, which produces glycolipid biosurfactants, mannosylerythritol lipids.

Pseudozyma antarctica T-34 secretes a large amount of biosurfactants (BS), mannosylerythritol lipids (MEL), from different carbon sources such as hydrocarbons and vegetable oils. The detailed biosynthetic pathway of MEL remained unknown due to lack of genetic information on the anamorphic basidiomycetous yeasts, including the genus Pseudozyma. Here, in order to obtain genetic information on P. antarctica T-34, we constructed a cDNA library from yeast cells producing MEL from soybean oil and identified the genes expressed through the creation of an expressed sequence tags (EST) library. We generated 398 ESTs, assembled into 146 contiguous sequences. Based upon a BLAST search similarity cut-off of E<or=10(-5), 21.4% of all contigs were orphan, while 78.6% showed similarity to sequences in the protein database; 60.3% of all contiguous sequences shared significant identities to hypothetical protein of Ustilago maydis, which is a smut fungus and BS producer. Based on the gene expression study using real-time reverse transcriptase-PCR, the predicted genes, such as mannosyltranferase and acyltransferase, were demonstrated to be highly involved in MEL biosynthesis in soybean oil-grown cells.

Acyltransferases↗

Comparative proteomic analysis with postmortem prefrontal cortex tissues of suicide victims versus controls.

BACKGROUND: The origin of suicidal behaviour is multifactorial including genetic, neurobiological and psychosocial correlates. Although there is no doubt that serotonin has a central role, the overall genetic findings with candidate genes of the serotonergic pathway are relatively inconsistent and suggests that other, yet unidentified, genes and gene products are also contributing to the vulnerability of suicidality. Proteomics is a powerful method to investigate modifications in protein expression. METHODS: We performed comparative proteomic analysis with prefrontal cortex tissues of 17 suicide victims and 9 controls. RESULTS: Applying two dimensional gel electrophoresis and image analysis we detected five protein spots to differ significantly in intensities between both groups. Three of them appeared only in suicide victims and could be identified by means of MALDI-TOF-MS analysis and protein database search as alpha crystallin chain B (CRYAB), glial fibrillary acidic protein (GFAP) and manganese superoxide dismutase (SOD2). CRYAB belongs to the low molecular heat shock proteins and GFAP is known as a marker of astrocytic activation in gliosis. SOD2 is a major antioxidant enzyme protecting cells against oxidative injury. Two further spots revealed higher intensities in the control group but had no unambiguous protein to match. CONCLUSIONS: Our findings suggest that proteins, being involved in glial function, neurodegeneration and oxidative stress neuronal injury, might also have an impact upon the neurobiological cascade leading to suicidality. As animal data provide evidence for an up-regulation of GFAP synthesis in astrocytes due to alterations in 5-HT levels, similar mechanisms of interaction might also be relevant in humans.

Adolescent↗

Classification of 29 families of secondary transport proteins into a single structural class using hydropathy profile analysis.

A classification scheme for membrane proteins is proposed that clusters families of proteins into structural classes based on hydropathy profile analysis. The averaged hydropathy profiles of protein families are taken as fingerprints of the 3D structure of the proteins and, therefore, are able to detect more distant evolutionary relationships than amino acid sequences. A procedure was developed in which hydropathy profile analysis is used initially as a filter in a BLAST search of the NCBI protein database. The strength of the procedure is demonstrated by the classification of 29 families of secondary transporters into a single structural class, termed ST[3]. An exhaustive search of the database revealed that the 29 families contain 568 unique sequences. The proteins are predominantly from prokaryotic origin and most of the characterized transporters in ST[3] transport organic and inorganic anions and a smaller number are Na(+)/H(+) antiporters. All modes of energy coupling (symport, antiport, uniport) are found in structural class ST[3]. The relevance of the classification for structure/function prediction of uncharacterised transporters in the class is discussed.

Membrane Transport Proteins↗

Protein expression signatures: an application of proteomics.

The methods of proteomics, the study of the protein complement of the genome, are applicable to environmental testing. Sets of proteins specific to different stressors can be isolated using computer imaging software. Individual proteins can be identified by mass spectrometry. The Protein Expression Signatures (PES) obtained have potential in diagnosing adverse factors in the environment. The challenge is to demonstrate their feasibility in complex environments. We have shown that PES for three endocrine disrupting compounds in trout (Onchorhynchus mykiss), can be detected in mixed sewage effluent. Other studies support these results. As protein databases expand, identification becomes routine, and capture molecules specific to each protein are developed, the possibility of simple field tests for multiple stressors becomes real.

Animals↗

Bovine herpesvirus 4: genomic organization and relationship with two other gammaherpesviruses, Epstein-Barr virus and herpesvirus saimiri.

Bovine herpesvirus 4 (BHV-4) belongs to the gammaherpesvirinae subfamily. Although the whole sequence of BHV-4 genome is not known it was possible, based on random sequencing, to assume that its genomic organization consists of genes clustered in blocks whose orientation and location in the genome are conserved within a herpesvirus subfamily. Between these blocks lie genes which are specific to either a particular virus or a virus subfamily. BHV-4 genome consists of 5 gene blocks conserved among the gammaherpesviruses and particularly within the Epstein-Barr virus (EBV) and the herpesvirus saimiri (HVS) genomes. Analysis of the regions located outside the gene blocks showed the presence of 12 open reading frames (ORFs). Protein database comparisons showed that no ORF translation products were similar to proteins encoded by alpha- or beta-herpesviruses. Nevertheless, 5 ORFs were homologous in amino acid sequences to proteins encoded by HVS and one was similar to a protein encoded by both HVS and EBV. On the basis of the molecular data BHV-4 is more closely related to HVS than to EBV. Genes homologous to cellular genes have been described in both HVS and EBV genomes. No genes homologous to presently sequenced cellular genes were found among those found in the BHV-4 genome to date.

Animals↗

Characterization of the major proteins of tubers of yam bean (Pachyrhizus ahipa).

Tubers of six accessions of ahipa (Pachyrhizus ahipa) contained between 0.77 and 1.34% nitrogen on a dry weight basis. This corresponds to 4.8 to 8.4% crude protein based on a nitrogen to protein conversion factor of 6.25; but detailed analysis of AC230 showed that although 93% of the total N was extracted with buffer containing 1.0 M NaCl, about a third of this was lost on dialysis. It was calculated, therefore, that salt-soluble proteins comprise about 60% of the total tuber nitrogen, with low-molecular-mass nitrogenous components comprising a further 30%. Electophoretic analysis of the salt-soluble proteins showed similar patterns of components in the six accessions, with none being present in amounts sufficiently high to suggest a role as storage proteins. Furthermore, light microscopy failed to show significant deposits of protein within the tuber cells. Five "major" protein bands, which together accounted for about 19% of the total salt-soluble protein fraction were purified and subjected to N-terminal amino acid sequencing. Comparison of these with sequences in protein databases revealed similarities to alpha-amylases, chitinases and chitin binding proteins, cysteine proteinases (including major components from P. erosus tubers), a tuberization-specific protein from potato, and proteins induced in soybean and pea by stress or the plant hormone abscisic acid, respectively. It was concluded that the primary roles of these proteins are probably in aspects of tuber metabolism and development and/or conferring protection to pests and pathogens, and that true storage proteins are not present. The absence of storage proteins is consistent with the biological role of the tubers as storage organs for carbohydrates (cf cassava tuberous roots) rather than as propagules (cf yam and potato tubers).

Amino Acid Sequence↗

Characterization of somatic embryogenesis-related cDNAs from alfalfa (Medicago sativa L.).

Messenger RNAs from cultures of embryogenic and non-embryogenic alfalfa (Medicago sativa L.) genotypes were used to differentially screen a cDNA library prepared from embryogenic cell masses of somatic embryo cultures to identify early-stage embryo transcripts. The three alfalfa somatic embryogenesis-specific transcripts cDNAs (ASET1, ASET2 and ASET3) identified by this screen were enriched in RNA samples from embryogenic tissues of the embryogenic genotype but were absent from petioles or mature embryos of an embryogenic genotype as well from tissue cultures of a nonembryogenic genotype. The ASET clones did not cross-hybridize and showed different patterns of expression in northerns of RNA from various fractions of alfalfa somatic embryo cultures. The ASET clones did not hybridize with the soybean embryogenesis-specific clone (Sbh1) which was shown to be expressed in embryogenic and non-embryogenic alfalfa tissue cultures. Sequencing showed ASET1 to be a partial transcript 595 nucleotides long. ASET2 was a complete transcript of 1193 nucleotides. From a comparison of the predicted open reading frame with the GenBank protein database it was concluded that ASET2 was a novel transcript. The protein predicted by the ASET2 sequence has several potential membrane-spanning domains and a potential phosphorylation site. In addition, the ASET2 cDNA had a long 5' region that contained two upstream reading frames (URFs) which could potentially code for 30 and 6 amino acid polypeptides.

Amino Acid Sequence↗

Cloning of tobacco genes that elicit the hypersensitive response.

We used a functional screening method to isolate genes whose products elicit the hypersensitive response (HR) pathway of defense against plant pathogens. A cDNA library derived from tobacco leaves undergoing the HR was cloned into a tobacco mosaic virus (TMV)-based expression vector. Infectious transcripts were generated and used to inoculate tobacco plants lacking the N resistance gene (genotype Xanthi nn). Approximately 1/1000 of the infectious transcripts produced local lesions, and may thus elicit the HR. The cDNA inserts from 50 lesion-forming clones were recovered by RT-PCR, and 12 unique clones were sequenced. Comparisons with protein databases revealed homologies to (a) ubiquitin, (b) tobacco tumor-related protein, similar to Kunitz-type trypsin inhibitors and (c) ribosomal protein S14. The remaining nine clones revealed no homology to known proteins and are thus considered novel. Five clones were able to induce the expression of PR2, a gene which is specifically activated in the tobacco HR. Northern and western blot analyses of leaves infected by the clone encoding ubiquitin strongly suggest that the infection produced a co-suppression response; the endogenous level of ubiquitin mRNA and protein in infected leaves are ca. 50% less than those found in healthy leaves. This observation supports a previous report on the involvement of the ubiquitin system in the tobacco HR [2], and validates and utility of the functional cloning method.

Base Sequence↗

Identification of a novel cDNA, encoding a cytoskeletal associated protein, differentially expressed in diffuse large B cell lymphomas.

Diffuse large B-cell lymphomas (DLBL) constitute an heterogeneous clinico-pathological entity. To characterize molecular events related to histological subtypes, clinical presentation or outcome, we compared the mRNAs expressed in a limited series of DLBL by Differential display-reverse transcription (DDRT) and cloned a differential cDNA, that we called LB1. LB1 open reading frame encodes a 683 amino-acid polypeptide that does not show significant homology upon comparison to protein databases, nor any structural domain relating LB1 to an already known protein family. Immunofluorescence analysis of transfected COS cells showed a cytoplasmic filamentous staining, indicating that LB1 protein is tightly associated with cytoskeletal fibers. Two LB1 transcripts, a major 3.6-3.9 Kb and a minor 2.2 Kb transcripts, were detected among human haematopoietic and non-haematopoietic lines and tissues. LB1 transcripts were abundant in testis, thymus and in tumour derived cell lines, while barely detectable in liver, prostate and kidney. Concerning DLBL, LB1 expression was high in two cases of DLBL, and low or undetectable in four others, confirming the differential expression previously observed in the DDRT experiment. Furthermore, LB1 gene mapped to chromosome 13q14, a region that has been involved as a chromosomal breakpoint in DLBL. The cellular function of LB1 and its relationship with B cell maturation and/or oncogenesis remain to be established.

Amino Acid Sequence↗