PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Interrogating the human genome using uninterpreted mass spectrometry data.

The public availability of a draft assembly of the human genome has enabled us to demonstrate, for the first time, the feasibility of searching a complete, unmasked eukaryotic genome using uninterpreted mass spectrometry data. A complex LC-MS/MS data set, containing peptides from at least 22 human proteins, was searched against a comprehensive, nonidentical protein database, an expressed sequence tag (EST) database, and the International Human Genome Project draft assembly of the human genome. The results from the three searches are compared in detail, and the merits of the different databases for this application are discussed. In the case of the EST database, the UniGene index provided a method of simplifying and summarising the search results. In the case of the genomic DNA, the presence of introns prevented matching of roughly one quarter of the spectra, but the technique can provide primary experimental verification of predicted coding sequences, and has the potential to identify novel coding sequences.

Algorithms↗

Hepatocellular carcinoma: from bedside to proteomics.

Hepatocellular carcinoma (HCC or hepatoma) is the most common primary cancer of the liver. It is responsible for approximately one million deaths each year, mainly in underdeveloped and developing countries. The aetiological factors identified in the development of HCC included persistent infection by hepatitis B and hepatitis C viruses, and exposure to aflatoxins. Although immunization can protect individuals from being infected by the hepatitis B virus, the early detection of HCC in those who have been infected by the virus remains a challenge. Thus most HCCs present late and are not suitable for curative treatment. Hence there is a tremendous interest and urgency to identify novel HCC diagnostic marker(s) for early detection, and tumour specific disease associated proteins as potential therapeutic targets in the treatment of HCC. Screening for these HCC proteins has been facilitated by proteomics, a key technology in the global analysis of protein expression and understanding gene function. Present and earlier proteome analyses of HCC have used predominantly experimental in vitro systems. The protein expression profiles of several hepatoma cell lines such as HepG2, Huh7, SK-Hep1, and Hep3B have been compared with normal liver, and nontransformed cell lines (Chang and WRL-68), while a comprehensive proteome analysis to create a protein database was carried out for the cell line HCC-M. In the future, proteome analyses utilizing tumour tissues, which reflect the pathological state of HCC more closely, will be undertaken. This work will complement the gene expression studies of HCC which are already underway. Efforts have also been directed at the proteome analysis of hepatic stellate cells, as these cells play an important role in liver fibrosis. Since liver fibrosis is reversible but not cirrhosis, it is of considerable importance to identify therapeutic targets that can slow its progression.

Amino Acid Sequence↗

Identification of macrophage activation associated proteins by two-dimensional gel electrophoresis and microsequencing.

To understand activation in monocytes and macrophages we have studied changes in protein synthesis using the human monocytoid U937 cell line and two-dimensional polyacrylamide gel electrophoresis (2D PAGE) and protein sequencing. U937 cells that had been metabolically labeled during treatment with PMA, LPS, or IFN-gamma showed appreciable increases or decreases in synthesis of 14 proteins when analyzed by 2D PAGE. Although some 20 proteins are reported to be affected by these agents in U937 cells, none of them correspond with the 14 proteins studied here. Of the 14 observed changes, four spots (p41/65, p35/65, p26/44, p20/53) were up-regulated by PMA only, one (p16/44) by LPS only, five spots (p29/47, p26/45, p26/48, p12/47, p10/45) by both LPS and PMA, and, finally, one (p29/45) by all three agents. Two spots (p20/59 and p20/61) were down-regulated by IFN-gamma and one of these spots (p20/59) was up-regulated by LPS. Only one spot (p20/48) was up-regulated by IFN-gamma. Eleven spots with matching mobilities (both M(r) and pI) to those identified in U937 were observed on 2D PAGE gels from human culture derived macrophages. Ten spots from U937 were sequenced by Edman degradation. Two were could not identified from information contained in the available DNA and protein databases and thus represent novel proteins, whereas a further six of the proteins were N-terminally blocked. The remaining two (29/47 and 12/47, respectively) were identified from existing protein databases as translationally controlled tumor protein (TCTP) and cytokeratin. This is the first report of the presence of TCTP in hemopoietic cells and its modulation by PMA or LPS in any cell type. We believe that 2D PAGE and sequencing is a powerful approach for identifying key proteins in macrophage cellular activation.

Amino Acid Sequence↗

Proteome analysis on an early transformed human bronchial epithelial cell line, BEP2D, after alpha-particle irradiation.

To probe the mechanism of carcinogenesis of lung cancer at the molecular level and to find potential protein markers involved in the early phase of tumorgenesis, differential proteome analysis on primary passage cell line R15H, and early transformed cell line R15H20 derived from (238)Pu alpha-particle irradiation of human papillomavirus (HPV) 18-immortalized human bronchial epithelial cell line (BEP2D), was carried out using two-dimensional electrophoresis (2-DE) and peptide mass fingerprinting (PMF) with matrix-assisted laser desorption/ionisation-time of flight mass spectrometry. Image analysis and Student's t-test (p < 0.05) showed that three protein spots were only expressed in R15H, intensities of 43 protein spots on the gels were altered between R15H and R15H20. Two of the three spots that were only expressed in R15H were identified as high mobility group protein 1. Two proteins decreased in abundance in R15H20 were identified as maspin precursor, a tumor suppressor and aminoacylase-1. Ornithine aminotransferase and peptidyl-prolyl cis-trans isomerase A that were increased in R15H20, were also identified. Relationships between these differentially expressed proteins and the carcinogenesis mechanism of lung cancer are discussed. The protein expression profile of the R15H cell line was also constructed during the study as a reference map for further comparative proteome analysis of the irradiation induced BEP2D cell line. Of the 90 spots analyzed with PMF in the 2-DE gel of R15H cell line, 50 proteins were identified by searching the nonredundant protein database SWISS-PROT/TrEMBL.

Alpha Particles↗

Proteomics of human umbilical vein endothelial cells applied to etoposide-induced apoptosis.

We have undertaken to continue the proteomic study of human umbilical vein endothelial cells (HUVECs) using the combination of 2-DE, automated trypsin digestion, and PMF analysis after MALDI-TOF MS and peptide sequencing using nano LC-ESI-MS/MS. The overall functional characterization of the 162 identified proteins from primary cultures of HUVECs confirms the metabolic capabilities of endothelium and illustrates various cellular functions more related to cell motility and angiogenesis, protein folding, anti-oxidant defenses, signal transduction, proteasome pathway and resistance to apoptosis. In comparison with controls cells, the differential proteomic analysis of HUVECs treated by the pro-apoptotic topoisomerase inhibitor etoposide further revealed the variation of eight proteins, namely, GRP78, GRP94, valosin-containing protein, proteinase inhibitor 9, cofilin, 37-kDa laminin receptor protein, bovine apolipoprotein, and tropomyosin. These data suggest that etoposide-induced apoptosis of human vascular endothelial cells results from the intricate involvement of multiple apoptosis processes including at least the mitochondrial and the ER stress pathways. The presented 2-D pattern and protein database, as well as the data related to apoptosis of HUVECs, are available at http://www.huvec.com.

Apoptosis↗

Inverse 15N-metabolic labeling/mass spectrometry for comparative proteomics and rapid identification of protein markers/targets.

The inverse labeling/mass spectrometry strategy has been applied to protein metabolic (15)N labeling for gel-free proteomics to achieve the rapid identification of protein markers/targets. Inverse labeling involves culturing both the perturbed (by disease or by a drug treatment) and control samples each in two separate pools of normal and (15)N-enriched culture media such that four pools are produced as opposed to two in a conventional labeling approach. The inverse labeling is then achieved by combining the normal (14)N-control with the (15)N-perturbed sample, and the (15)N-control with the (14)N-perturbed sample. Both mixtures are then proteolyzed and analyzed by mass spectrometry (coupled with on-line or off-line separation). Inverse labeling overcomes difficulties associated with protein metabolic labeling with regard to isotopic peak correlation and data interpretation in the single-experiment approach (due to the non-predictable/variable mass difference). When two data sets from inverse labeling are compared, proteins of differential expression are readily recognized by a characteristic inverse labeling pattern or apparent qualitative mass shifts between the two inverse labeling analyses. MS/MS fragmentation data provide further confirmation and are subsequently used to search protein databases for protein identification. The methodology has been applied successfully to two model systems in this study. Utilizing the inverse labeling strategy, one can use any mass spectrometer of standard unit resolution, and acquire only the minimum, essential data to achieve the rapid and unambiguous identification of differentially expressed protein markers/targets. The strategy permits quick focus on the signals of differentially expressed proteins. It eliminates the detection ambiguities caused by the dynamic range of detection. Finally, inverse labeling enables the detection of covalent changes of proteins responding to a perturbation that one might fail to distinguish with a conventional labeling experiment.

Amino Acid Sequence↗

Specific peptide patterns of follicular fluids at different growth stages analyzed by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry.

Human follicular fluid (HFF) has been suggested to influence oocyte development potential, and some of HFF proteins may be potential markers for oocyte maturation during follicular development. Using matrix-assisted laser desorption /ionization-time of flight mass spectrometry (MALDI-TOF MS), the presence of specific peptide peaks in HFF which could represent the follicle development potential was evaluated. HFF from different developmental stages were first digested and the resultant peptide mixtures were directly analyzed by MALDI-TOF MS. It was shown that the frequencies of specific peaks demonstrated higher reproducibility than peak intensities after multiple measurements (>or=6 times) per sample. Using this approach, a reliable peak list for each different sample could be generated by combining the information from multiple measurements. By comparing the peak lists from different samples at different growth stages, we found that 5 specific peaks appeared in the 100% frequency category of 6 replicates in all the HFF samples containing mature oocyte. Similarly, such 25 peptide peaks were also identified for HFF containing immature oocyte. These specific peaks could be used to distinguish HFF from different stages as biomarkers related to follicle development and maturation. After searching the protein database, some proteins that are known to be involved in the development and maturation of oocyte were identified, such as apolipoprotein A-I, collagen type IV, integrin, et al. Identification of such proteins in our experiment further proved that the direct analysis of tryptic digests could be of practical value.

Female↗

Statistical analysis of intrahelical ionic interactions in alpha-helices and coiled coils.

There are many controversies concerning whether ionic interactions in alpha-helices and coiled coils actually contribute to the stabilisation and formation of these structures. Here we used a statistical approach to probe this question. We extracted unique alpha-helical and coiled coil structures from the protein database and analysed the ionic interactions between positively and negatively charged residues. The ionic interactions were categorized according to the type, spacing and order of the residues involved. Separate datasets were produced depending on the number of alpha-helices in the coiled coils and the mutual orientation of the helices. We compared the frequency of residue configurations able to form ionic interactions with their probability to form the interaction. We found a correlation between the two variables in alpha-helices, antiparallel two-stranded coiled coils and parallel two-stranded coiled coils. This indicates that some ionic interactions are indeed important for the formation and stabilisation of alpha-helices and coiled coils. We concluded that the configurations, which have simultaneously a large probability to form the ionic interaction and a frequent occurrence, are those, which have the most stabilising effect. These are the 4RE, 3ER and 4ER interactions.

Algorithms↗

Direct atmospheric pressure coupling of polyacrylamide gel electrophoresis to mass spectrometry for rapid protein sequence analysis.

Using laser desorption-atmospheric pressure chemical ionization we describe a novel approach for coupling mass spectrometry to polyacrylamide gel electrophoresis. In contrast to other approaches, the method allows for the direct sampling of a polyacrylamide gel-embedded protein without the addition of any exogenous matrixes and is performed at atmospheric pressure. After electrophoresis and enzymatic digestion, the gel is analyzed at AP by photons that desorb neutral peptide molecules, followed by corona discharge ionization in the gas-phase, and subsequent mass analysis. Our experimental results demonstrate the method to (1) rapidly identify electrophoresed proteins via "peptide fingerprinting" using protein databases, (2) detect single-amino acid polymorphisms, and (3) has potential for sub-picomole sensitivity while still maintaining in situ gel desorption-ionization at ambient conditions.

Atmospheric Pressure↗

Analysis of lectin-bound glycoproteins in snake venom from the Elapidae and Viperidae families.

This paper describes an efficient method of studying the glycoproteins found in snake venom. The glycosylation profiles of the Elapidae and Viperidae snake families were analyzed using FITC-labeled lectin glycoconjugates. The Con A-agarose affinity enrichment technique was used to fractionate glycoproteins from the N. naja kaouthia venom. The results revealed a large number of Con A binding glycoproteins, most of which have moderate to high molecular weights. To identify the proteins, the isolated glycoprotein fractions were subjected to two-dimensional electrophoresis and MALDI-TOF MS. Protein sequences were compared with published protein databases to determine for their biological functions.

Animals↗

Identification of differentially expressed proteins during larval molting of Helicoverpa armigera.

Insect molting involves many molecular processes, such as protein degradation and protein synthesis in the epidermis. Various proteins have been implicated in these processes. The differentially expressed proteins during larval molting of Helicoverpa armigera were investigated using two-dimensional electrophoresis (2-D-PAGE) and matrix-assisted laser desorption/ionization-time-of-flight-mass spectrometry (MALTI-TOF-MS). Four larval tissues sampled during molting and feeding were examined. Seventy-seven differentially expressed proteins were identified in these tissues, including 20 proteins from the fifth-molting epidermis (fifth instar molting to sixth instar), 36 proteins from the fifth-molting hemolymph, and 21 from the fifth-molting fat bodies. No obviously different spots were identified from the fifth-molting midgut under these experimental conditions. After application of MALTI-TOF-MS and similarity analysis comparing results to a Drosophila protein database, 30 proteins were identified: 10 proteins from the fifth-molting epidermis, 11 proteins from the hemolymph, and 9 proteins from fat bodies. These proteins were separated into 5 groups according to their probable functions, such as enzymes, regulators, protein hydrolases, receptors, and proteins with unknown functions. These differentially expressed proteins were proposed to be involved in the Helicoverpa molting cascade.

Animals↗

Designability of alpha-helical proteins.

A typical protein structure is a compact packing of connected alpha-helices and/or beta-strands. We have developed a method for generating the ensemble of compact structures a given set of helices and strands can form. The method is tested on structures composed of four alpha-helices connected by short turns. All such natural four-helix bundles that are connected by short turns seen in nature are reproduced to closer than 3.6 A per residue within the ensemble. Because structures with no natural counterpart may be targets for ab initio structure design, the designability of each structure in the ensemble-defined as the number of sequences with that structure as their lowest-energy state-is evaluated using a hydrophobic energy. For the case of four alpha-helices, a small set of highly designable structures emerges, most of which have an analog among the known four-helix fold families; however, several packings and topologies with no analogs in protein database are identified.

Amino Acid Sequence↗

Molecular dynamics simulations of beta-turn forming tetra- and hexapeptides.

It was previously shown that the structural ensemble of model peptides DDKG and GKDG (H. Ishii et al. Biopolymers 24, 2045-2056, 1985), DEKS (A. Otter et al. J. Biomol. Struct. Dyn. 7, 455-476, 1989) NPGQ (F. R. Carbone et al. Int. J. Pept. Protein. Res. 26, 498-508, 1985), SALN (H. Santa et al. J. Biomol. Struct. Dyn. 16, 1033-1041, 1999), SYPFDV and SYPYDV (J. Yao et al. J. Mol. Biol. 243, 736-753, 1994), VP(D)AH and VP(D)SH (B. Imperiali et al. J. Am. Chem. Soc. 114, 3182-3188, 1992) in solution contains a significant - or in some cases dominant - proportion of beta-turn conformation. In this study, a protein database was searched for the above, unprotected sequences which incorporate only L-amino acid residues. Simulated annealing and 25 ns MD simulations of structures were also performed. The DSSP and STRIDE secondary structure-assigning algorithms and clustering were used to analyze trajectories and i, i+3 hydrogen bonds were also sought. The DSSP analysis showed a fluctuation between beta-turn and random meander structure, although bend structures were not detected because of the insufficient length of peptide chains. This alternating trend was confirmed when the STRIDE algorithm was used to analyze trajectories, but STRIDE assigned more turn structures. The population of the strongest clusters was above 40% and the middle structures adopted beta-turn structure for most sequences. These results are in good agreement with previous experimental results and support the idea of the ultra-marginal stability of turns in the absence of stabilizing long-range interactions of the neighboring segments of a polypeptide chain. However, interactions between the side-chains in tetrapeptides could also contribute to turn stability and result in unusual stability in some cases. Our observations suggest that such interactions are the consequence rather than the driving force of turn formation.

Amino Acid Sequence↗

An insight into domain combinations.

Domains are the building blocks of all globular proteins, and are units of compact three-dimensional structure as well as evolutionary units. There is a limited repertoire of domain families, so that these domain families are duplicated and combined in different ways to form the set of proteins in a genome. Proteins are gene products. The processes that produce new genes are duplication and recombination as well as gene fusion and fission. We attempt to gain an overview of these processes by studying the structural domains in the proteins of seven genomes from the three kingdoms of life: Eubacteria, Archaea and Eukaryota. We use here the domain and superfamily definitions in Structural Classification of Proteins Database (SCOP) in order to map pairs of adjacent domains in genome sequences in terms of their superfamily combinations. We find 624 out of the 764 superfamilies in SCOP in these genomes, and the 624 families occur in 585 pairwise combinations. Most families are observed in combination with one or two other families, while a few families are very versatile in their combinatorial behaviour. This type of pattern can be described by a scale-free network. Finally, we study domain repeats and we compare the set of the domain combinations in the genomes to those in PDB, and discuss the implications for structural genomics.

Computational Biology↗

NQ-Flipper: validation and correction of asparagine/glutamine amide rotamers in protein crystal structures.

The error rate of asparagine (Asn) and glutamine (Gln) amide rotamers in protein crystal structures is in the order of 20% and as a consequence the current Protein Database (PDB) contains approximately half a million incorrect Asn and Gln side-chain rotamers. Here we present NQ-Flipper, a web service based on knowledge-based potentials of mean force to automatically detect and correct erroneous rotamers. We achieve excellent agreement with expert curated data.

Asparagine↗

Expressed sequence tags from the Closterium peracerosum-strigosum-littorale complex, a unicellular charophycean alga, in the sexual reproduction process.

We obtained genetic information on sexual reproduction of the Closterium peracerosum-strigosum-littorale complex, a unicellular charophycean alga. Normalized cDNA libraries were constructed from cells in the sexual reproduction process, and a total of 1190 5'-end expressed sequence tags were established. Since 604 of these ESTs were classified into 174 non-redundant sequences, these 1190 ESTs include 760 unique sequences. Similarity search against a public non-redundant protein database indicated that 390 unique sequences had significant similarity to registered sequences. Among these 390 sequences, 3 were identical to and 4 were homologous to previously identified sex-pheromone genes. According to our study, 370 of 760 unique sequences are likely to be novel transcripts. These cDNA clones and EST sequence information may would be helpful for future functional analyses using DNA array technologies.

Amino Acid Sequence↗

Comparative analysis of proteins with a mucus-binding domain found exclusively in lactic acid bacteria.

Lactic acid bacteria (LAB) are frequently encountered inhabitants of the human intestinal tract. A protective layer of mucus covers the epithelial cells of the intestine, offering an attachment site for these bacteria. In this study bioinformatics tools were used to identify and characterize proteins containing one type of mucus-binding domain, called MUB, that is postulated to play an important role in the adherence of LAB to this mucus layer. By searching in all protein databases 48 proteins containing at least one of these MUB domains in nine LAB species were identified. These MUB domains varied in size, ranging from approximately 100 to more than 200 residues per domain. Complete MUB domains were found exclusively in LAB. The number of MUB domains present in a single protein varied from 1 to 15. In some cases, orthologous proteins in closely related species contained a different number of domains, indicating that repeats of the domain undergo rapid duplication and deletion. Proteins containing the MUB domain were often encoded by gene clusters that encode multiple extracellular proteins. In addition to one or more copies of the MUB domain, many of these proteins contained other domains that are predicted to be involved in binding to and degradation of extracellular components. These findings strongly suggest that the MUB domain is an LAB-specific functional unit that performs its task in various domain contexts and could fulfil an important role in host-microbe interactions in the gastrointestinal tract.

Amino Acid Sequence↗