PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Sequence analysis of a mannitol dehydrogenase cDNA from plants reveals a function for the pathogenesis-related protein ELI3.

Mannitol is the most abundant sugar alcohol in nature, occurring in bacteria, fungi, lichens, and many species of vascular plants. Celery (Apium graveolens L.), a plant that forms mannitol photosynthetically, has high photosynthetic rates thought to results from intrinsic differences in the biosynthesis of hexitols vs. sugars. Celery also exhibits high salt tolerance due to the function of mannitol as an osmoprotectant. A mannitol catabolic enzyme that oxidizes mannitol to mannose (mannitol dehydrogenase, MTD) has been identified. In celery plants, MTD activity and tissue mannitol concentration are inversely related. MTD provides the initial step by which translocated mannitol is committed to central metabolism and, by regulating mannitol pool size, is important in regulating salt tolerance at the cellular level. We have now isolated, sequenced, and characterized a Mtd cDNA from celery. Analyses showed that Mtd RNA was more abundant in cells grown on mannitol and less abundant in salt-stressed cells. A protein database search revealed that the previously described ELI3 pathogenesis-related proteins from parsley and Arabidopsis are MTDs. Treatment of celery cells with salicylic acid resulted in increased MTD activity and RNA. Increased MTD activity results in an increased ability to utilize mannitol. Among other effects, this may provide an additional source of carbon and energy for response to pathogen attack. These responses of the primary enzyme controlling mannitol pool size reflect the importance of mannitol metabolism in plant responses to divergent types of environmental stress.

Amino Acid Sequence↗

Molecular characterization of bsg25D: a blastoderm-specific locus of Drosophila melanogaster.

The blastoderm stage of Drosophila embryogenesis is a time of crucial transitions in RNA transcription, the cell cycle and segment determination. We have previously identified three loci encoding RNAs specific to this stage (Roark et al., Dev. Biol. 109, 476-488, 1985). We present here the complete nucleotide sequence of one of these loci, bsg25D, which encodes a 2.7 kb blastoderm-specific RNA. The primary structure of this RNA, and that of an overlapping 4.5 kb RNA, has been determined. The amino acid sequence of the predicted bsg25D protein has been compared to the NBRF protein database. Structural similarities between domains in the bsg25D, fos, and tropomyosin proteins, and their possible significance for early embryogenesis are discussed.

Amino Acid Sequence↗

A 127 kDa component of a UV-damaged DNA-binding complex, which is defective in some xeroderma pigmentosum group E patients, is homologous to a slime mold protein.

A cDNA which encodes a approximately 127 kDa UV-damaged DNA-binding (UV-DDB) protein with high affinity for (6-4)pyrimidine dimers [Abramic', M., Levine, A.S. & Protic', M., J. Biol. Chem. 266: 22493-22500, 1991] has been isolated from a monkey cell cDNA library. The presence of this protein in complexes bound to UV-damaged DNA was confirmed by immunoblotting. The human cognate of the UV-DDB gene was localized to chromosome 11. UV-DDB mRNA was expressed in all human tissues examined, including cells from two patients with xeroderma pigmentosum (group E) that are deficient in UV-DDB activity, which suggests that the binding defect in these cells may reside in a dysfunctional UV-DDB protein. Database searches have revealed significant homology of the UV-DDB protein sequence with partial sequences of yet uncharacterized proteins from Dictyostelium discoideum (44% identity over 529 amino acids) and Oryza sativa (54% identity over 74 residues). According to our results, the UV-DDB polypeptide belongs to a highly conserved, structurally novel family of proteins that may be involved in the early steps of the UV response, e.g., DNA damage recognition.

Amino Acid Sequence↗

Nucleotide sequence of 42 kbp of vaccinia virus strain WR from near the right inverted terminal repeat.

The nucleotide sequence of 42090 bp of vaccinia virus strain WR is presented. The sequence includes the SalI L, F, G and I fragments and starts near the centre of the HindIII A fragment and extends rightwards towards the genomic terminus, finishing approximately 0.5 kb internal of the inverted terminal repeat (ITR). Translation of this region has identified 65 open reading frames (ORFs) of greater than 65 amino acids in length. Fifty-one of these which do not extensively overlap other larger ORFs have been subjected to further analysis; the other 14 are termed minor ORFs. In the rightmost 28.7 kb, the genes are, with one exception, transcribed towards the genomic terminus, similar to the arrangement of genes at the left end of the virus genome. Internal of this region the genes are expressed off either DNA strand but still predominately rightwards. ORFs are tightly packed with few intergenic non-coding regions of greater than 250 bp. Protein sequence comparisons have established a remarkably high number of homologies with entries in existing protein databases. Of these, DNA ligase, thymidylate kinase, two serine-threonine protein kinases, two serine proteinase inhibitors (serpins), two interleukin-1 receptor homologous and a discontinuous ORF related to tumour necrosis factor receptor have been reported. Other homologies include lectins, profilin, 3 beta-hydroxy steroid dehydrogenase, superoxide dismutase, guanylate kinase, ankyrin and complement factor H. In addition, there are a number of polypeptides with predicted properties of membrane-associated, secretory or glyco-proteins. Twelve gene families are described here and elsewhere. There is considerable similarity between genes from the right and left end of the virus genome that may have arisen by terminal transposition events. Several differences from the corresponding region of vaccinia virus strain Copenhagen sequence are noted. Near the right terminus the sequences diverge completely, and internal of this there are multiple examples of deletion of short sequences (eight to 10 nucleotides) that lie within penta- or hexanucleotide direct repeats.

Amino Acid Sequence↗

Analysis of bovine herpesvirus 4 genomic regions located outside the conserved gammaherpesvirus gene blocks.

Bovine herpesvirus 4 (BHV-4) DNA sequences located outside the gene blocks conserved among the gammaherpesviruses BHV-4, herpesvirus saimiri (HVS) and Epstein-Barr virus (EBV) were analysed. Twelve potential open reading frames (ORFs) were found. Protein database comparisons showed that no ORF translation products were similar to proteins encoded by alpha- or betaherpesviruses. Nevertheless, six of the ORFs were homologous in amino acid sequences to proteins encoded by HVS but apparently not to those encoded by EBV. Furthermore, the location and orientation of these six ORFs in the BHV-4 genome were similar to the corresponding ORFs in the HVS genome. No genes homologous to known cellular genes were found in the BHV-4 genome; this feature is the major difference between the BHV-4 and HVS genomes with regards to the overall gene content.

Amino Acid Sequence↗

Transposon mutations in the flagella biosynthetic pathway of the solvent-tolerant Pseudomonas putida S12 result in a decreased expression of solvent efflux genes.

Fourteen solvent-sensitive transposon mutants were generated from the solvent-tolerant Pseudomonas putida strain S12 by applying the TnMod-KmO mutagenesis system. These mutants were unable to grow in the presence of octanol and toluene. By cloning the region flanking the transposon insertion point a partial sequence of the interrupted genes was determined. Comparison of the deduced amino acid sequences with a protein database revealed the following interrupted putative gene products: organic solvent efflux proteins SrpA and SrpB, the flagellar structural proteins FlgK, FlaG, FliI, FliC, and FliH, the transcriptional activator FleQ, the alternative RNA polymerase sigma factor RpoN, and the flagellum-specific RNA polymerase sigma factor FliA (RpoF). The transposon mutants, except for the organic solvent efflux mutants, were nonmotile as determined by a swarm assay and the formation of the flagellum was totally impaired. Expression studies with a srp promoter probe showed a decreased expression of the SrpABC efflux pump in the nonmotile mutants.

Bacterial Proteins↗

A random survey of the Cryptosporidium parvum genome.

Cryptosporidium parvum is an obligate intracellular pathogen responsible for widespread infections in humans and animals. The inability to obtain purified samples of this organism's various developmental stages has limited the understanding of the biochemical mechanisms important for C. parvum development or host-parasite interaction. To identify C. parvum genes independent of their developmental expression, a random sequence analysis of the 10.4-megabase genome of C. parvum was undertaken. Total genomic DNA was sheared by nebulization, and fragments between 800 and 1,500 bp were gel purified and cloned into a plasmid vector. A total of 442 clones were randomly selected and subjected to automated sequencing by using one or two primers flanking the cloning site. In this way, 654 genomic survey sequences (GSSs) were generated, corresponding to >320 kb of genomic sequence. These sequences were assembled into 408 contigs containing >250 kb of unique sequence, representing approximately 2.5% of the C. parvum genome. Comparison of the GSSs with sequences in the public DNA and protein databases revealed that 107 contigs (26%) displayed similarity to previously identified proteins and rRNA and tRNA genes. These included putative genes involved in the glycolytic pathway, DNA, RNA, and protein metabolism, and signal transduction pathways. The repetitive sequence elements identified included a telomere-like sequence containing hexamer repeats, 57 microsatellite-like elements composed of dinucleotide or trinucleotide repeats, and a direct repeat sequence. This study demonstrates that large-scale genomic sequencing is an efficient approach to analyze the organizational characteristics and information content of the C. parvum genome.

Amino Acid Sequence↗

Identification of a novel gene, aut, involved in autotrophic growth of Alcaligenes eutrophus.

The aerobic facultative chemoautotroph Alcaligenes eutrophus was found to possess a novel gene, designated aut, required for both lithoautotrophic (hydrogen plus carbon dioxide) and organoautotrophic (formate) growth (Aut+ phenotype). Insertional mutagenesis by transposon Tn5-Mob localized the gene on a chromosomal 13-kbp EcoRI fragment. Physiological characterization of various Aut- mutants revealed pleiotropic effects caused by the transposon insertion. Heterotrophic growth of the mutants on substrates catabolized via the glycolytic pathway was slower than that of the parent strains, and the colony morphology of the mutants was altered when grown on nutrient agar. The heterotrophic derepression of the cbb operons encoding Calvin cycle enzymes was abolished, although their expression was still inducible in the presence of formate. Apparently, the mutation did not affect the cbb genes directly but impaired the autotrophic growth in a more general manner. The conjugally transferred wild-type EcoRI fragment allowed phenotypic in trans complementation of the mutants. Further subcloning and sequencing identified a single open reading frame (aut) of 495 bp that was sufficient for complementation. The monocistronic aut gene was constitutively transcribed into a 0.65-kb mRNA. However, its expression appeared to be low. Heterologous expression of aut was achieved in Escherichia coli, resulting in overproduction of an 18-kDa protein. Database searches yielded weak partial sequence similarities of the deduced Aut protein sequence to some cytidylyltransferases, but no indication for the exact function of the aut gene was obtained. Hybridizing DNA sequences that might be similar to the aut gene were detected by Southern hybridization in the genome of two other autotrophic bacteria.

Alcaligenes↗

Proteomic analysis of differently expressed proteins in a mouse model for allergic asthma.

Allergic asthma is associated with persistent functional and structural changes in the airways and involves many different cell types. Many proteins involved in allergic asthma have been identified individually, but complete protein profiles (proteome) have not yet been reported. Here we have used a differential proteome mapping strategy to identify tissue proteins that are differentially expressed in mice with allergic asthma and in normal mice. Mouse lung tissue proteins were separated using two-dimensional gel electrophoresis over a pH range between 4 and 7, digested, and then analyzed by matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MS). The proteins were identified using automated MS data acquisition. The resulting data were searched against a protein database using an internal Mascot search routine. This approach identified 15 proteins that were differentially expressed in the lungs of mice with allergic asthma and normal mice. All 15 proteins were identified by MS, and 9 could be linked to asthma-related symptoms, oxidation, or tissue remodeling. Our data suggest that these proteins may prove useful as surrogate biomarkers for quantitatively monitoring disease state progression or response to therapy.

Animals↗

Mouse embryonic fibroblasts derived from Odin deficient mice display a hyperproliiferative phenotype.

Odin is a recently identified cytosolic phosphotyrosine binding (PTB) domain containing negative regulatory protein, that was discovered on the basis of its ability to undergo tyrosine phosphorylation upon stimulation by epidermal growth factor in HeLa cells. The protein was originally obtained as a KIAA clone (KIAA 0229) from the Kazusa DNA Research Institute which maintains the HUGE protein database--a database devoted to the analysis of long cDNA clones encoding large proteins (>50 kDa). Odin has been demonstrated to cause downregulation of c-Fos promoter activity and to inhibit PDGF-induced mitogenesis in cell lines. To further investigate the role of Odin in growth factor receptor signaling and to elucidate its biological function in vivo, we have generated mice deficient in Odin by gene targeting. Odin-deficient mice do not display any obvious phenotype, and histological examination of the kidney, lung and liver does not show any major abnormalities as compared to wild-type controls. However, mouse embryonic fibroblasts (MEFs) generated from Odin-deficient mice exhibit a hyperproliferative phenotype compared to wild-type-derived MEFs, consistent with its role as a negative regulator of growth factor receptor signaling. Our results confirm that although Odin expression in mice is not essential for any major developmental pathway, it could play a significant functional role to negatively regulate growth factor receptor signaling pathways.

Adaptor Proteins, Signal Transducing↗

Protein family classification using sparse markov transducers.

We present a method for classifying proteins into families based on short subsequences of amino acids using a new probabilistic model called sparse Markov transducers (SMT). We classify a protein by estimating probability distributions over subsequences of amino acids from the protein. Sparse Markov transducers, similar to probabilistic suffix trees, estimate a probability distribution conditioned on an input sequence. SMTs generalize probabilistic suffix trees by allowing for wild-cards in the conditioning sequences. Since substitutions of amino acids are common in protein families, incorporating wild-cards into the model significantly improves classification performance. We present two models for building protein family classifiers using SMTs. As protein databases become larger, data driven learning algorithms for probabilistic models such as SMTs will require vast amounts of memory. We therefore describe and use efficient data structures to improve the memory usage of SMTs. We evaluate SMTs by building protein family classifiers using the Pfam and SCOP databases and compare our results to previously published results and state-of-the-art protein homology detection methods. SMTs outperform previous probabilistic suffix tree methods and under certain conditions perform comparably to state-of-the-art protein homology methods.

Algorithms↗

The PSSH database of alignments between protein sequences and tertiary structures.

We introduce the PSSH ('Protein Sequence-to-Structure Homologies') database derived from HSSP2, an improved version of the HSSP ('Homology-derived Secondary Structure of Proteins') database [Dodge et al. (1998) Nucleic Acids Res., 26, 313-315]. Whereas each HSSP entry lists all protein sequences related to a given 3D structure, PSSH is the 'inverse', with each entry listing all structures related to a given sequence. In addition, we introduce two other derived databases: HSSPchain, in which each entry lists all sequences related to a given PDB chain, and HSSPalign, in which each entry gives details of one sequence aligned onto one PDB chain. This re-organization makes it easier to navigate from sequence to structure, and to map sequence features onto 3D structures. Currently (September 2002), PSSH provides structural information for over 400 000 protein sequences, covering 48% of SWALL and 61% of SWISS-PROT sequences; HSSPchain provides sequence information for over 25 000 PDB chains, and HSSPalign gives over 14 million sequence-to-structure alignments. The databases can be accessed via SRS 3D, an extension to the SRS system, at http://srs3d.ebi.ac.uk/.

Animals↗

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

MitoP2, an integrated database on mitochondrial proteins in yeast and man.

The aim of the MitoP2 database (http://ihg.gsf.de/mitop2) is to provide a comprehensive list of mitochondrial proteins of yeast and man. Based on the current literature we created an annotated reference set of yeast and human proteins. In addition, data sets relevant to the study of the mitochondrial proteome are integrated and accessible via search tools and links. They include computational predictions of signalling sequences, and summarize results from proteome mapping, mutant screening, expression profiling, protein-protein interaction and cellular sublocalization studies. For each individual approach, specificity and sensitivity for allocating mitochondrial proteins was calculated. By providing the evidence for mitochondrial candidate proteins the MitoP2 database lends itself to the genetic characterization of human mitochondriopathies.

Computational Biology↗

Het-PDB Navi.: a database for protein-small molecule interactions.

The genomes of more than 100 species have been sequenced, and the biological functions of encoded proteins are now actively being researched. Protein function is based on interactions between proteins and other molecules. One approach to assuming protein function based on genomic sequence is to predict interactions between an encoded protein and other molecules. As a data source for such predictions, knowledge regarding known protein-small molecule interactions needs to be compiled. We have, therefore, surveyed interactions between proteins and other molecules in Protein Data Bank (PDB), the protein three-dimensional (3D) structure database. Among 20,685 entries in PDB (April, 2003), 4,189 types of small molecules were found to interact with proteins. Biologically relevant small molecules most often found in PDB were metal ions, such as calcium, zinc, and magnesium. Sugars and nucleotides were the next most common. These molecules are known to act as cofactors for enzymes and/or stabilizers of proteins. In each case of interactions between a protein and small molecule, we found preferred amino acid residues at the interaction sites. These preferences can be the basis for predicting protein function from genomic sequence and protein 3D structures. The data pertaining to these small molecules were collected in a database named Het-PDB Navi., which is freely available at http://daisy.nagahama-i-bio.ac.jp/golab/hetpdbnavi.html and linked to the official PDB home page.

Adenosine Triphosphate↗

ARAMEMNON, a novel database for Arabidopsis integral membrane proteins.

A specialized database (DB) for Arabidopsis membrane proteins, ARAMEMNON, was designed that facilitates the interpretation of gene and protein sequence data by integrating features that are presently only available from individual sources. Using several publicly available prediction programs, putative integral membrane proteins were identified among the approximately 25,500 proteins in the Arabidopsis genome DBs. By averaging the predictions from seven programs, approximately 6,500 proteins were classified as transmembrane (TM) candidate proteins. Some 1,800 of these contain at least four TM spans and are possibly linked to transport functions. The ARAMEMNON DB enables direct comparison of the predictions of seven different TM span computation programs and the predictions of subcellular localization by eight signal peptide recognition programs. A special function displays the proteins related to the query and dynamically generates a protein family structure. As a first set of proteins from other organisms, all of the approximately 700 putative membrane proteins were extracted from the genome of the cyanobacterium Synechocystis sp. and incorporated in the ARAMEMNON DB. The ARAMEMNON DB is accessible at the URL http://aramemnon.botanik.uni-koeln.de.

Amino Acid Transport Systems↗

Identification and characterization of a cell surface protein of Prevotella intermedia 17 with broad-spectrum binding activity for extracellular matrix proteins.

Prevotella intermedia binds and invades a variety of host cells. This binding is most probably mediated through cell surface proteins termed adhesins. To identify proteins binding to the host extracellular matrix (ECM) component, fibronectin, and study the molecular mechanism underlying bacterial colonization, we applied proteomic approaches to perform a global investigation of P. intermedia strain 17 outer membrane proteins. 2-DE followed by Far Western Blot analysis using fibronectin as a probe revealed a 29-kDa fibronectin-binding protein, designated here AdpB. The molecular identity of the protein was determined using PMF followed by a search of the P. intermedia 17 protein database. Database searches revealed the similarity of AdpB to multiple bacterial outer membrane proteins including the fibronectin-binding protein from Campylobacter jejuni. A recombinant AdpB protein bound fibronectin as well as other host ECM components, including fibrinogen and laminin, in a saturable, dose-dependent manner. Binding of AdpB to immobilized fibronectin was also inhibited by soluble fibronectin, laminin, and fibrinogen, indicating the binding was specific. Finally, immunoelectron microscopy with anti-AdpB demonstrated the cell surface location of the protein. This is the first cell surface protein with a broad-spectrum ECM-binding abilities identified and characterized in P. intermedia 17.

Amino Acid Sequence↗

Polar zippers: their role in human disease.

Ascaris hemoglobin consists of 8 subunits, each of which contains a C-terminal peptide with the sequence Glu-Glu-Lys-His repeated 4 times. When plotted on a beta-strand, this sequence leads to alternate lysines and glutamates on one side of the strand, and alternate glutamates and histidines on the other side, suggestive of a polar zipper that links the subunits together. A computer search of the protein database showed that the same or similar sequences also occur in other proteins. Some contain long repeats of Asp-Arg or Glu-Arg, among them the small nuclear ribonucleo-U1 70K protein, which is an autoantigen in systemic lupus erythematosis. These repeats appear to constitute the dominant epitopes in the autoimmune reaction. Single chains with Asp-Arg repeats may form alpha-helices in which alternate positively charged ridges and negatively charged grooves compensate each other. Several separate chains with Asp-Arg repeats could compensate each other's charges optimally by zipping together to beta-sheets. Several homeodomains of Drosophila, as well as the human transcription factor SP1, contain repeats of glutamines. Molecular modeling, circular dichroism, and electron and X-ray diffraction studies of a synthetic poly(L-glutamine) showed that it forms beta-sheets held together by hydrogen bonds between the main-chain and side-chain amides. Published data suggest that the function of these glutamine repeats consists of joining essential transcription factors bound to distant segments of DNA.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗