PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

MELDB: a database for microbial esterases and lipases.

MELDB is a comprehensive protein database of microbial esterases and lipases which are hydrolytic enzymes important in the modern industry. Proteins in MELDB are clustered into groups according to their sequence similarities based on a local pairwise alignment algorithm and a graph clustering algorithm (TribeMCL). This differs from traditional approaches that use global pairwise alignment and joining methods. Our procedure was able to reduce the noise caused by dubious alignment in the distantly related or unrelated regions in the sequences. In the database, 883 esterase and lipase sequences derived from microbial sources are deposited and conserved parts of each protein are identified. HMM profiles of each cluster were generated to classify unknown sequences. Contents of the database can be keyword-searched and query sequences can be aligned to sequence profiles and sequences themselves.

Amino Acid Sequence↗

The Hansenula polymorpha (strain CBS4732) genome sequencing and analysis.

The methylotrophic yeast Hansenula polymorpha is a recognised model system for investigation of peroxisomal function, special metabolic pathways like methanol metabolism, of nitrate assimilation or thermostability. Strain RB11, an odc1 derivative of the particular H. polymorpha isolate CBS4732 (synonymous to ATCC34438, NRRL-Y-5445, CCY38-22-2) has been developed as a platform for heterologous gene expression. The scientific and industrial significance of this organism is now being met by the characterisation of its entire genome. The H. polymorpha RB11 genome consists of approximately 9.5 Mb and is organised as six chromosomes ranging in size from 0.9 to 2.2 Mb. Over 90% of the genome was sequenced with concomitant high accuracy and assembled into 48 contigs organised on eight scaffolds (supercontigs). After manual annotation 4767 out of 5933 open reading frames (ORFs) with significant homologies to a non-redundant protein database were predicted. The remaining 1166 ORFs showed no significant similarity to known proteins. The number of ORFs is comparable to that of other sequenced budding yeasts of similar genome size.

Base Sequence↗

The TetR family of transcriptional repressors.

We have developed a general profile for the proteins of the TetR family of repressors. The stretch that best defines the profile of this family is made up of 47 amino acid residues that correspond to the helix-turn-helix DNA binding motif and adjacent regions in the three-dimensional structures of TetR, QacR, CprB, and EthR, four family members for which the function and three-dimensional structure are known. We have detected a set of 2,353 nonredundant proteins belonging to this family by screening genome and protein databases with the TetR profile. Proteins of the TetR family have been found in 115 genera of gram-positive, alpha-, beta-, and gamma-proteobacteria, cyanobacteria, and archaea. The set of genes they regulate is known for 85 out of the 2,353 members of the family. These proteins are involved in the transcriptional control of multidrug efflux pumps, pathways for the biosynthesis of antibiotics, response to osmotic stress and toxic chemicals, control of catabolic pathways, differentiation processes, and pathogenicity. The regulatory network in which the family member is involved can be simple, as in TetR (i.e., TetR bound to the target operator represses tetA transcription and is released in the presence of tetracycline), or more complex, involving a series of regulatory cascades in which either the expression of the TetR family member is modulated by another regulator or the TetR family member triggers a cell response to react to environmental insults. Based on what has been learned from the cocrystals of TetR and QacR with their target operators and from their three-dimensional structures in the absence and in the presence of ligands, and based on multialignment analyses of the conserved stretch of 47 amino acids in the 2,353 TetR family members, two groups of residues have been identified. One group includes highly conserved positions involved in the proper orientation of the helix-turn-helix motif and hence seems to play a structural role. The other set of less conserved residues are involved in establishing contacts with the phosphate backbone and target bases in the operator. Information related to the TetR family of regulators has been updated in a database that can be accessed at www.bactregulators.org.

Adaptation, Physiological↗

A set-theoretic approach to database searching and clustering.

MOTIVATION: In this paper, we introduce an iterative method of database searching and apply it to design a database clustering algorithm applicable to an entire protein database. The clustering procedure relies on the quality of the database searching routine and further improves its results based on a set-theoretic analysis of a highly redundant yet efficient to generate cluster system. RESULTS: Overall, we achieve unambiguous assignment of 80% of SWISS-PROT sequences to non-overlapping sequence clusters in an entirely automatic fashion. Our results are compared to an expert-generated clustering for validation. The database searching method is fast and the clustering technique does not require time-consuming all-against-all comparison. This allows for fast clustering of large amounts of sequences. AVAILABILITY: The resulting clustering for the PIR1 (Release 51) and SWISS-PROT (Release 34) databases is available over the Internet from http://www.dkfz-heidelberg.de/tbi/services/modest/b rowsesysters.pl. CONTACT: a.krause@dkfz-heidelberg.de; m.vingron@dkfz-heidelberg.de

Algorithms↗

The CATH Database provides insights into protein structure/function relationships.

We report the latest release (version 1.4) of the CATH protein domains database (http://www.biochem.ucl.ac.uk/bsm/cath). This is a hierarchical classification of 13 359 protein domain structures into evolutionary families and structural groupings. We currently identify 827 homologous families in which the proteins have both structual similarity and sequence and/or functional similarity. These can be further clustered into 593 fold groups and 32 distinct architectures. Using our structural classification and associated data on protein functions, stored in the database (EC identifiers, SWISS-PROT keywords and information from the Enzyme database and literature) we have been able to analyse the correlation between the 3D structure and function. More than 96% of folds in the PDB are associated with a single homologous family. However, within the superfolds, three or more different functions are observed. Considering enzyme functions, more than 95% of clearly homologous families exhibit either single or closely related functions, as demonstrated by the EC identifiers of their relatives. Our analysis supports the view that determining structures, for example as part of a 'structural genomics' initiative, will make a major contribution to interpreting genome data.

Algorithms↗

Mapping and identification of Brucella melitensis proteins by two-dimensional electrophoresis and microsequencing.

Two-dimensional (2-D) gel electrophoresis was used to map Brucella melitensis proteins. The 2-D proteins map of B. melitensis B115 revealed 595 silver-stained protein spots separated by both isoelectric point and molecular mass. Twenty-five proteins were identified either by immunoblotting using monoclonal antibodies (MAbs) or by N-terminal microsequencing. The protein spots identified by MAbs were the 89 kDa outer membrane protein, DnaK, bacterioferritin, CP24, and BP26. Some spots were identified by N-terminal microsequencing as proteins whose sequences had been reported previously from Brucella, such as three heat-shock proteins, namely DnaK, GroEL and GroES; bacterioferritin; Cu-Zn superoxide dismutase; and the 50S ribosomal protein L7/L12. Other proteins had amino acid sequences homologous with those of various proteins from other bacteria found in protein databases: ClpP; the 10K-S protein; the ORFU phosphoprotein; succinyl-CoA synthetase alpha sub-unit; an inorganic pyrophosphatase; the Fe and/or Mn superoxide dismutase; the nucleoside diphosphate kinase, an amino acid ABC type transporter, and an electron transfer flavoprotein small subunit. Seven proteins were identified with N-terminal sequences not yet reported in databases. The 2-D map established in this study will be the basis for comparative studies of protein expression in Brucella.

Amino Acid Sequence↗

Proteome analysis reveals elevated serum levels of clusterin in patients with preeclampsia.

Preeclampsia is a pregnancy-specific syndrome and a major cause of maternal mortality. The pathophysiology of preeclampsia is unknown, and no proteome analysis of preeclampsia has been reported. We sought to identify proteins associated with preeclampsia using a proteomic technique and performed two-dimensional electrophoresis (2-DE) on sera from six patients with preeclampsia and six normal pregnant women, followed by comparison of the SYPRO Ruby-stained 2-DE profiles. A group of overexpressed spots was identified in the limited study set. Overexpressed spots were identified as clusterin by matrix-assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) followed by peptide mass fingerprinting, a protein database search, and Western blot analysis. Additionally, sera of 80 preeclamptic women and 80 normal pregnant women were processed by immunoassay methods to confirm changes in clusterin concentrations quantitatively. Immunoassays showed that clusterin levels in the 80 preeclamptic women were significantly higher than those in the 80 controls (mean +/- SD; 1.62 +/- 0.46 times reference level in preeclamptic women vs. 1.30 +/- 0.46 times reference level in controls, P < 0.001). Proteomic analysis of serum proteins is a promising tool for studying preeclampsia pathophysiology and identifying proteins associated with preeclampsia.

Blood Proteins↗

NotI clones in the analysis of the human genome.

Not I linking clones contain sequences flanking Not I recognition sites and were previously shown to be tightly associated with CpG islands and genes. To directly assess the value of Not I clones in genome research, high density grids with 50 000 Not I linking clones originating from six representative Not I linking libraries were constructed. Altogether, these libraries contained nearly 100 times the total number of Not I sites in the human genome. A total of 3437 sequences flanking Not I sites were generated. Analysis of 3265 unique sequences demonstrated that 51% of the clones displayed significant protein similarity to SWISSPROT and TREMBL database proteins based on MSPcrunch filtering with stringent parameters. Of the 3265 sequences, 1868 (57.2%) were new sequences, not present in the EMBL and EST databases (similarity < or =90%). Among these new sequences, 795 (24.3%) showed similarity to known proteins and 712 (21.8%) displayed an identity of >75% at the nucleotide level to sequences from EMBL or EST databases. The remaining 361 (11.1%) sequences were completely new, i.e. <75% identical. The work also showed tight, specific association of Not I sites with the first exon and suggest that the so-called 3' ESTs can actually be generated from 5'-ends of genes that contain Not I sites in their first exon.

Base Sequence↗

P45, an extracellular 45 kDa protein of Listeria monocytogenes with similarity to protein p60 and exhibiting peptidoglycan lytic activity.

A monoclonal antibody obtained by immunization of mice with heat-killed cells of Listeria monocytogenes serotype 4d showed reactivity towards a protein (P45) from L. monocytogenes with an apparent molecular mass of 45 kDa. This protein was detected in the culture supernatant and at the cell surface of L. monocytogenes. Proteins cross-reacting with the monoclonal antibody were present in all Listeria strains investigated, except L. grayi. The structural gene was cloned in Escherichia coli and sequenced. Translation of the gene starts at a TTG initiation codon. The gene was found to code for a protein of 402 amino acid residues with a predicted molecular mass of 42.7 kDa. It has a signal peptide of 27 amino acid residues, resulting in a molecular mass for the mature polypeptide of 39.9 kDa. Protein database searches showed that this protein has 55% similarity and 38% identity to protein p60 of L. monocytogenes and exhibits significant sequence similarities to p54 from Enterococcus faecium and Usp45 from Lactococcus lactis. P45 was shown to have peptidoglycan lytic activity and the encoding gene was named spl (secreted protein with lytic property).

Amino Acid Sequence↗

Isolation and characterization of a cell-associated protein of Bacillus pumilus PH-01.

A cell-associated protein released from Bacillus pumilus PH-01 showed an affinity for some dioxins, like 1,2,3,4-tetrachlorodibenzo-p-dioxin (TCDD) and 1,2,3,4-tetrachlorodibenzofuran (TCDF), and the concentration of the protein increased when B. pumilus PH-01 was boiled in minimal salts medium. Sodium dodecyl sulfate-polyacrylamide gel electrophoresis and matrix-assisted laser desorption ionization-mass spectrometry revealed that the boiled culture supernatant contained a major protein with a molecular mass of 5,313.4 Da. The adsorption behavior of the protein for 1,2,3,4-TCDD and 1,2,3,4-TCDF was examined by digesting it with proteinase K and trypsin, showing that the proteolyzed protein lost the ability to adsorb the compounds. The amino acid sequence of the protein was determined by automated Edman degradation and tandem mass spectrometry. A search of the protein databases showed no existence of proteins with an homologous sequence.

Adsorption↗

Peptide electroextraction for direct coupling of in-gel digests with capillary LC-MS/MS for protein identification and sequencing.

An electrophoretic method has been developed for the extraction of peptides following in-gel digests of SDS-PAGE separated proteins. During electroextraction, the peptides are trapped on a strong cation-exchange microcartridge, before analysis by capillary LC--ESI-tandem mass spectrometry. The spectra obtained by tandem mass spectrometry are searched directly against a protein database for identification of the protein from which the peptide originated. By minimizing surface exposure of the peptides during electroextraction, a reduction of the detection limits for protein identification is realized. The performance of the peptide electroextraction was compared directly with the standard extraction method for in-gel protein digests, using a standard dilution series of phosphorylase B and carbonic anhydrase, separated by SDS-PAGE. The lowest gel loading in which phosphorylase B was identified using the standard extraction method was 2.5 ng or 25 fmol, and the lowest gel loading in which phosphorylase B was identified using electroextraction was 1.25 ng or 12.5 fmol. The design of the microextraction cartridge allows for direct interfacing with capillary LC, which is crucial for maintaining low detection limits. Furthermore, this method can be used for high-throughput proteomics since it can be easily multiplexed and requires only voltage control and low pressures (approximately 15 psi) for operation. We believe that peptide electroextraction is a significant advance for identification of proteins separated by one-dimensional or two-dimensional gel electrophoresis, as it can be easily automated and requires less protein than conventional methods.

Amino Acid Sequence↗

Trypsin-based monolithic bioreactor coupled on-line with LC/MS/MS system for protein digestion and variant identification in standard solutions and serum samples.

The applicability of a trypsin-based monolithic bioreactor coupled on-line with LC/MS/MS for rapid proteolytic digestion and protein identification is here described. Dilute samples are passed through the bioreactor for generation of proteolytic fragments in less than 10 min. After digestion and peptide separation, electrospray ionization tandem mass spectrometry is used to generate a peptide map and to identify proteolytic peptides by correlating their fragmentation spectra with amino acid sequences from a protein database. By digesting picomoles of proteins sufficient data from ESI and MS/MS were obtained to unambiguously identify proteins alone and in serum samples. This approach was also extended to locate mutation sites in beta-lactoglobulin A and B variants.

Amino Acid Sequence↗

Acetyltransferase machinery conserved in p300/CBP-family proteins.

CREB-binding protein (CBP) and p300 are highly conserved and functionally related transcription coactivators and histone/protein acetyltransferases. They are tumor suppressors, participate in a wide variety of physiological events, and serve as integrators among different signal transduction pathways. In this study, 11 distinct proteins that have a high degree of homology with the amino acid sequence of p300 have been identified in current protein databases. All of these 11 proteins belong to either animal or plant multicellular organisms (higher eucaryotes). Conservation of p300/CBP domains among these proteins was examined further by sequence alignment and pattern search. The domains of p300/CBP that are required for the HAT function, including PHD, putative CoA-binding, and ZZ domains, are conserved in all of these 11 proteins. This observation is consistent with the previous functional assays and indicates that they are a family of acetyltransferases, i.e. p300/CBP acetyltransferases (PCAT). TAZ domains (TAZ1 and/or TAZ2) of PCAT proteins may allow them to participate in transcription regulation by either directly recruiting transcription factors, acetylating them subsequently, or directing targeted acetylation of nucleosomal histones.

Acetyltransferases↗

The sulphur oxygenase reductase from Acidianus ambivalens is a multimeric protein containing a low-potential mononuclear non-haem iron centre.

The SOR (sulphur oxygenase reductase) is the initial enzyme in the sulphur-oxidation pathway of Acidianus ambivalens. Expression of the sor gene in Escherichia coli resulted in active, soluble SOR and in inclusion bodies from which active SOR could be refolded as long as ferric ions were present in the refolding solution. Wild-type, recombinant and refolded SOR possessed indistinguishable properties. Conformational stability studies showed that the apparent unfolding free energy in water is approx. 5 kcal x mol(-1) (1 kcal=4.184 kJ), at pH 7. The analysis of the quaternary structures showed a ball-shaped assembly with a central hollow core probably consisting of 24 subunits in a 432 symmetry. The subunits form homodimers as the building blocks of the holoenzyme. Iron was found in the wild-type enzyme at a stoichiometry of one iron atom/subunit. EPR spectroscopy of the colourless SOR resulted in a single isotropic signal at g=4.3, characteristic of high-spin ferric iron. The signal disappeared upon reduction with dithionite or incubation with sulphur at elevated temperature. Thus both EPR and chemical analysis indicate the presence of a mononuclear iron centre, which has a reduction potential of -268 mV at pH 6.5. Protein database inspection identified four SOR protein homologues, but no other significant similarities. The spectroscopic data and the sequence comparison led to the proposal that the Acidianus ambivalens SOR typifies a new type of non-haem iron enzyme containing a mononuclear iron centre co-ordinated by carboxylate and/or histidine ligands.

Acidianus↗

Spectral clustering of protein sequences.

An important problem in genomics is automatically clustering homologous proteins when only sequence information is available. Most methods for clustering proteins are local, and are based on simply thresholding a measure related to sequence distance. We first show how locality limits the performance of such methods by analysing the distribution of distances between protein sequences. We then present a global method based on spectral clustering and provide theoretical justification of why it will have a remarkable improvement over local methods. We extensively tested our method and compared its performance with other local methods on several subsets of the SCOP (Structural Classification of Proteins) database, a gold standard for protein structure classification. We consistently observed that, the number of clusters that we obtain for a given set of proteins is close to the number of superfamilies in that set; there are fewer singletons; and the method correctly groups most remote homologs. In our experiments, the quality of the clusters as quantified by a measure that combines sensitivity and specificity was consistently better [on average, improvements were 84% over hierarchical clustering, 34% over Connected Component Analysis (CCA) (similar to GeneRAGE) and 72% over another global method, TribeMCL].

Algorithms↗

Removal of N-terminal methionine from recombinant proteins by engineered E. coli methionine aminopeptidase.

The removal of N-terminal translation initiator Met by methionine aminopeptidase (MetAP) is often crucial for the function and stability of proteins. On the basis of crystal structure and sequence alignment of MetAPs, we have engineered Escherichia coli MetAP by the mutation of three residues, Y168G, M206T, Q233G, in the substrate-binding pocket. Our engineered MetAPs are able to remove the Met from bulky or acidic penultimate residues, such as Met, His, Asp, Asn, Glu, Gln, Leu, Ile, Tyr, and Trp, as well as from small residues. The penultimate residue, the second residue after Met, was further removed if the antepenultimate residue, the third residue after Met, was small. By the coexpression of engineered MetAP in E. coli through the same or a separate vector, we have successfully produced recombinant proteins possessing an innate N terminus, such as onconase, an antitumor ribonuclease from the frog Rana pipiens. The N-terminal pyroglutamate of recombinant onconase is critical for its structural integrity, catalytic activity, and cyto-toxicity. On the basis of N-terminal sequence information in the protein database, 85%-90% of recombinant proteins should be produced in authentic form by our engineered MetAPs.

Amino Acid Substitution↗

Functional mapping of Autographa california nuclear polyhedrosis virus genes required for late gene expression.

A plasmid containing the bacterial chloramphenicol acetyltransferase (CAT) gene under the control of an Autographa california nuclear polyhedrosis virus (AcNPV) late gene promoter was constructed. This plasmid (pL2cat) also contained the AcNPV hr5 enhancer element. Transient-expression assay experiments indicated that the late promoter was active in Spodoptera frugiperda cells cotransfected with pL2cat and AcNPV DNA but not when pL2cat was transfected alone. Low levels of CAT activity were observed in cells cotransfected with pL2cat and pIE-1 DNAs. However, CAT activity was not induced in a similar plasmid which lacked the cis-linked enhancer element, indicating that the enhancer was required for expression of the late gene. Cotransfection mapping of pPstI clones of AcNPV DNA indicated that the pPstI-G clone of viral DNA contained a factor which further stimulated late gene expression 3- to 10-fold. Transient-expression assay analysis of subclones of pPstI-G localized the trans-active factor to a 3.0-kilobase XbaI fragment. The nucleotide sequence of this fragment was determined and found to contain three potential open reading frames. A computer-assisted search of a protein database revealed no closely related proteins. One of the predicted amino acid sequences contained potential metal-binding domains similar to those found in nucleic acid-binding proteins. Subcloning and subsequent CAT assay indicated that two of the open reading frames were required for the activation of pL2cat. Nuclease S1 mapping of infected and transfected RNAs indicated that the two open reading frames were transcribed as delayed-early genes. Quantitative nuclease S1 analysis and differential DNA digestion of recovered plasmids indicated that the activation of pL2cat was not due to an increase in steady-state levels of mRNA replication of the viral DNA.

Amino Acid Sequence↗

Proteomic identification of a large complement of rat urinary proteins.

The characterization of urinary proteins is an important tool to identify disease-related biomarkers and to better understand renal physiology. Expression of urinary proteins has been previously studied by Western blotting and other immunological methods. The scope of such studies, however, is limited to previously identified proteins for which specific antibodies are existed. We used proteomic analysis to identify proteins and to construct a proteome map for Sprague-Dawley (SD) rat urine isolated by ultracentrifugation. Urinary proteins were separated by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and visualized by silver staining. Proteins were identified by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS), followed by peptide mass fingerprinting using the NCBI protein database. A total of 350 protein spots were visualized. From 250 excised spots, 111 protein components were identified including transporters, transport regulators, chaperones, enzymes, signaling proteins, cytoskeletal proteins, pheromone-binding proteins, receptors, and novel gene products. The presence of a number of these identified proteins was unexpected and had not previously been identified in the urine. 2-D Western blot analyses for randomly selected proteins (ezrin, HSP70, beta- and gamma-actin, Rho-GDI, and l-myc) clearly confirmed the proteomic identification. Several potential posttranslational modifications were predicted by bioinformatic analyses. These data indicate that a large complement of expected and unexpected urinary proteins can be simultaneously studied by proteomic analysis. This approach may lead to better understanding of renal physiology and pathophysiology, and to biomarker discovery.

Actins↗