PubMed HealthSearch

Biomedical subjects

E Gasteiger

Publications and source records attributed to E Gasteiger.

5 recordsLinked to original sources

Protein identification with N and C-terminal sequence tags in proteome projects.

Genome sequences are available for increasing numbers of organisms. The proteomes (protein complement expressed by the genome) of many such organisms are being studied with two-dimensional (2D) gel electrophoresis. Here we have investigated the application of short N-terminal and C-terminal sequence tags to the identification of proteins separated on 2D gels. The theoretical N and C termini of 15, 519 proteins, representing all SWISS-PROT entries for the organisms Mycoplasma genitalium, Bacillus subtilis, Escherichia coli, Saccharomyces cerevisiae and human, were analysed. Sequence tags were found to be surprisingly specific, with N-terminal tags of four amino acid residues found to be unique for between 43% and 83% of proteins, and C-terminal tags of four amino acid residues unique for between 74% and 97% of proteins, depending on the species studied. Sequence tags of five amino acid residues were found to be even more specific. To utilise this specificity of sequence tags for protein identification, we created a world-wide web-accessible protein identification program, TagIdent (http://www.expasy.ch/www/tools.html), which matches sequence tags of up to six amino acid residues as well as estimated protein pI and mass against proteins in the SWISS-PROT database. We demonstrate the utility of this identification approach with sequence tags generated from 91 different E. coli proteins purified by 2D gel electrophoresis. Fifty-one proteins were unambiguously identified by virtue of their sequence tags and estimated pI and mass, and a further 11 proteins identified when sequence tags were combined with protein amino acid composition data. We conlcude that the TagIdent identification approach is best suited to the identification of proteins from prokaryotes whose complete genome sequences are available. The approach is less well suited to proteins from eukaryotes, as many eukaryotic proteins are not amenable to sequencing via Edman degradation, and tag protein identification cannot be unambiguous unless an organism's complete sequence is available.

Amino Acid Sequence

Two-dimensional gel electrophoresis for proteome projects: the effects of protein hydrophobicity and copy number.

Two-dimensional (2-D) gel electrophoresis is often used in proteome projects to provide a global view of the proteins expressed in any cell or tissue type. Here we have investigated the effects of protein hydrophobicity and cellular protein copy number on a protein's presence or absence on a two-dimensional gel. The average hydropathy values of all known proteins from Bacillus subtilis, Escherichia coli and Saccharomyces cerevisiae were calculated, thus defining the range of protein hydrophobicity and hydrophilicity in these organisms. The average hydropathy values were then calculated for a total of 427 proteins from these species, which had been identified elsewhere on 2-D gels. Strikingly, it was seen that no highly hydrophobic proteins, as defined by average hydrophobicity values, have been found to date on 2-D gel separations of whole cell lysates. A clear hydrophobicity cutoff point was seen, above which current 2-D electrophoresis methods appear not to be useful for protein separation. The effect of cellular protein copy number on a protein's presence on a 2-D gel was investigated by means of a graphical model. This model showed how variations in protein loading and copy number per cell interact to determine the quantity of a protein that will be present on a 2-D gel. Considering the current maximum in 2-D gel loading capacity, it was found that 2-D probably can not visualize or produce analytical quantities of proteins present at less than 1000 copies per cell. We conclude that further developments of 2-D electrophoresis techniques are desirable to enable the visualization and analysis of all proteins expressed by a cell or tissue.

Bacillus subtilis

LALNVIEW: a graphical viewer for pairwise sequence alignments.

LALNVIEW is a graphical program for visualising local alignments between two sequences (protein or nucleic acids). Sequences are represented by coloured rectangles to give an overall picture of their similarities. LALNVIEW can display sequence features (exon, intron, active site, domain, propeptide, etc.) along with the alignment. When using LALNVIEW through our Web servers, sequence features are automatically extracted from database annotations (SWISS-PROT, GenBank, EMBL or HOVERGEN) and displayed with the alignment. LALNVIEW is a useful tool for analysing pairwise sequence alignments and for making the link between sequence homology and what is known about the structure or function of sequences. LALNVIEW executables for UNIX, Macintosh and PC computers are freely available from our server (http:// expasy.hcuge.ch/sprot/lalnview.html).

Acyltransferases

Detailed peptide characterization using PEPTIDEMASS--a World-Wide-Web-accessible tool.

In peptide mass fingerprinting, there are frequently peptides whose masses cannot be explained. These are usually attributed to either a missed cleavage site during the chemical or enzymatic cutting process, the lack of reduction and alkylation of a protein, protein modifications like the oxidation of methionine, or the presence of protein post-translational modifications. However, they could equally be due to database errors, unusual splicing events, variants of a protein in a population, or artifactual protein modifications. Unfortunately the verification of each of these possibilities can be tedious and time-consuming. To better utilize annotated protein databases for the understanding of peptide mass fingerprinting data, we have written the program "PEPTIDEMASS". This program generates the theoretical peptide masses of any protein in the SWISS-PROT database, or of any sequence specified by the user. If the sequence is derived from the SWISS-PROT database, the program takes into account any annotations for that protein in order to generate the peptide masses. In this manner, the user can obtain the predicted masses of peptides from proteins which are known to have signal sequences, propeptides, transit peptides, simple post-translational modifications, and disulfide bonds. Users are also warned if any peptide masses are subject to change from protein isoforms, database conflicts, or an mRNA splicing variation. The program is freely accessible to the scientific community via the ExPASy World Wide Web server, at the URL address: http://www.expasy.ch/www/tools.html.

Amino Acid Sequence