PubMed Health⌕ Search

Biomedical subjects

Elisabeth Gasteiger

Publications and source records attributed to Elisabeth Gasteiger.

15 recordsLinked to original sources

ScanProsite: detection of PROSITE signature matches and ProRule-associated functional and structural residues in proteins.

ScanProsite--http://www.expasy.org/tools/scanprosite/--is a new and improved version of the web-based tool for detecting PROSITE signature matches in protein sequences. For a number of PROSITE profiles, the tool now makes use of ProRules--context-dependent annotation templates--to detect functional and structural intra-domain residues. The detection of those features enhances the power of function prediction based on profiles. Both user-defined sequences and sequences from the UniProt Knowledgebase can be matched against custom patterns, or against PROSITE signatures. To improve response times, matches of sequences from UniProtKB against PROSITE signatures are now retrieved from a pre-computed match database. Several output modes are available including simple text views and a rich mode providing an interactive match and feature viewer with a graphical representation of results.

Amino Acids↗

The Universal Protein Resource (UniProt): an expanding universe of protein information.

The Universal Protein Resource (UniProt) provides a central resource on protein sequences and functional annotation with three database components, each addressing a key need in protein bioinformatics. The UniProt Knowledgebase (UniProtKB), comprising the manually annotated UniProtKB/Swiss-Prot section and the automatically annotated UniProtKB/TrEMBL section, is the preeminent storehouse of protein annotation. The extensive cross-references, functional and feature annotations and literature-based evidence attribution enable scientists to analyse proteins and query across databases. The UniProt Reference Clusters (UniRef) speed similarity searches via sequence space compression by merging sequences that are 100% (UniRef100), 90% (UniRef90) or 50% (UniRef50) identical. Finally, the UniProt Archive (UniParc) stores all publicly available protein sequences, containing the history of sequence data with links to the source databases. UniProt databases continue to grow in size and in availability of information. Recent and upcoming changes to database contents, formats, controlled vocabularies and services are described. New download availability includes all major releases of UniProtKB, sequence collections by taxonomic division and complete proteomes. A bibliography mapping service has been added, and an ID mapping service will be available soon. UniProt databases can be accessed online at http://www.uniprot.org or downloaded at ftp://ftp.uniprot.org/pub/databases/.

Databases, Protein↗

The Universal Protein Resource (UniProt).

The Universal Protein Resource (UniProt) provides the scientific community with a single, centralized, authoritative resource for protein sequences and functional information. Formed by uniting the Swiss-Prot, TrEMBL and PIR protein database activities, the UniProt consortium produces three layers of protein sequence databases: the UniProt Archive (UniParc), the UniProt Knowledgebase (UniProt) and the UniProt Reference (UniRef) databases. The UniProt Knowledgebase is a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase with extensive cross-references. This centrepiece consists of two sections: UniProt/Swiss-Prot, with fully, manually curated entries; and UniProt/TrEMBL, enriched with automated classification and annotation. During 2004, tens of thousands of Knowledgebase records got manually annotated or updated; we introduced a new comment line topic: TOXIC DOSE to store information on the acute toxicity of a toxin; the UniProt keyword list got augmented by additional keywords; we improved the documentation of the keywords and are continuously overhauling and standardizing the annotation of post-translational modifications. Furthermore, we introduced a new documentation file of the strains and their synonyms. Many new database cross-references were introduced and we started to make use of Digital Object Identifiers. We also achieved in collaboration with the Macromolecular Structure Database group at EBI an improved integration with structural databases by residue level mapping of sequences from the Protein Data Bank entries onto corresponding UniProt entries. For convenient sequence searches we provide the UniRef non-redundant sequence databases. The comprehensive UniParc database stores the complete body of publicly available protein sequence data. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). New releases are published every two weeks.

Amino Acid Sequence↗

UniProt: the Universal Protein knowledgebase.

To provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information, the Swiss-Prot, TrEMBL and PIR protein database activities have united to form the Universal Protein Knowledgebase (UniProt) consortium. Our mission is to provide a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and query interfaces. The central database will have two sections, corresponding to the familiar Swiss-Prot (fully manually curated entries) and TrEMBL (enriched with automated classification, annotation and extensive cross-references). For convenient sequence searches, UniProt also provides several non-redundant sequence databases. The UniProt NREF (UniRef) databases provide representative subsets of the knowledgebase suitable for efficient searching. The comprehensive UniProt Archive (UniParc) is updated daily from many public source databases. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). The scientific community is encouraged to submit data for inclusion in UniProt.

Animals↗

The Swiss-Prot variant page and the ModSNP database: a resource for sequence and structure information on human protein variants.

Missense mutation leading to single amino acid polymorphism (SAP) is the type of mutation most frequently related to human diseases. The Swiss-Prot protein knowledgebase records information on such mutations in various sections of a protein entry, namely in the "feature," "comment," and "reference" fields. To facilitate users in obtaining the most relevant information about each human SAP recorded in the knowledgebase, the Swiss-Prot Variant web pages were created to provide a summary of available sequence information, as well as additional structural information on each variant. In particular, the ModSNP database was set up to store information related to SAPs and to manage the modeling of SAPs onto protein structures via an automatic homology modeling pipeline. Currently, among the 16,566 human SAPs recorded in the Swiss-Prot knowledgebase (release 42.5, 21 November 2003), more than 25% have corresponding 3D-models. Of these variants, 47% are related to disease, 26% are polymorphisms, and 27% are not yet clearly classified. The ModSNP database is updated and the subsequent model construction pipeline is launched with each weekly Swiss-Prot release. Thus, the ModSNP database represents a valuable resource for the structural analysis of protein variation. The Swiss-Prot variant pages are accessible from the NiceProt view of a Swiss-Prot entry on the ExPASy server (www.expasy.org/), via a hyperlink created for the stable and unique identifier FTId of each human SAP.

Amino Acid Substitution↗

Annotation of post-translational modifications in the Swiss-Prot knowledge base.

High-throughput proteomic studies produce a wealth of new information regarding post-translational modifications (PTMs). The Swiss-Prot knowledge base is faced with the challenge of including this information in a consistent and structured way, in order to facilitate easy retrieval and promote understanding by biologist expert users as well as computer programs. We are therefore standardizing the annotation of PTM features represented in Swiss-Prot. Indeed, a controlled vocabulary has been associated with every described PTM. In this paper, we present the major update of the feature annotation, and, by showing a few examples, explain how the annotation is implemented and what it means. Mod-Prot, a future companion database of Swiss-Prot, devoted to the biological aspects of PTMs (i.e., general description of the process, identity of the modification enzyme(s), taxonomic range, mass modification) is briefly described. Finally we encourage once again the scientific community (i.e., both individual researchers and database maintainers) to interact with us, so that we can continuously enhance the quality and swiftness of our services.

Computational Biology↗

Swiss-Prot: juggling between evolution and stability.

We describe some of the aspects of Swiss-Prot that make it unique, explain what are the developments we believe to be necessary for the database to continue to play its role as a focal point of protein knowledge, and provide advice pertinent to the development of high-quality knowledge resources on one aspect or the other of the life sciences.

Amino Acid Sequence↗

ExPASy: The proteomics server for in-depth protein knowledge and analysis.

The ExPASy (the Expert Protein Analysis System) World Wide Web server (http://www.expasy.org), is provided as a service to the life science community by a multidisciplinary team at the Swiss Institute of Bioinformatics (SIB). It provides access to a variety of databases and analytical tools dedicated to proteins and proteomics. ExPASy databases include SWISS-PROT and TrEMBL, SWISS-2DPAGE, PROSITE, ENZYME and the SWISS-MODEL repository. Analysis tools are available for specific tasks relevant to proteomics, similarity searches, pattern and profile searches, post-translational modification prediction, topology prediction, primary, secondary and tertiary structure analysis and sequence alignment. These databases and tools are tightly interlinked: a special emphasis is placed on integration of database entries with related resources developed at the SIB and elsewhere, and the proteomics tools have been designed to read the annotations in SWISS-PROT in order to enhance their predictions. ExPASy started to operate in 1993, as the first WWW server in the field of life sciences. In addition to the main site in Switzerland, seven mirror sites in different continents currently serve the user community.

Databases, Protein↗

The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.

The SWISS-PROT protein knowledgebase (http://www.expasy.org/sprot/ and http://www.ebi.ac.uk/swissprot/) connects amino acid sequences with the current knowledge in the Life Sciences. Each protein entry provides an interdisciplinary overview of relevant information by bringing together experimental results, computed features and sometimes even contradictory conclusions. Detailed expertise that goes beyond the scope of SWISS-PROT is made available via direct links to specialised databases. SWISS-PROT provides annotated entries for all species, but concentrates on the annotation of entries from human (the HPI project) and other model organisms to ensure the presence of high quality annotation for representative members of all protein families. Part of the annotation can be transferred to other family members, as is already done for microbes by the High-quality Automated and Manual Annotation of microbial Proteomes (HAMAP) project. Protein families and groups of proteins are regularly reviewed to keep up with current scientific findings. Complementarily, TrEMBL strives to comprise all protein sequences that are not yet represented in SWISS-PROT, by incorporating a perpetually increasing level of mostly automated annotation. Researchers are welcome to contribute their knowledge to the scientific community by submitting relevant findings to SWISS-PROT at swiss-prot@expasy.org.

Animals↗

Automated annotation of microbial proteomes in SWISS-PROT.

Large-scale sequencing of prokaryotic genomes demands the automation of certain annotation tasks currently manually performed in the production of the SWISS-PROT protein knowledgebase. The HAMAP project, or 'High-quality Automated and Manual Annotation of microbial Proteomes', aims to integrate manual and automatic annotation methods in order to enhance the speed of the curation process while preserving the quality of the database annotation. Automatic annotation is only applied to entries that belong to manually defined orthologous families and to entries with no identifiable similarities (ORFans). Many checks are enforced in order to prevent the propagation of wrong annotation and to spot problematic cases, which are channelled to manual curation. The results of this annotation are integrated in SWISS-PROT, and a website is provided at http://www.expasy.org/sprot/hamap/.

Amino Acid Sequence↗

FindPept, a tool to identify unmatched masses in peptide mass fingerprinting protein identification.

FindPept (http://www.expasy.org/tools/findpept.html) is a software tool designed to identify the origin of peptide masses obtained by peptide mass fingerprinting which are not matched by existing protein identification tools. It identifies masses resulting from unspecific proteolytic cleavage, missed cleavage, protease autolysis or keratin contaminants. It also takes into account post-translational modifications derived from the annotation of the SWISS-PROT database or supplied by the user, and chemical modifications of peptides. Based on a number of experimental examples, we show that the commonly held rules for the specificity of tryptic cleavage are an oversimplification, mainly because of effects of neighboring residues, experimental conditions, and contaminants present in the enzyme sample.

Algorithms↗

Hydrogen/deuterium exchange for higher specificity of protein identification by peptide mass fingerprinting.

Genome sequencing projects produce large amounts of information that could be translated into potential protein sequences. Such amounts of material continuously increase protein database sizes. At present, 22 times more protein sequences are available in the SWISS-PROT and TrEMBL databases than 8 years ago in SWISS-PROT. One of the methods of choice for protein identification makes use of specific endoproteolytic cleavage followed by matrix-assisted laser desorption/ionisation mass spectrometric (MALDI-MS) analysis of the digested product. Since 1993, when this technique was first demonstrated, the conditions required for a correct identification have changed dramatically. Whilst 4-5 peptides with an uncertainty of 2-3 Da were sufficient for a correct identification in 1993, 10-13 peptides with less than 60 ppm mass error are now required for human and E. coli proteins. This evolution is directly related to the continuous increase in protein database sizes, which causes an increase in the number of false positive matches in identification results. Use of an information complement deduced from the primary protein sequence, in the process of identification by peptide mass fingerprints, can help to increase confidence in the identification results. In this article, we propose the exchange of labile hydrogen atoms with deuterium atoms to provide an alternative information complement. The exchange reaction with optimised techniques has shown an average 95% of hydrogen/deuterium (H/D) exchange on tryptic peptides. This level of exchange was sufficient to single out one or more peptides from a list of potential candidate proteins due to the dependence of H/D exchange on the peptide primary structure. This technique also has clear advantages in the identification of small proteins where direct protein identification is impaired by the limited number of endoproteolytic peptides. Then, information related to primary sequence obtained with this technique could help to identify proteins with high confidence without any expensive tandem mass spectrometry instruments.

Amino Acid Sequence↗

High-quality protein knowledge resource: SWISS-PROT and TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and a high level of integration with other databases. Together with its automatically annotated supplement TrEMBL, it provides a comprehensive and high-quality view of the current state of knowledge about proteins. Ongoing developments include the further improvement of functional and automatic annotation in the databases including evidence attribution with particular emphasis on the human, archaeal and bacterial proteomes and the provision of additional resources such as the International Protein Index (IPI) and XML format of SWISS-PROT and TrEMBL to the user community.

Amino Acid Sequence↗

The Sulfinator: predicting tyrosine sulfation sites in protein sequences.

UNLABELLED: Protein tyrosine sulfation is an important post-translational modification of proteins that go through the secretory pathway. No clear-cut acceptor motif can be defined that allows the prediction of tyrosine sulfation sites in polypeptide chains. The Sulfinator is a software tool that can be used to predict tyrosine sulfation sites in protein sequences with an overall accuracy of 98%. Four different Hidden Markov Models were constructed, each of them specialized to recognize sulfated tyrosine residues depending on their location within the sequence: near the N-terminus, near the C-terminus, in the center of a window with a size of at least 25 amino acids, as well as in windows containing several tyrosine residues. AVAILABILITY: The Sulfinator is accessible at (http://www.expasy.org/tools/sulfinator/). SUPPLEMENTARY INFORMATION: Sulfinator documentation is accessible at (http://www.expasy.org/tools/sulfinator/sulfinator-doc.html).

Amino Acid Sequence↗

ScanProsite: a reference implementation of a PROSITE scanning tool.

Many different software tools are available publicly to scan the PROSITE database of protein families. However, none of them, to our knowledge, wholly implements the PROSITE syntax, or satisfies all the rules for scanning a pattern against a sequence. We hereby propose a strict definition of how a PROSITE pattern is to be scanned against a sequence, and provide a reference implementation of a tool to scan PROSITE patterns, rules and profiles against protein sequences.

Computational Biology↗