PubMed Health⌕ Search

Biomedical subjects

Marco Pagni

Publications and source records attributed to Marco Pagni.

10 recordsLinked to original sources

MyHits: a new interactive resource for protein annotation and domain identification.

The MyHits web server (http://myhits.isb-sib.ch) is a new integrated service dedicated to the annotation of protein sequences and to the analysis of their domains and signatures. Guest users can use the system anonymously, with full access to (i) standard bioinformatics programs (e.g. PSI-BLAST, ClustalW, T-Coffee, Jalview); (ii) a large number of protein sequence databases, including standard (Swiss-Prot, TrEMBL) and locally developed databases (splice variants); (iii) databases of protein motifs (Prosite, Interpro); (iv) a precomputed list of matches ('hits') between the sequence and motif databases. All databases are updated on a weekly basis and the hit list is kept up to date incrementally. The MyHits server also includes a new collection of tools to generate graphical representations of pairwise and multiple sequence alignments including their annotated features. Free registration enables users to upload their own sequences and motifs to private databases. These are then made available through the same web interface and the same set of analytical tools. Registered users can manage their own sequences and annotations using only web tools and freeze their data in their private database for publication purposes.

Computer Graphics↗

trome, trEST and trGEN: databases of predicted protein sequences.

We previously introduced two new protein databases (trEST and trGEN) of hypothetical protein sequences predicted from EST and HTG sequences, respectively. Here, we present the updates made on these two databases plus a new database (trome), which uses alignments of EST data to HTG or full genomes to generate virtual transcripts and coding sequences. This new database is of higher quality and since it contains the information in a much denser format it is of much smaller size. These new databases are in a Swiss-Prot-like format and are updated on a weekly basis (trEST and trGEN) or every 3 months (trome). They can be downloaded by anonymous ftp from ftp://ftp.isrec.isb-sib.ch/pub/databases.

Animals↗

Swiss EMBnet node web server.

EMBnet is a consortium of collaborating bioinformatics groups located mainly within Europe (http://www.embnet.org). Each member country is represented by a 'node', a group responsible for the maintenance of local services for their users (e.g. education, training, software, database distribution, technical support, helpdesk). Among these services a web portal with links and access to locally developed and maintained software is essential and different for each node. Our web portal targets biomedical scientists in Switzerland and elsewhere, offering them access to a collection of important sequence analysis tools mirrored from other sites or developed locally. We describe here the Swiss EMBnet node web site (http://www.ch.embnet.org), which presents a number of original services not available anywhere else.

Databases, Protein↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

Automated annotation of microbial proteomes in SWISS-PROT.

Large-scale sequencing of prokaryotic genomes demands the automation of certain annotation tasks currently manually performed in the production of the SWISS-PROT protein knowledgebase. The HAMAP project, or 'High-quality Automated and Manual Annotation of microbial Proteomes', aims to integrate manual and automatic annotation methods in order to enhance the speed of the curation process while preserving the quality of the database annotation. Automatic annotation is only applied to entries that belong to manually defined orthologous families and to entries with no identifiable similarities (ORFans). Many checks are enforced in order to prevent the propagation of wrong annotation and to spot problematic cases, which are channelled to manual curation. The results of this annotation are integrated in SWISS-PROT, and a website is provided at http://www.expasy.org/sprot/hamap/.

Amino Acid Sequence↗

The PROSITE database, its status in 2002.

PROSITE [Bairoch and Bucher (1994) Nucleic Acids Res., 22, 3583-3589; Hofmann et al. (1999) Nucleic Acids Res., 27, 215-219] is a method of identifying the functions of uncharacterized proteins translated from genomic or cDNA sequences. The PROSITE database (http://www.expasy.org/prosite/) consists of biologically significant patterns and profiles designed in such a way that with appropriate computational tools it can rapidly and reliably help to determine to which known family of proteins (if any) a new sequence belongs, or which known domain(s) it contains.

Amino Acid Motifs↗

InterPro: an integrated documentation resource for protein families, domains and functional sites.

The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.

Algorithms↗

PROSITE: a documented database using patterns and profiles as motif descriptors.

Among the various databases dedicated to the identification of protein families and domains, PROSITE is the first one created and has continuously evolved since. PROSITE currently consists of a large collection of biologically meaningful motifs that are described as patterns or profiles, and linked to documentation briefly describing the protein family or domain they are designed to detect. The close relationship of PROSITE with the SWISS-PROT protein database allows the evaluation of the sensitivity and specificity of the PROSITE motifs and their periodic reviewing. In return, PROSITE is used to help annotate SWISS-PROT entries. The main characteristics and the techniques of family and domain identification used by PROSITE are reviewed in this paper.

Amino Acid Motifs↗

Bacillus subtilis 168 gene lytF encodes a gamma-D-glutamate-meso-diaminopimelate muropeptidase expressed by the alternative vegetative sigma factor, sigmaD.

A gamma-D-glutamate-meso-diaminopimelate muropeptidase was detected in the vegetative growth phase of Bacillus subtilis 168. It is encoded by the monocistronic lytF operon expressed by the alternative vegetative sigma factor, sigmaD. Sequence analysis of LytF revealed two domains, an organization common to exoproteins of B. subtilis as well as to those from other organisms. The N-terminal domain contains a fivefold-repeated motif attributed to cell wall binding, whilst the C-terminal domain is probably endowed with the catalytic activity. Overexpression of LytF allowed its purification and biochemical characterization. Inactivation of lytF led to the loss of the cell-wall-bound protein 49' (CWBP49') and of the corresponding lytic activity as revealed by renaturation gel assay. Native cell walls prepared from the multiple lytC lytD lytE lytF-deficient mutant did not exhibit any autolysis, whereas walls prepared from a strain endowed with LytF but not with the other three enzymes underwent a slight lysis. Analysis of degradation products of cell wall devoid of teichoic-acid-bound O-esterified D-alanine unambiguously confirmed that LytF cuts the gamma-D-glutamate-mesodiaminopimelate bond.

Amino Acid Sequence↗

Assay for UDPglucose 6-dehydrogenase in phosphate-starved cells: gene tuaD of Bacillus subtilis 168 encodes the UDPglucose 6-dehydrogenase involved in teichuronic acid synthesis.

A novel assay permitting the detection of UDPglucose 6-dehydrogenase activity in cell-free extracts obtained from phosphate-starved cultures of Bacillus subtilis is described. The critical step, the separation of phosphate-starvation-induced exo-enzymes, phosphatases and phosphodiesterases from the cytoplasmic fraction containing the UDPglucose dehydrogenase, was achieved by protoplasting and removal of the periplasmic fraction by protoplast washing. Using this method, the following were unambiguously demonstrated: (i) the presence in the cytoplasm of an enzymic activity oxidizing UDPglucose to UDPglucuronic acid, and (ii) that detection of the activity in whole-cell-free extracts is prevented by the presence of 'periplasmic' enzymes catalysing the degradation of the sugar nucleotides. With this method, several B. subtilis 168 mutants unable to synthesize teichuronic acid were examined. Strains inactivated in gene tuaD, whose product shares homology with UDPglucose 6-dehydrogenase and GDPmannose 6-dehydrogenase from other organisms, were shown to lack UDPglucose 6-dehydrogenase activity. Anion exchange chromatography revealed that mutants deficient in tuaD lacked a cytoplasmic UDPglucuronate pool.

Bacillus subtilis↗