PubMed Health⌕ Search

Biomedical subjects

P Nicodème

Publications and source records attributed to P Nicodème.

5 recordsLinked to original sources

Proteome analysis based on motif statistics.

MOTIVATION: Even for the amino acid motifs collected in the Prosite database there may be chance occurences as opposed to those occurences where the motif is involved in fold or function of a protein. With recent mathematical advances in assessing the significance of observing such a motif a particular number of times, we can now study the over- or under-representation of particular motifs in a complete genome and attempt to make functional deductions. RESULTS: We demonstrate that statistical over- or under-representation of motifs in complete proteomes may be an indicator of whether, in that organism, we are looking at chance occurrences of the motif or whether the occurrences are sufficiently numerous to suggest a systematic, and thus functionally important occurrence. This has important implications on databank annotations. AVAILABILITY: The complete dataset comprising the plotted statistics of 266 Prosite motifs on 42 proteomes is available at http://algo.inria.fr/nicodeme/proteomes/proteocomp.html. The software used to compute this data has been described by Nicodème (2000, 2001). They are available either by web access as mentioned in these articles or by direct request from Pierre Nicodème.

Amino Acid Motifs↗

Fast approximate motif statistics.

We present in this article a fast approximate method for computing the statistics of a number of non-self-overlapping matches of motifs in a random text in the nonuniform Bernoulli model. This method is well suited for protein motifs where the probability of self-overlap of motifs is small. For 96% of the PROSITE motifs, the expectations of occurrences of the motifs in a 7-million-amino-acids random database are computed by the approximate method with less than 1% error when compared with the exact method. Processing of the whole PROSITE takes about 30 seconds with the approximate method. We apply this new method to a comparison of the C. elegans and S. cerevisiae proteomes.

Amino Acid Motifs↗

WWW access to the SYSTERS protein sequence cluster set.

SUMMARY: We present a Web server where the SYSTERS cluster set of the non-redundant protein database consisting of sequences from SWISS-PROT and PIR is being made available for querying and browsing. The cluster set can be searched with a new sequence using the SSMAL search tool. Additionally, a multiple alignment is generated for each cluster and annotated with domain information from the Pfam protein family database. AVAILABILITY: The server address is http://www.dkfz-heidelberg.de/tbi/services/cluster/ systersform

Algorithms↗

SSMAL: similarity searching with alignment graphs.

MOTIVATION: We want to provide biologists with a fast and sensitive scanning tool for searching local alignments of a protein query sequence against databases of protein multiple alignments, such as ProDom. Conversely, we want to provide a tool for locally aligning a protein multiple alignment query against a protein database such as SWISSPROT. RESULTS: We developed the program SSMAL (Shuffling Similarities with Multiple Alignments) which utilizes features of the Blast (Altschul et al., J. Mol. Biol., 215, 403-410, 1990) algorithm and part of the Blast code. Our software allows both scanning of multiple alignments and searching with a multiple alignment. Deletions in the multiple alignment only are handled and a SSMAL search may miss some similarities found by a profile search. However, an SSMAL scan of a database such as ProDom would be 20-30 times faster that a profile scan. In the worst case, a SSMAL search is approximately 9 times faster than a profile search. AVAILABILITY: http://www.dkfz-heidelberg.de/tbi/ people/nicodeme and follow the hyperlink SSMAL. CONTACT: p.nicodeme@DKFZ-Heidelberg.de

Amino Acid Sequence↗

Selecting optimal oligonucleotide primers for multiplex PCR.

We investigate the problem of designing efficient multiplex PCR for medical applications. We show that the problem is NP-complete by transformation to the Multiple Choice Matching problem and give an efficient approximation algorithm. We developed this algorithm in a computer program that predicts which genomic regions may be simultaneously amplified by PCR. Practical use of the software shows that the method can treat 250 non-polymorphic loci with less than 5 simultaneous experiments.

Algorithms↗