PubMed Health⌕ Search

Biomedical subjects

J J Codani

Publications and source records attributed to J J Codani.

5 recordsLinked to original sources

Applications of the pyramidal clustering method to biological objects.

In conventional hierarchical clustering methods, any object can belong to only one class or cluster. We present here an application of the pyramidal classification method to biological objects, which illustrates the intuitively appealing idea that some objects may belong simultaneously to two classes. In a first step, we performed an all-by-all comparison of all the open reading frames in the genomes from S. cerevisiae, M. jannaschii, E. coli, H. influenzae and Synechocystis. In a second step, a series of connex classes was built, each connex class containing all those sequences that were linked by a Z-value (obtained after 100 sequence shufflings) greater than a given threshold. Finally, each connex class was submitted to a pyramidal classification. Three examples of such classifications are given, concerning two sets of multi-domains protein sequences and a family of aminoacyl-tRNA synthetases. They make it clear that the linear order among the classified objects that results from the pyramidal classification is useful in deciphering the multiple relationships that can exist between the objects under study. A program for calculating and displaying a pyramidal classification from a dissimilarity matrix is available from http:/(/)genome.genetique.uvsq.fr/Pyramids. The pyramidal classifications of the connex classes from the five organisms (intra- and inter-genomic comparisons) are available from http:/(/)www.gene-it.com under the family item.

Algorithms↗

Significance of Z-value statistics of Smith-Waterman scores for protein alignments.

The Z-value is an attempt to estimate the statistical significance of a Smith-Waterman dynamic alignment score (SW-score) through the use of a Monte-Carlo process. It partly reduces the bias induced by the composition and length of the sequences. This paper is not a theoretical study on the distribution of SW-scores and Z-values. Rather, it presents a statistical analysis of Z-values on large datasets of protein sequences, leading to a law of probability that the experimental Z-values follow. First, we determine the relationships between the computed Z-value, an estimation of its variance and the number of randomizations in the Monte-Carlo process. Then, we illustrate that Z-values are less correlated to sequence lengths than SW-scores. Then we show that pairwise alignments, performed on 'quasi-real' sequences (i.e., randomly shuffled sequences of the same length and amino acid composition as the real ones) lead to Z-value distributions that statistically fit the extreme value distribution, more precisely the Gumbel distribution (global EVD, Extreme Value Distribution). However, for real protein sequences, we observe an over-representation of high Z-values. We determine first a cutoff value which separates these overestimated Z-values from those which follow the global EVD. We then show that the interesting part of the tail of distribution of Z-values can be approximated by another EVD (i.e., an EVD which differs from the global EVD) or by a Pareto law. This has been confirmed for all proteins analysed so far, whether extracted from individual genomes, or from the ensemble of five complete microbial genomes comprising altogether 16956 protein sequences.

Computing Methodologies↗

Removing redundancy in SWISS-PROT and TrEMBL.

SUMMARY: One of the distinguishing criteria of the SWISS-PROT protein sequence data bank is minimal redundancy. The introduction of TrEMBL as a supplementary database ensured the comprehensiveness of SWISS-PROT and TrEMBL but introduced some degree of redundancy. We developed a strategy to identify the redundancy present within and between SWISS-PROT and TrEMBL and its subsequent removal. AVAILABILITY: The tools mentioned in this paper are available on request.

Algorithms↗

The complete genome sequence of the gram-positive bacterium Bacillus subtilis.

Bacillus subtilis is the best-characterized member of the Gram-positive bacteria. Its genome of 4,214,810 base pairs comprises 4,100 protein-coding genes. Of these protein-coding genes, 53% are represented once, while a quarter of the genome corresponds to several gene families that have been greatly expanded by gene duplication, the largest family containing 77 putative ATP-binding transport proteins. In addition, a large proportion of the genetic capacity is devoted to the utilization of a variety of carbon sources, including many plant-derived molecules. The identification of five signal peptidase genes, as well as several genes for components of the secretion apparatus, is important given the capacity of Bacillus strains to secrete large amounts of industrially important enzymes. Many of the genes are involved in the synthesis of secondary metabolites, including antibiotics, that are more typically associated with Streptomyces species. The genome contains at least ten prophages or remnants of prophages, indicating that bacteriophage infection has played an important evolutionary role in horizontal gene transfer, in particular in the propagation of bacterial pathogenesis.

Bacillus subtilis↗

LASSAP, a LArge Scale Sequence compArison Package.

MOTIVATION: This paper presents LASSAP, a new software package for sequence comparison. LASSAP is a programmable, high-performance system designed to raise current limitations of sequence comparison programs in order to fit the needs of large-scale analysis. LASSAP provides an API (Application Programming Interface) allowing the integration of any generic pairwise-based algorithm. RESULTS: Whatever pairwise algorithm is used in LASSAP, it shares with all other algorithms numerous enhancements such as: (i) intra- and inter-databank comparisons; (ii) computational requests (selections and computations are achieved on the fly); (iii) frame translations on queries and databanks; (iv) structured results allowing easy and powerful post-analysis; (v) performance improvements by parallelization and the driving of specialized hardware. LASSAP currently implements all major sequence comparison algorithms (Fasta, Blast, Smith/Waterman), and other string matching and pattern matching algorithms. LASSAP is both an integrated software for end-users and a framework allowing the integration and the combination of new algorithms. LASSAP is used in different projects such as the building of PRODOM, the exhaustive comparison of yeast sequences, and the subfragments matching problem of TREMBL.

Algorithms↗