PubMed HealthSearch

Biomedical subjects

G A Fichant

Publications and source records attributed to G A Fichant.

4 recordsLinked to original sources

A frameshift error detection algorithm for DNA sequencing projects.

During the determination of DNA sequences, frameshift errors are not the most frequent but they are the most bothersome as they corrupt the amino acid sequence over several residues. Detection of such errors by sequence alignment is only possible when related sequences are found in the databases. To avoid this limitation, we have developed a new tool based on the distribution of non-overlapping 3-tuples or 6-tuples in the three frames of an ORF. The method relies upon the result of a correspondence analysis. It has been extensively tested on Bacillus subtilis and Saccharomyces cerevisiae sequences and has also been examined with human sequences. The results indicate that it can detect frameshift errors affecting as few as 20 bp with a low rate of false positives (no more than 1.0/1000 bp scanned). The proposed algorithm can be used to scan a large collection of data, but it is mainly intended for laboratory practice as a tool for checking the quality of the sequences produced during a sequencing project.

Algorithms

Fast identification of repetitive elements in biological sequences.

We have developed a fast filtering method for searching repetitive sequences in databases that allows the simultaneous identification of different families of repetitive elements during the same scanning. It discriminates between repetitive elements and non-related sequences by comparing the frequencies of k-words found in both groups of sequences. The distance used to sort out the sequences is based on a weighting of the k-words, which is obtained by performing a correspondence analysis on learning sets of correctly chosen sequences. The identification of Alu elements in human sequences is given as an illustration of the method. The Alu sequences are divided in four distinct groups of elements: the left and right monomers located on the direct and on the complementary strands. The results obtained on the test sets show that a very good discrimination is achieved with a word length of 6 b.p. Indeed, only 0.5% of the non-Alu sequences were incorrectly predicted as Alu elements for a threshold value allowing the identification of all Alu monomers. The misclassification of the different Alu monomers (1.4%) in the four groups of examples occurs only when the left and the right monomers are in the same orientation. Moreover, during the scanning of 63 GenBank sequences longer than 10 Kb, all the Alu elements were correctly identified (616 elements) and only a few non-Alu sequences were wrongly predicted as Alu elements (22 fragments). There is a real need for this kind of method since most of the repetitive elements are not annotated in the database entries. This method can then be used for a systematic screening of new sequences before their insertion in databases. It can also allow the creation of specific databases devoted to repetitive elements, which is a required step for any further analysis of those elements.

Algorithms

Constraints acting on the exon positions of the splice site sequences and local amino acid composition of the protein.

The exon positions located at the 5' and 3' splice sites are involved in two functions: in the accurate removal of introns from the nuclear pre-messenger RNAs and in coding for amino acids. Therefore, at least two constraints will act on the exon positions: the splicing constraint and the protein constraint. In the present study we investigate the effect of those constraints on a set of splice sites extracted from GenBank. The consensus matrices computed for each intron location in the reading frame present striking differences at the exon positions of the 5' splice sites. The results obtained can not be explained by the action of a single constraint but rather by the competition between the splicing and protein constraints. Out of eight sites corresponding to codons located in the vicinity of the intron, three present an amino acid distribution that differs greatly from the average amino acid composition of the proteins. Each of these three sites can be characterized by specific amino acids. Results show that the splicing constraint has an effect on the local amino acid composition of the protein as long as the function of the protein is not disrupted.

Amino Acid Sequence

Identifying potential tRNA genes in genomic DNA sequences.

We have developed an algorithm that automatically and reproducibly identifies potential tRNA genes in genomic DNA sequences, and we present a general strategy for testing the sensitivity of such algorithms. This algorithm is useful for the flagging and characterization of long genomic sequences that have not been experimentally analyzed for identification of functional regions, and for the scanning of nucleotide sequence databases for errors in the sequences and the functional assignments associated with them. In an exhaustive scan of the GenBank database, 97.5% of the 744 known tRNA genes were correctly identified (true-positives), and 42 previously unidentified sequences were predicted to be tRNAs. A detailed analysis of these latter predictions reveals that 16 of the 42 are very similar to known tRNA genes, and we predict that they do, in fact, code for tRNA, yielding a false-positive rate for the algorithm of 0.003%. The new algorithm and testing strategy are a considerable improvement over any previously described strategies for recognizing tRNA genes, and they allow detections of genes (including introns) embedded in long genomic sequences.

Algorithms