PubMed HealthSearch

Biomedical subjects

P A Pevzner

Publications and source records attributed to P A Pevzner.

9 recordsLinked to original sources

Statistical distance between texts and filtration methods in sequence comparison.

Upon searching local similarities in long sequences, the necessity of a 'rapid' similarity search becomes acute. Quadratic complexity of dynamic programming algorithms forces the employment of filtration methods that allow elimination of the sequences with a low similarity level. The paper is devoted to the theoretical substantiations of the filtration method based on the statistical distance between texts. The notion of the filtration efficiency is introduced and the efficiency of several filters is estimated. It is shown that the efficiency of the statistical l-tuple filtration upon DNA database search is associated with a potential extension of the original four-letter alphabet and grows exponentially with increasing l. The formula that allows one to estimate the filtration parameters is presented.

Amino Acid Sequence

Extendable words in nucleotide sequences.

Previous statistical analyses revealed several peculiarities of nucleotide sequences that preclude their description by existing models and thus allow one to distinguish DNA and RNA sequences from random A,T,G,C-texts. This is a consequence of the unusual distribution of certain words in nucleotide sequences: while the distribution of (most) words is consistent with Markov models of small orders, the distribution of certain words cannot be described by any previous model (anomalies in distribution of homonucleotide/homopurine/homopyrimidine runs, complementary and mirror palindromes, and non-stationary words). In this work we introduce a probabilistic approach that is partly motivated by analogy with linguistics. We also describe another important feature of DNA/RNA sequences: anomalies in distribution of words of poor nucleotide composition. We show that some classes of these words are the major obstacle for the simple Markov description of nucleotide sequences.

Base Sequence

Improved chips for sequencing by hybridization.

The SHOM method (Sequencing by Hybridization with Oligonucleotide Matrix) developed in 1988 is a new approach to nucleic acid sequencing by hybridization to an oligonucleotide matrix composed of an array of immobilized oligonucleotides. The original matrix proposed for sequencing by SHOM had to contain at least 65,536 octanucleotides. The present work describes a new family of matrices, which allows one to reduce the number of synthesized oligonucleotides 5-15 times without essentially decreasing the resolving power of the method.

Algorithms

Genome inhomogeneity is determined mainly by WW and SS dinucleotides.

According to the hypothesis of the modular structure of DNA, genomes consist of modules of various nature which may differ in statistical characteristics. Statistical analysis helps in revealing the differences in statistical characteristics and predicting the modular structure. In this connection the question about the contribution of each word of length l (l-tuple) to the inhomogeneity of genetic text arises. The notion of stationary (i.e. relatively evenly distributed over a genome) versus non-stationary l-tuples has been introduced previously. In this paper, the dinucleotide distributions for all long sequences from GenBank were analyzed and it was shown that non-stationary dinucleotides are closely associated with polyW and polyS tracts (W denotes 'weak' nucleotides A or T, while S stands for the 'strong' nucleotides G or C). Thus, genome inhomogeneity is shown to be determined mainly by AA, TT, GG, CC, AT, TA, GC and CG dinucleotides. It has been demonstrated that neither 'codon usage' nor the 'isochore model' can account for this phenomenon.

Algorithms

Linguistics of nucleotide sequences. I: The significance of deviations from mean statistical characteristics and prediction of the frequencies of occurrence of words.

Mathematical models of the generation of genetic texts appeared simultaneously with the first sequencing DNA. They are used to establish functional and evolutionary relations between genetic texts, to predict the number and distribution of specific sites in a sequence and to identify "meaningful" words. The present paper deals with two problems: 1) The significance of deviations from the mean statistical characteristics in a genetic text. Anyone who has addressed himself to the statistical analysis of sequenced DNA is familiar with the question: what deviations from the expected frequencies of occurrence of particular words testify to the "biological" significance of those words? We propose a formula for the variance of the number of word's occurrences in the text, with allowance for word overlaps, making it possible to assess the significance of the deviations from the expected statistical characteristics. 2) A new method for predicting the frequencies of occurrence of particular words in a genetic text using the statistical characteristics of "spaced" L-grams. The method can be used for predicting the number of restriction sites in human DNA and in planning experiments on the physical mapping and sequencing of the human genome.

Bacteriophage lambda

Linguistics of nucleotide sequences. II: Stationary words in genetic texts and the zonal structure of DNA.

Words are irregularly distributed in genetic texts. The analysis of this irregularity leads to the notion of stationary and non-stationary words. The polyW and polyS tracts are shown to be the most non-stationary words in genetic texts (here W-[A,T], S-[G,C], a polyW tract is a sequence of A,T nucleotides and a polyS tract is a sequence of G,C nucleotides. The distribution of stationary words suggests a method for partitioning DNA into zones. The zones obtained in the case of the phage are interpreted in the light of the Dowe hypothesis of the modular structure of bacteriophage genomes.

Adenoviridae

1-Tuple DNA sequencing: computer analysis.

A new method of DNA reading was proposed at the end of 1988 by Lysov et al. According to the authors' claims it has certain advantages as compared to the Maxam-Gilbert and Sanger methods, which are revealed by automation and rapidity of DNA sequencing. Nevertheless its employment is hampered by a number of biological and mathematical problems. The present study proposes an algorithm that allows to overcome the computational difficulties occurring in the course of the method during reconstruction of the DNA sequence by its l-tuple composition. It is shown also that the biochemical problems connected with the loss of information about the l-tuple DNA composition during hybridization are not crucial and can be overcome by finding the maximal flow of minimal cost in the special graph.

Algorithms

[Optimal chips for megabase DNA sequencing].

The SHOM method (Sequencing by Hybridization with Oligonucleotide Matrix) developed in 1988 is a new approach to nucleic acid sequencing by hybridization to a octanucleotide matrix composed of an array of immobilized oligonucleotides. The original matrix proposed for sequencing by SHOM had to contain at least 65,536 octanucleotides. The present work describes a new family of matrices for sequencing, which allows one to reduce the number of synthesized oligonucleotides 5-15 times without essentially decreasing the resolving power of the method.

Base Sequence

[Paths in graphs and selection of oligonucleotide linkers].

The problem of search of an universal linker with minimal length containing all restriction endonuclease recognition sites is considered. We reduce this problem to the search of Eiler's and Hamilton's paths in graph. The use of the discrete optimization methods allows to construct the linker with the length closed to the minimum.

Base Sequence