PubMed Health⌕ Search

Biomedical subjects

S Schbath

Publications and source records attributed to S Schbath.

9 recordsLinked to original sources

Numerical comparison of several approximations of the word count distribution in random sequences.

The exact distribution of word counts in random sequences and several approximations have been proposed in the past few years. The exact distribution has no theoretical limit but may require prohibitive computation time. On the other hand, approximate distributions can be rapidly calculated but, in practice, are only accurate under specific conditions. After making a survey of these distributions, we compare them according to both their accuracy and computational cost. Rules are suggested for choosing between Gaussian approximations, compound Poisson approximation, and exact distribution. This work is illustrated with the detection of exceptional words in the phage Lambda genome.

Bacteriophage lambda↗

Compound Poisson and Poisson process approximations for occurrences of multiple words in Markov chains.

We derive a Poisson process approximation for the occurrences of clumps of multiple words and a compound Poisson process approximation for the number of occurrences of multiple words in a sequence of letters generated by a stationary Markov chain. Using the Chen-Stein method, we provide a bound on the error in the approximations. For rare words, these errors tend to zero as the length of the sequence increases to infinity. Modeling a DNA sequence as a stationary Markov chain, we show as an application that the compound Poisson approximation is efficient for the number of occurrences of rare stem-loop motifs.

Base Sequence↗

An efficient statistic to detect over- and under-represented words in DNA sequences.

In this note, we point out a very efficient statistic to detect over- and under-represented words in DNA sequences, when Markov chain models are used to represent the sequences. This statistic is missing from the recent review done on this important problem and appears to be a better measure of rarity and abundance of words in DNA sequences.

Algorithms↗

Coverage processes in physical mapping by anchoring random clones.

The aim of this paper is to provide general results for predicting progress in a physical mapping project by anchoring random clones, when clones and anchors are not homogeneously distributed along the genome. A complete physical map of the DNA of an organism consists of overlapping clones spanning the entire genome. Several schemes can be used to construct such a map, depending on the way that clones overlap. We focus here on the approach consisting of assembling clones sharing a common random short sequence called an anchor. Some mathematical analyses providing statistical properties of anchored clones have been developed in the stationary case. Modeling the clone and anchor processes as nonhomogeneous Poisson processes provides such an analysis in a general nonstationary framework. We apply our results to two natural nonhomogeneous models to illustrate the effect of inhomogeneity. This study reveals that using homogeneous processes for clones and anchors provides an overly optimistic assessment of the progress of the mapping project.

Cloning, Molecular↗

Exceptional motifs in different Markov chain models for a statistical analysis of DNA sequences.

Identifying exceptional motifs is often used for extracting information from long DNA sequences. The two difficulties of the method are the choice of the model that defines the expected frequencies of words and the approximation of the variance of the difference T(W) between the number of occurrences of a word W and its estimation. We consider here different Markov chain models, either with stationary or periodic transition probabilities. We estimate the variance of the difference T(W) by the conditional variance of the number of occurrences of W given the oligonucleotides counts that define the model. Two applications show how to use asymptotically standard normal statistics associated with the counts to describe a given sequence in terms of its outlying words. Sequences of Escherichia coli and of Bacillus subtilis are compared with respect to their exceptional tri- and tetranucleotides. For both bacteria, exceptional 3-words are mainly found in the coding frame. E. coli palindrome counts are analyzed in different models, showing that many overabundant words are one-letter mutations of avoided palindromes.

Bacillus subtilis↗

Characteristics of Chi distribution on different bacterial genomes.

The availability of full genome sequences provides the bases for analyzing global properties of the genetic text. For example, oligonucleotide sequences that are over- or underrepresented can be identified by taking into account the overall genome composition and organization. One of the most overrepresented oligonucleotides in Escherichia coli is the Chi site, an octanucleotide that stimulates DNA repair by homologous recombination. Here we analyze the genomic distribution of Chi in E. coli and in the three other bacteria where a Chi sequence has been identified; note that Chi is a different sequence in each organism. For each bacterial genome, Chi sequences are frequent, regularly distributed, and overrepresented. This suggests that selection for Chi may have occurred during evolution to favor efficient repair of a damaged chromosome. Other characteristics of Chi distribution are not conserved and might reflect specific features of DNA repair in each host. The different sequence and characteristics of Chi in each microorganism suggest that selection for Chi occurred independently in different bacteria.

Bacillus subtilis↗

Probabilistic and statistical properties of words: an overview.

In the following, an overview is given on statistical and probabilistic properties of words, as occurring in the analysis of biological sequences. Counts of occurrence, counts of clumps, and renewal counts are distinguished, and exact distributions as well as normal approximations, Poisson process approximations, and compound Poisson approximations are derived. Here, a sequence is modelled as a stationary ergodic Markov chain; a test for determining the appropriate order of the Markov chain is described. The convergence results take the error made by estimating the Markovian transition probabilities into account. The main tools involved are moment generating functions, martingales, Stein's method, and the Chen-Stein method. Similar results are given for occurrences of multiple patterns, and, as an example, the problem of unique recoverability of a sequence from SBH chip data is discussed. Special emphasis lies on disentangling the complicated dependence structure between word occurrences, due to self-overlap as well as due to overlap between words. The results can be used to derive approximate, and conservative, confidence intervals for tests.

Base Sequence↗

The effect of nonhomogeneous clone length distribution on the progress of an STS mapping project.

We provide both theoretical and simulation results on the progress of an STS mapping project in the presence of clone length inhomogeneity. For an example in which the genome comprises alternating regions of clones with short and long average length, the main conclusion is that the efficiency of the project is clearly decreased in the presence of such inhomogeneity. The case of deterministic clone length gives the worst progress. The general simulation algorithm we propose shows that strategies that space the anchors as regularly as possible do best: fewer contigs of larger average length are expected. The simulation algorithm can be used to study many statistical properties of the progress of any anchoring project.

Algorithms↗

An overview on the distribution of word counts in Markov chains.

In this paper, we give an overview about the different results existing on the statistical distribution of word counts in a Markovian sequence of letters. Results concerning the number of overlapping occurrences, the number of renewals and the number of clumps will be presented. Counts of single words and also multiple words are considered. Most of the results are approximations as the length of the sequence tends to infinity. We will see that Gaussian approximations switch to (compound) Poisson approximations for rare words. Modeling DNA sequences or proteins by stationary Markov chains, these results can be used to study the statistical frequency of motifs in a given sequence.

Biometry↗