PubMed Health⌕ Search

Biomedical subjects

S Pasupathy

Publications and source records attributed to S Pasupathy.

2 recordsLinked to original sources

Models of local behavior of DNA electrophoresis peak parameters.

Many base calling algorithms implicitly or explicitly rely on predictions of local sequence parameters such as amplitude, peak time and peak width. For example, an algorithm may search for the next peak about a predicted peak time formed by adding the mean peak separation to the last position measurement. In this paper, covariance models are presented which characterize the dependence of peak parameters on those of other peaks. Based on experimental measurements, the model features an exponential decay in peak time jitter covariance with respect to base separation. Both peak amplitude and peak width are modelled as being uncorrelated with those of adjacent bases. In the model, linear expressions are given to describe the growth in peak time jitter and peak width as a function of base position while other parameters, such as amplitude variance, are modeled by constants. Together, these results form a simple model which may be used in the derivation of new sequencing algorithms or in simulations for the testing of such algorithms. We suggest that the correlation of the peak times is related to the Kuhn length of the single-stranded DNA fragments.

Algorithms↗

Optimal structure for automatic processing of DNA sequences.

The faithful recovery of the base sequence in automatic DeoxyriboNucleic Acid (DNA) sequencing fundamentally depends on the underlying statistics of the DNA electrophoresis time series. Current DNA sequencing algorithms are heuristic in nature and modest in their use of statistical information. In this paper, a formal statistical model of the DNA time series is presented and then used to construct the optimal maximum-likelihood (ML) processor. The DNA-ML algorithm that is derived in this paper features Kalman prediction of peak locations, peak parameter estimation, whitened waveform comparison and multiple hypothesis processing using the M-algorithm. Properties of the algorithm are examined using both simulated and real data. Model parameters of critical importance and their impact on different types of error mechanisms, such as insertions and deletions, are pointed out. The statistical model of the DNA time-series and the structure of the DNA-ML algorithm provides a basis for future investigation and refinement of DNA sequencing techniques.

Algorithms↗