PubMed HealthSearch

Biomedical subjects

M C Giddings

Publications and source records attributed to M C Giddings.

3 recordsLinked to original sources

A software system for data analysis in automated DNA sequencing.

Software for gel image analysis and base-calling in fluorescence-based sequencing consisting of two primary programs, BaseFinder and GelImager, is described. BaseFinder is a framework for trace processing, analysis, and base-calling. BaseFinder is highly extensible, allowing the addition of trace analysis and processing modules without recompilation. Powerful scripting capabilities combined with modularity and multilane handling allow the user to customize BaseFinder to virtually any type of trace processing. We have developed an extensive set of data processing and analysis modules for use with the program in fluorescence-based sequencing. GelImager is a framework for gel image manipulation. It can be used for gel visualization, lane retracking, and as a front end to the Washington University Getlanes program. The programs were designed using a cross-platform development environment, currently allowing them to run in Windows NT, Windows 95, Openstep/Mach, and Rhapsody. Work is ongoing to deploy the software on additional platforms, including Solaris, Linux, and MacOS. This software has been thoroughly tested and debugged in the analysis of >2 million bp of raw sequence data from human chromosome 19 region q13. Overall sequencing accuracy was measured using a significant subset of these data, consisting of approximately 600 sequences, by comparing the individual shotgun sequences against the final assembled contigs. Also, results are reported from experiments that analyzed the accuracy of the software and two other well-known base-calling programs for sequencing the M13mp18 vector sequence. [The sequence data described in this paper have been submitted to the GenBank data library under accession no. AF025422]

Algorithms

Automatic matrix determination in four dye fluorescence-based DNA sequencing.

The four dye fluorescence detection strategy is a widely used approach to automated DNA sequence analysis. An important aspect of data processing in this approach is the multicomponent analysis to deduce the concentrations of four fluorophores from fluorescence emission intensities at four different wavelengths. This requires knowledge of the correct transformation matrix M. The matrix M is a function both of the fluorophores employed and the fluorescence detection system. M is typically determined either by a calibration process with individual dyes, or by choosing four well-separated individual peaks corresponding to the four different dyes. Both are time-consuming and complicated procedures for routine use. An automatic scheme for finding M directly from raw sequence data is presented here. This facilitates data analysis and the underlying algorithm may also find utility in other multispectral applications.

Algorithms

An adaptive, object oriented strategy for base calling in DNA sequence analysis.

An algorithm has been developed for the determination of nucleotide sequence from data produced in fluorescence-based automated DNA sequencing instruments employing the four-color strategy. This algorithm takes advantage of object oriented programming techniques for modularity and extensibility. The algorithm is adaptive in that data sets from a wide variety of instruments and sequencing conditions can be used with good results. Confidence values are provided on the base calls as an estimate of accuracy. The algorithm iteratively employs confidence determinations from several different modules, each of which examines a different feature of the data for accurate peak identification. Modules within this system can be added or removed for increased performance or for application to a different task. In comparisons with commercial software, the algorithm performed well.

Algorithms