PubMed HealthSearch

Biomedical subjects

C Schwager

Publications and source records attributed to C Schwager.

At least 19 recordsLinked to original sources

Efficient low redundancy large-scale DNA sequencing at EMBL.

An efficient low redundancy DNA sequencing strategy should allow high accuracy determination of the consensus sequence on both strands of a DNA fragment from a minimal number of sequencing reactions with minimal overlap. At EMBL we developed a directed strategy for cosmid-scale sequencing based on primer walking, whereas most other sequencing projects of this scale rely on the random 'shotgun' strategy. In our strategy, highly accurate raw data are obtained from automated double-stranded Sanger dideoxy sequencing with inexpensive walking primers (8 to 10 $ per primer), T7 DNA polymerase and internal labelling by fluorescein-15- dATP on A.L.F. DNA sequencers (Pharmacia Biotech). The use of 60-cm long glass plates enables reading length of up to 1000 bases. Comparing various random and directed sequencing strategies in the course of the European Community yeast genome sequencing project on cosmids from chromosomes IX, XI and XV, primer walking was found to be the strategy resulting in the lowest possible redundancy of 2.6 to 2.8. Future development of the sequencing strategy is based on the new EMBL 2-dye sequencing device for simultaneous sequencing on both strands, and implementation of an initial limited random sequencing phase to reduce the number of walking primers required by a factor of 3, while still maintaining a low redundancy of approx. 3.

Animals

Simultaneous on-line DNA sequencing on both strands with two fluorescent dyes.

We describe an automated DNA-sequencing technique which allows both the simultaneous sequencing from the two strands of double-stranded templates and the subsequent detection of the sequencing products online and in parallel. The technique is based on hardware technology also used in the ALF DNA sequencer (Pharmacia, Uppsala). A helium-neon laser was mounted into the sequencing device additionally to the standard argon laser. Two different primers, labeled with either fluorescein or Texas red, are used in a single sequencing reaction resulting in an output of two sequences. Both sequencing products are then analyzed on-line in the same lanes of a gel. This technique is especially useful for the complete sequencing of DNA fragments up to 1 kb. High accuracy sequencing of PCR products in double-stranded form can now be accomplished in only one sequencing reaction.

Fluorescent Dyes

Nucleotide sequence and analysis of the centromeric region of yeast chromosome IX.

We have determined the nucleotide sequence of a cosmid (pIX338) containing the centromere region of yeast (Saccharomyces cerevisiae) chromosome IX. The complete nucleotide sequence of 33.8 kb was obtained by using an efficient directed sequencing strategy in combination with automated DNA sequencing on the A.L.F. DNA sequencer. Sequence analysis revealed the presence of 17 open reading frames (ORFs), four of them previously known yeast genes (sly12, pan1, sts1 and prl1), a tRNA gene and the centromere motif. Exhaustive database searches detected sequence homologues of known function for as many as 14 of the 17 ORFs. These include a mammalian tyrosine kinase substrate; the Escherichia coli cell cycle protein MinD; the human inositol polyphosphate-5-phosphatase (gene OCRL) involved in Lowe's syndrome, a developmental disorder; and helicases, for which the new yeast member defines a distinct DEAD/H-box subfamily. A surprisingly large fraction of the ORFs (at least six out of 17) in the centromeric region are apparently involved in RNA or DNA binding.

Adenosine Triphosphatases

Primer design for automated DNA sequencing utilizing T7 DNA polymerase and internal labeling with fluorescein-15-dATP.

We have identified additional criteria for the walking primer design that improve the success rate of automated fluorescent DNA sequencing using the internal labeling technique and T7 DNA polymerase. These criteria resulted from the evaluation of over 220 sequences generated with walking primers and fluorescein-15-dATP as internal label in the course of the European Community (EC) yeast genome sequencing project. In this project primers were designed using standard commercial software. Intensities of sequencing signals varied over a broad range from very strong to very weak, depending on the primers used. This led us to evaluate primer performance relative to (i) the template sequence immediately downstream of the primer binding site and (ii) the primer sequence itself. Our experiments show that the position of the first labeled dATP to be incorporated downstream of the primer into the growing strand is substantial for the signal intensity of the sequence. The closer to the primer that the first 'A' is incorporated, the stronger the peak intensities are. An additional feature of sequencing with native T7 DNA polymerase is its ability to remove a 3'-terminal 'A' of the primer by the 3'-->5' exonuclease activity and to exchange the nucleotide with a labeled dATP by the polymerase activity.

Binding Sites

Structure of the gene encoding human casein kinase II subunit beta.

Casein kinase II (CKII) is a ubiquitous serine/threonine protein kinase with numerous key functions in cell metabolism and growth. The human CKII has a tetrameric structure; two catalytic subunits (alpha and alpha') form the holoenzyme together with two presumably regulatory subunits (beta). The gene encoding CKII subunit beta was isolated from human genomic DNA and analyzed for its primary structure using exclusively nonradioactive procedures. The gene was found to span 4.2 kilobase pairs and to be composed of seven exons. Exon sizes range from 76 (exon 5) to 329 base pairs (bp) (exon 1), intron sizes from 145 (intron V) to 965 bp (intron II). All exon-intron junctional sequences conform to the canonical GT-AG rule. Primer extension analysis determined three transcription initiation sites, at 951, 919, and (minor) 840 bp upstream of the translation start site. The translation start is located early in the second exon; exon 1 is untranslated. The 3'-cleavage/polyadenylation signal sequence (AA-TAAA) is in the last exon at position 4173 bp relative to the first transcription initiation site. The coding sequence for CKII beta comprises 648 nucleotides identical to the published CKII beta-cDNA sequence (Jakobi, R., Voss, H., and Pyerin, W. (1989) Eur. J. Biochem. 183, 227-233). The upstream promoter region of the CKII beta gene contains multiple potential gene regulatory sequence elements, noticeable DNA structures, and the characteristics of a housekeeping gene (more than one transcription initiation site, lack of a TATA-box, presence of a CpG island, occurrence of multiple GC boxes and of nonstandard positioned CCAAT boxes). The CKII beta gene promoter shares common features with that of mammalian protein kinases and is closely related to the regulatory subunit gene promoter of cAMP-dependent protein kinase.

Amino Acid Sequence

Automated DNA sequencing of the human HPRT locus.

The complete sequence of 57 kb of the human HPRT locus has been determined using automated fluorescent DNA sequencing. The strategy employed increasingly directed sequencing methods: A randomly generated M13 library was sequenced to generate contiguous overlapping sets of sequences (contigs). M13 clones at the ends of these contigs were further sequenced using M13 (universal and reverse) and custom oligonucleotide primers to order the contigs and to complete the sequencing project. The human HPRT sequence includes 1676 bp 5' and 15,238 bp 3' to exons 1 and 9, respectively. The sequence contains 49 representatives of the Alu repeat, along with several other types of repetitive sequences. The Alu sequences exhibit a biased orientation, with those sequences in the first half of the locus oriented in the minus direction relative to transcription of the gene (3'----5' = 77%, P less than 0.005) and those sequences in the latter half of the locus oriented randomly (5'----3' = 67%, P less than 0.5). The development and performance of the sequencing strategy and the features of the human HPRT gene are presented.

Amino Acid Sequence

Automated sequencing of fluorescently labelled DNA by chemical degradation.

A new general method for sequencing fluorescently labelled DNA by chemical degradation has been developed. It is based on the observation that fluorescein attached via a mercaptopropyl or aminopropyl linker arm to the 5'-phosphate of an oligonucleotide is stable during the reactions commonly used in chemical cleavage procedures. DNA to be degraded is first enzymatically synthesized in vitro by annealing and extending a fluorescently labelled primer thereby introducing the fluorescent label at the 5'-end of the fragment. The newly synthesized fluorescently labelled DNA is then chemically degraded using: (a) a set of four different cleavage reactions; or (b) only one reaction comprising methylation of G-residues followed by a partial cleavage with piperidine in the presence of sodium chloride. The fluorescent degradation products are loaded on either four lanes or one lane of the gel, respectively, and the emitted fluorescence detected online during electrophoresis. In the 'four reactions/four lanes' method 200-350 bp (base pairs) can be read from the labelled end. The 'one reaction/one lane' method, in which the nucleotide sequence is determined by measuring different signal intensities following the rule G greater than A greater than C greater than T, currently yields around 100-200 bp of sequence per sample.

Automation

Direct genomic fluorescent on-line sequencing and analysis using in vitro amplification of DNA.

In vitro amplification of genomic DNA and total RNA, as well as recombinant DNA, using one fluorescently labelled and one unlabelled primer during amplification, together with on-line analysis of the products on the EMBL fluorescent DNA sequencer, is described. Further is reported direct sequencing of fluorescently labelled amplified probes by solid-phase chemical degradation, without subcloning and purification steps involved. At present up to 350 bases in 4 hours are determined with this technique. The fluorescent dye and its bond to the oligonucleotide are stable during the amplification cycles, and do not interfere with the enzymatic polymerization. High sensitivity of the detection device, down to 10(-18) moles, corresponding to less than 10(6) molecules makes possible analyses of the non-radioactive amplified probes after only 10 amplification cycles, starting with about 5 x 10(4) copies of recombinant DNA.

Autoanalysis

Automated Sanger DNA sequencing with one label in less than four lanes on gel.

Novel Sanger dideoxy sequencing with only one fluorescent dye label for the four bases of one clone and sequence determination in two lanes on polyacrylamide gel is presented, loading A greater than G in one lane and T greater than C in the other. Sequencing reactions for the two bases in each lane are carried out in one tube. At present the ratio of ddATP:ddGTP and ddTTP:ddCPT is set to 5:1 in the two tubes. Distinction between the two bases in one lane is done by comparing the different magnitudes of the peaks. This method increases the capacity since more clones may be run simultaneously on one gel, while keeping the reliability and simplicity that comes with the use of only one fluorescent dye for the four bases of one clone. At present about 200 bases are determined with the one-dye two-lane method on the EMBL's automated fluorescent DNA sequencer, using T7 DNA polymerase. The error rate in the deduced sequence is about 1%. The technique is used for the determination of overlaps in mapping projects. In principle, it is possible to determine the sequence with one dye in only one lane on the gel by choosing the proper ddNTP ratios for all four bases, carrying out reactions in one tube and applying the product in one lane, but the error rate for this one-lane method seems too high at present and further improvements in the uniformity of peaks obtainable with the T7 DNA polymerase or other enzymes are required.

Base Sequence

Automated Sanger dideoxy sequencing reaction protocol.

The protocol for Sanger dideoxy chain termination reactions in DNA sequencing is tedious and prone to errors due to the repetitive character of the pipetting steps. An industrial robot, with the addition of a few simple parts, was programmed to automate the dideoxy sequencing reactions. The system is set up in a short time for routine operation and it is faster and more reliable than a human operator. It is flexible and allows variations and optimization of the standard procedure. Disposable microtiter plates at a controlled temperature are used. In one reaction cycle (about 50 min) up to 48 templates are processed. Up to 450 bases were resolved in automated DNA sequencing on samples prepared by the robot. The protocol is applicable to fluorescent as well as to radioactive labeling.

Autoanalysis

T7 DNA polymerase in automated dideoxy sequencing.

T7 DNA polymerase with chemically inactivated 3'-5' exo-nuclease activity, as well as unmodified T7 DNA polymerase, were used for sequencing by the dideoxy method in an automated system with fluorescence labelled primer and on-line detection of laser-excited reaction products. An analysis of signal intensity variations in the C track revealed that low C signals were usually preceded by a T in the sequence. This effect was modified by surrounding nucleotides. Signal intensities were more uniform with T7 polymerase than with the Klenow fragment of DNA polymerase I. Some sequences ambiguous with the Klenow enzyme could easily be evaluated with the T7 enzyme. One sequence could only be read by the unmodified T7 polymerase, while both the Klenow fragment and the chemically modified T7 enzyme gave uninterpretable data.

Base Sequence

Non-radioactive automated sequencing of oligonucleotides by chemical degradation.

A non-radioactive sequencing of fluorescently labelled oligonucleotides by solid-phase chemical degradation is described. Although non-radioactive methods have been reported for the dideoxy chain termination technique, such a method has not yet been developed for the chemical degradation sequencing of DNA fragments. A 21-mer fluorescein labelled M13 sequencing primer was sequenced in an on-line automated system in about 30 minutes. The fluorescent dye and its bond to the oligonucleotide were stable during the chemical reactions used for the base specific degradations. As the sequence is determined on-line during electrophoresis, reloading and running 10 fragments simultaneously allows us to use one gel for sequencing of about 50 different oligonucleotides.

Automation

Two different mRNAs are transcribed from a single genomic locus encoding the chicken erythrocyte anion transport proteins (band 3).

The chicken erythrocyte anion transport protein (band 3 of the erythrocyte cytoskeleton) is a central component taking part in two widely divergent functions of erythroid cells; it is a primary determinant of cytoskeletal architecture and responsible for electroneutral Cl-/HCO3- exchange across the plasma membrane. To analyze interesting aspects of the developmental regulation of this gene, we have cloned the cDNA and genomic counterparts of the erythroid-specific anion transport protein. We show that a single genetic locus for band 3 encodes two different erythroid cell-specific mRNAs, with different translational initiation sites, which predict polypeptides of sizes very close to those observed in vivo. In vitro translation and immune precipitation of synthetic mRNA derived from one putative fully encoding cDNA clone demonstrate that this clone gives rise to a protein which is identical in size and antigenicity to bona fide chicken erythroid band 3.

Amino Acid Sequence