PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Sequence of the cDNA and 5'-flanking region for human acid alpha-glucosidase, detection of an intron in the 5' untranslated leader sequence, definition of 18-bp polymorphisms, and differences with previous cDNA and amino acid sequences.

Acid maltase or acid alpha-glucosidase (GAA) is a lysosomal enzyme that hydrolyzes glycogen to glucose and is deficient in glycogen storage disease type II. Previously, we isolated a partial cDNA (1.9 kb) for human GAA; we have now used this cDNA to isolate and determine sequence in longer cDNAs from four additional independent cDNA libraries. Primer extension studies indicated that the mRNA extended approximately 200 bp 5' of the cDNA sequence obtained. Therefore, we isolated a genomic fragment containing 5' cDNA sequences that overlapped the previous cDNA sequence and extended an additional 24 bp to an initiation codon within a Kozak consensus sequence. The sequence of the genomic clone revealed an intron-exon junction 32 bp 5' to the ATG, indicating that the 5' leader sequence was interrupted by an intron. The remaining 186 bp of 5' untranslated sequence was identified approximately 3 kb upstream. The promoter region upstream from the start site of transcription was GC rich and contained areas of homology to Sp1 binding sites but no identifiable CAAT or TATA box. The combined data gave a nucleotide sequence of 2,856 bp for the coding region from the ATG to a stop codon, predicting a protein of 952 amino acids. The 3' untranslated region contained 555 bp with a polyadenylation signal at 3,385 bp followed by 16 bp prior to a poly(A) tail. This sequence of the GAA coding region differs from that reported by Hoefsloot et al. (1988) in three areas that change a total of 42 amino acids. Direct determination of the amino acid sequence in one of these areas confirmed the nucleotide sequence reported here but also disagreed with the directly determined amino acid sequence reported by Hoefsloot et al. (1988). At two other areas, changes in base pairs predicted new restriction sites that were identified in cDNAs from several independent libraries. The amino acid changes in all three ares increased the homology to rabbit-human isomaltase. Therefore, we believe that our nucleotide sequence for GAA is more precise. We have also identified single base-pair polymorphisms at 18 sites for human GAA, some of which are not silent.

Amino Acid Sequence

Defining the sequence recognized with BmFTZ-F1, a sequence specific DNA binding factor in the silkworm, Bombyx mori, as revealed by direct sequencing of bound oligonucleotides and gel mobility shift competition analysis.

BmFTZ-F1 is a Bombyx mori homologue of FTZ-F1, a positive regulator of the fushi tarazu gene of Drosophila melanogaster. In order to determine the sequence recognized with this factor, we made three sets of oligonucleotide mixture which contain 4 possible nucleotides at different positions within the previously proposed 12-bp binding consensus sequence. Oligonucleotides which bound to purified BmFTZ-F1 were separated by a gel mobility shift procedure and a binding sequence was determined by direct sequencing through Maxam-Gilbert method. By this analysis, 7 positions showed clear sequence preference and 5 positions showed weak or no sequence preference. The importance of each nucleotide at each position was confirmed by a gel mobility shift competition analysis and results were presented as a quantitative difference in the binding affinity. From these analyses, we conclude that the best binding sequence of BmFTZ-F1 is 5'-PyCAAGGPyCPu-3'. This method may be useful for the determination of a binding sequence of other sequence specific DNA binding factor.

Animals

Amino-acid sequence of lac repressor from Escherichia coli. Isolation, sequence analysis and sequence assembly of tryptic peptides and cyanogen-bromide fragments.

The lac repressor from Escherichia coli, composed of four identical subunits with a molecular weight of 37160, was carboxymethylated and fragmented by tryptic digestion and cyanogen bromide treatment. Using ion-exchange chromatography, gel filtration and preparative thin-layer electrophoresis and chromatography 29 of the 30 tryptic peptides were isolated in pure form. Direct Edman degradation and the dansyl-Edman technique were used to determine the sequence of the small tryptic peptides. Special emphasis was put on the sequence determination of the six large tryptic fragments which together account for 177 residues, corresponding to 51% of the repressor subunit with its 347 residues. The large tryptic fragments were analyzed after fragmentation with chymotrypsin, thermolysin and dipeptidyl aminopeptidase I. Thus the sequence of all 30 tryptic peptides could be deduced. The complete sequences of all cyanogen bromide fragments were deduced from peptides obtained by tryptic, chymotryptic and thermolytic digestion of the individual fragments and by automated stepwise Edman degradation of lac repressor and of the large cyanogen bromide fragments. The order of the cyanogen bromide fragments was given by overlapping tryptic peptides. The resulting amino acid composition of the monomer is Asp15, Asn11, Thr18, Ser30, Glu14, Gln27, Pro13, Gly22, Ala44, Cys3, Val33, Met9, Ile17, Leu40, Tyr8, Phe4, Trp2, Lys11, His7, Arg19. The sequence of lac repressor shows no similarities with that of other proteins known to bind to DNA or RNA. The N-terminal 55 residues contain two homologous regions. This part of the sequence which is involved in lac operator binding might have been formed by gene duplication.

Amino Acid Sequence

Affinity labeling of 3 alpha-hydroxysteroid dehydrogenase with 3 alpha-bromoacetoxyandrosterone and 11 alpha-bromoacetoxyprogesterone. Isolation and sequence of active site peptides containing reactive cysteines; sequence confirmation using nucleotide sequence from a cDNA clone.

Homogeneous 3 alpha-hydroxysteroid dehydrogenase (3 alpha-HSD, EC 1.1.1.50) of rat liver cytosol is potently inhibited at its active site by nonsteroidal anti-inflammatory drugs (NSAIDs). Using 3 alpha-bromoacetoxy-5 alpha-androstan-17-one (BrAnd, a substrate analog) and 11 alpha-bromoacetoxyprogesterone (Br11P, a glucocorticoid analog) as affinity-labeling agents, kinetic evidence was obtained that these agents alkylate this site. Inactivation of 3 alpha-HSD with either [14C]BrAnd or [14C] Br11P led to the incorporation of 1 mol of affinity-labeling agent per enzyme monomer. Complete acid hydrolysis of 3 alpha-HSD radiolabeled with either agent followed by amino acid analysis led to the identification of [14C]carboxymethylcysteine indicating that [14C]BrAnd and [14C]Br11P covalently tag discrete reactive cysteine(s) at the enzyme active site. Trypsin digestion of [14C]BrAnd-inactivated 3 alpha-HSD followed by peptide mapping led to the purification of a single radiolabeled peptide (3A1) which gave the following sequence: H2N-Ser-Ile-Gly-Val-Ser-Asn-Phe-Asn-X-Arg-CO2H. Identical experiments on [14C] Br11P-inactivated 3 alpha-HSD led to the purification of three radiolabeled peptides (11P1-11P3). The major radiolabeled peptide (11P1) had an identical sequence to 3A1 which was tagged with [14C]BrAnd. The minor radiolabeled peptides had the following sequences: H2N-Ser-Lys-Asp-Ile-Ile-Leu-Val-Ser-Tyr-X-Thr-Leu-Gly-Ser-Ser-Arg-CO2H (11P2) and H2N-Ser-Pro-Val-Leu-Leu-Asp-Asp-Pro-Val-Leu-X-Ala-Ile-Ala-Lys-CO2H (11P3). In each peptide group X was identified as carboxymethylcysteine. Alignment of the peptide sequences with the primary structure of 3 alpha-HSD, deduced from its cDNA clone, assigned peptide 11P1 to residues 162-171, peptide 11P2 to residues 208-223, and peptide 11P3 to residues 232-246 of the amino acid sequence. The reactive cysteines correspond to Cys170, Cys217, and Cys242. We propose that Cys170 labeled by BrAnd may lie within the catalytic pocket of the enzyme. By contrast the 11 alpha-bromoacetoxy group in Br11P labeled several reactive cysteines which may be involved in the binding of glucocorticoids and NSAIDs.

3-Hydroxysteroid Dehydrogenases

Complete sequence of the Drosophila nonmuscle myosin heavy-chain transcript: conserved sequences in the myosin tail and differential splicing in the 5' untranslated sequence.

We have sequenced a cDNA that encodes the nonmuscle myosin heavy chain from Drosophila melanogaster. An alternatively spliced exon at the 5' end generates two distinct heavy-chain transcripts: the longer transcripts inserts an additional start codon upstream of the primary translation start site and encodes a myosin heavy chain with a 45-residue extension at its amino terminus. The remainder of the coding sequence reveals extensive homology with other conventional myosins, especially metazoan nonmuscle and smooth muscle myosin isoforms. Comparisons among available myosin heavy-chain sequences establish that characteristic differences in sequence throughout the length of both the globular myosin head and extended rod-like tail readily distinguish nonmuscle and smooth muscle myosins from striated muscle isoforms and predict a basis for their functional diversity.

Amino Acid Sequence

Sequence of the A-protein of coliphage MS2. I. Isolation of A-protein, determination of the NH2- and COOH-terminal sequences, isolation and amino acid sequence of the tryptic peptides.

The A-protein of coliphage MS2 was purified to a state of sufficient homogeneity to study its primary structure. The NH2-terminal sequence was determined for the first 8 residues. Comparison with the reported sequence of R17 protein (Weiner, A. M., Platt, T., and Weber, K. (1972) J. Biol. Chem. 247, 3242-3251) shows a difference at position 6 where alanine in R17 is replaced by threonine in MS2. The COOH-terminal sequence was shown to be -Arg-Leu-Ser-Arg, confirming the existence of UAG as the termination codon of the maturation protein (Comtreras, R., Ysebaert, M., Min Jou, W., and Fiers, W. (19731 Nature New Biol. 241, 99-101; Vandekerckhove, J., Nolf, F., and Van Montagu, M. C. (1973) Nature New Biol. 241, 102; Remaut E., and Fiers, W. (1972) J. Mol. Biol. 71, 243-261). Peptides obtained by enzymatic hydrolysis with trypsin were fractionated by a combination of gel filtration and paper electrophoresis and chromatography. Thirty-eight peptides were analyzed for amino acid composition and sequence. They provide information for 312 of the 393 residues of the A-protein polypeptide chain.

Amino Acid Sequence

Amino acid sequence studies on the alpha chain of human fibrinogen. Overlapping sequences providing the complete sequence.

The complete amino acid sequence of the alpha chain of human fibrinogen has been determined. It contains 610 amino acid residues and has a calculated molecular weight of 66,124. The chain has 10 methionines, and fragmentation with cyanogen bromide yields 11 peptides [Doolittle, R.F., Cassman, K.G., Cottrell, B.A., Friezner, S.J., Hucko, J.T., & Takagi, T. (1977) Biochemistry 16, 1703]. The arrangement of the 11 fragments was determined by the isolation of peptide overlaps from plasmic and staphylococcal protease digests of fibrinogen and/or alpha chains. In addition, certain of the cyanogen bromide fragments, preliminary reports of whose sequences have appeared previously, have been reexamined in order to resolve several discrepancies. The alpha chain is homologous with the beta and gamma chains of fibrinogen, although a large repetitive segment of unusual composition is absent from the latter two chains. The existence of this unusual segment divides the sequence of the alpha chain into three zones of about 200 residues each that are readily distinguishable on the basis of amino acid composition alone.

Amino Acid Sequence

The beta globin gene cluster of the prosimian primate Galago crassicaudatus: nucleotide sequence determination of the 41-kb cluster and comparative sequence analyses.

The nucleotide sequence of the beta globin gene cluster of the prosimian Galago crassicaudatus has been determined. A total sequence spanning 41,101 bp contains and links together previously published sequences of the five galago beta-like globin genes (5'-epsilon-gamma-psi eta-delta-beta-3'). A computer-aided search for middle interspersed repetitive sequences identified 10 LINE (L1) elements, including a 5' truncated repeat that is orthologous to the full-length L1 element found in the human epsilon-gamma intergenic region. SINE elements that were identified included one Alu type I repeat, four Alu type II repeats, and two methionine tRNA-derived Monomer (type III) elements. Alu type II and Monomer sequences are unique to the galago genome. Structural analyses of the cluster sequence reveals that it is relatively A+T rich (about 62%) and regions with high G+C content are associated primarily with globin coding regions. Comparative analyses with the beta globin cluster sequences of human, rabbit, and mouse reveal extensive sequence homologies in their genic regions, but only human, galago, and rabbit sequences share extensive intergenic sequence homologies. Divergence analyses of aligned intergenic and flanking sequences from orthologous human, galago, and rabbit sequences show a gradation in the rate of nucleotide sequence evolution along the cluster where sequences 5' of the epsilon globin gene region show the least sequence divergence and sequences just 5' of the beta globin gene region show the greatest sequence divergence.

Amino Acid Sequence

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition

Identification of peptides within a known protein sequence using COMSEQ analysis of data containing multiple sequences.

Modern methods of automated protein sequence analysis can provide high-quality data with which unambiguous amino-acid sequences can be determined, but analyses are more difficult when the sample is not pure. COMSEQ and auxillary programs were written to facilitate reconciliation of multiple amino-acid sequences potentially contained in noisy data with the known amino-acid sequence of the parent protein. The COMSEQ program prints a matrix in which the first vertical column represents the known amino-acid sequence of a selected protein. Each row of the matrix contains the sequencer yield corresponding to the amino acid in the first column, with each column corresponding to the sequencing reaction cycle. A diagonal which contains net increases of amino acids for each amino acid in the known sequence identifies a peptide potentially contained within the data. The number of matches for each diagonal over the entire known sequence are tabulated and presented as an aid to locating comparisons of greatest interest. The RNDSEQ program conducts multiple analyses using randomized versions of the known amino-acid sequence and tabulates the cumulative frequencies of potential sequence matches irrespective of the true known sequence. TRANSEQ is a utility program that translates edited sequence data from common databases into files that can be used by COMSEQ and RNDSEQ. The programs have been used successfully to identify two co-sequenced peptides from bovine serum albumin, an albumin peptide sequence in the presence of hemoglobin, and to identify two sequences of rat alpha-2u-globulin that differ in their amino termini.

Amino Acid Sequence

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans

Rapid and simple characterization of in vivo HIV-1 sequences using solid-phase direct sequencing.

Solid-phase direct sequencing was used to obtain in vivo sequence data of polymerase chain reaction (PCR)-amplified HIV-1 p25/p7 gene segments. The solid-phase sequencing method was compared to double-strand sequencing of preparative gel electrophoresis-purified amplification products and found to give more consistent results. Lysates of cells were compared to purified DNA as PCR template. HIV-1 sequences were as well amplified from lysates as from purified DNA and 7,780 bp of sequence from 41 samples were produced by direct sequencing. Sequence analysis revealed common sequence motifs relating the sequence to the Euroamerican and African groups of sequences previously described. The results indicate that many viruses of diverse origin circulate in Finland, although the majority seems to be of Euroamerican type. Solid-phase direct sequencing may provide a valuable tool for both epidemiological and pathogenic studies of in vivo HIV-1 infections.

Amino Acid Sequence

Detection of sequence variants in the gene for human type II procollagen (COL2A1) by direct sequencing of polymerase chain reaction-amplified genomic DNA.

The direct sequencing of the human type II procollagen (COL2A1) gene from polymerase chain reaction (PCR)-amplified genomic DNA is described. Thirty-two regions of the COL2A1 gene were asymmetrically amplified with intron primers which were specifically chosen to amplify a region spanning 500 to 800 bp of sequence encoding one or more exons and their accompanying intervening sequences. Primers for dideoxynucleotide sequencing of the PCR products were then designed to provide complete exon sequence information and to insure that intron:exon splice junction sequence data would be obtained. Amplification and sequencing reactions were performed on an automated workstation to facilitate the handling of multiple DNA templates. The procedure allowed efficient sequencing of over 25,000 bp of each allele of the COL2A1 gene per diploid genome. We used this method for the comparative analyses of COL2A1 sequences in DNA isolated from the blood of 42 unrelated individuals and we identified 21 neutral sequence variants in the gene. The sequence variations were confirmed by independent assays, including restriction enzyme digestion. The sequence variants described here will be important for identifying haplotypes of the type II procollagen gene that will be useful in defining a genetic etiology for diseases of cartilaginous tissues.

Alleles

The carboxylesterase family exhibits C-terminal sequence diversity reflecting the presence or absence of endoplasmic-reticulum-retention sequences.

Resident proteins of the endoplasmic reticulum lumen are continuously retrieved from an early Golgi compartment by a receptor-mediated mechanism. The sorting or retention sequence on the endoplasmic reticulum proteins is located at the C-terminus and was initially shown to be the tetrapeptide KDEL in mammalian cells and HDEL in Saccharomyces cerevisiae. The carboxylesterases are a large family of enzymes primarily localized to the lumen of the endoplasmic reticulum. Retention sequences in these proteins have been difficult to identify due to atypical and heterogeneous C-terminal sequences. Utilizing the polymerase chain reaction with degenerate primers, we have identified and characterized the C-termini of four members of the carboxylesterase family from rat liver. Three of the carboxylesterases sequences contained C-terminal sequences (HVEL, HNEL or HTEL) resembling the yeast sorting signal which were reported to be non-functional in mammalian cells. A fourth carboxylesterase contained a distinct C-terminal sequence, TEHT. A full-length esterase cDNA clone, terminating in the sequence HVEL, was isolated and was used to assess the retention capabilities of the various esterase C-terminal sequences. This esterase was retained in COS-1 cells, but was secreted when its C-terminal tetrapeptide, HVEL, was deleted. Addition of C-terminal sequences containing HNEL and HTEL resulted in efficient retention. However, the C-terminal sequence containing TEHT was not a functional retention signal. Both HDEL, the authentic yeast retention signal, and KDEL were efficient retention sequences for the esterase. These studies show that some members of the rat liver carboxylesterase family contain novel C-terminal retention sequences that resemble the yeast signal. At least one member of the family does not contain a C-terminal retention signal and probably represents a secretory form.

Amino Acid Sequence

Sequence arrangement in herpes simplex virus type 1 DNA: identification of terminal fragments in restriction endonuclease digests and evidence for inversions in redundant and unique sequences.

It has been proposed by Sheldrick and Berthelot (1974) that the terminal sequences of herpes simplex virus type 1 (HSV-1) DNA are repeated in an internal inverted form and that the inverted redundant sequences delimit and separate two unique sequences, S and L. In this study the sequence arrangement in HSV-1 DNA has been investigated with restriction endonuclease cleavage, end-labeling studies, and molecular hybridization experiments. The terminal fragments in digests with restriction endonucleases Hind III, Hpa-1, EcoRI and Bum were identified and shown to be consistent with the Sheldrick and Berthelot model. Inverted fragments which contain unique sequences as well as redundant sequences, and which the model predicts, were identified by DNA-DNA hybridization studies. Further cleavage of Bum fragments with Hpa-1 also revealed inversions of the terminal sequences that contained unique sequences. The results obtained showed that the unique sequences S and L are relatively inverted in different DNA molecules in the population, resulting in the presence of four related genomes with rearranged sequences in apparently equal amounts. The redundant sequences bounding S do not share complete sequence homology with those bounding L, but hybridization studies are presented which show that the terminal 0.3% of the genome is repeated in every redundant sequence.

Base Sequence

Sequence analysis of homogeneous peptides of shark immunoglobulin light chains by tandem mass spectrometry: correlation with gene sequence and homologies among variable and constant region peptides of sharks and mammals.

Morphologically, sharks are living fossils that are remarkably similar to their Devonian ancestors of ca. 400 million years ago. If a parallel conservation in biochemical properties characterizes shark evolution, knowledge of the properties of shark immunoglobulins should provide information on the structure of primordial immunoglobulins and their genes. The problem of polyclonality of shark immunoglobulins has precluded detailed analysis of shark immunoglobulin light polypeptide chains. Here, we approach the problem of obtaining direct sequence information on polyclonal light chains of shark immunoglobulins by isolating homogeneous peptides from tryptic digests of shark light chains and sequencing these by tandem mass spectrometry. To confirm the location of the peptides, we isolated a complementary DNA (cDNA) clone from a sandbar shark cDNA library in the expression vector lambda gt11, identifying the clone by its ability to produce a peptide serologically detectable using rabbit antibody to purified shark light chain. The correspondence between peptide sequence and that derived from gene sequence provided direct proof that the gene studied was that of a major expressed serum light chain. Using this combined approach, we isolated homogeneous peptides from both constant and variable regions. The variable region peptides showed homology to corresponding sequences of mammalian V lambda and V kappa sequences. The constant region gene sequence we obtained was homologous to mammalian C lambda sequence. The four constant region tryptic peptides we sequenced corresponded exactly to stretches of the C lambda sequence derived from the DNA sequence. The combined approach described here shows that shark light chains exhibit heterogeneity at both the protein and gene level, but that the constant regions of these chains can be identified as homologs of mammalian lambda chains and that evolutionary conservation has occurred in V region sequences ranging from elasmobranchs to man.

Amino Acid Sequence

A method of sequencing without subcloning and its application to the identification of a novel ORF with a sequence suggestive of a transcriptional regulator in the water mold Achlya ambisexualis.

Genomic amplification with transcript sequencing (GAWTS) is a method of direct sequencing that involves amplification with PCR using primers containing phage promoters, transcription of the amplified product, and sequencing with reverse transcriptase. GAWTS requires the generation of PCR primers that are specific for the sequences on both sides of a region. Here we describe promoter ligation and transcript sequencing (PLATS), a direct method for rapidly obtaining novel sequences that utilizes generic primers and only requires knowledge of the sequence on one side of a region. PLATS involves restriction digestion of the amplified vector insert, ligation with a phage promoter, and then GAWTS using phage promoter sequences as the PCR primers. The method is rapid and economical because it uses a limited set of oligonucleotides, and it is potentially amenable to automation because it does not require in vivo manipulations. PLATS facilitates the determination of a genomic sequence responsible for cross-hybridization in a Southern blot. Using PLATS, sequence has been obtained from a 1.1-kb segment in Achlya ambisexualis, which cross-hybridizes to the DNA-binding region of the chicken and Xenopus estrogen receptors. To our knowledge, this represents the first sequence reported from the Oomycetes, a large and widely distributed group of fungi. The sequence reveals a large, transcribed open reading frame that is markedly deficient in the dinucleotide TpA. A putative zinc finger containing three cysteines and one histidine (C-X2-C-X12-H-X3-C) and an acidic segment hint that this clone may be a member of a novel class of transcriptional regulators.

Amino Acid Sequence

Inverted repeat regions of Marek's disease virus DNA possess a structure similar to that of the a sequence of herpes simplex virus DNA and contain host cell telomere sequences.

The genomic structure of Marek's disease virus (MDV) is similar to those of the alphaherpesviruses herpes simplex virus (HSV) types 1 and 2. Sequence analysis of the junction region between the long component (L) and the short component (S) revealed the existence of an a-like sequence, similar in structure to the a sequence of HSV-1. Further study revealed that the MDV genome contains five copies of the a-like sequence within the long terminal repeat region as well as in the short terminal repeat region. The junction between the L and S components was found to contain 10 copies of the a-like sequence. Within the a-like sequence, a structure homologous to the DR2 of HSV was found to contain 17 copies of the telomeric sequence, GGGGTTA. There appears to be little to no sequence homology between the HSV a sequence and the MDV a-like sequence; however, the strong physical homology to its counterpart in HSV-1 suggests that the MDV a-like sequence may have the same functional homology (the domain for cleavage/packaging of the DNA into the viral capsids and for genomic inversion) as well.

Animals