PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “DNA sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Sequence assembly validation by multiple restriction digest fragment coverage analysis.

DNA sequence analysis depends on the accurate assembly of fragment reads for the determination of a consensus sequence. This report examines the possibility of analyzing multiple, independent restriction digests as a method for testing the fidelity of sequence assembly. A dynamic programming algorithm to determine the maximum likelihood alignment of error prone electrophoretic mobility data to the expected fragment mobilities given the consensus sequence and restriction enzymes is derived and used to assess the likelihood of detecting rearrangements in genomic sequencing projects. The method is shown to reliably detect errors in sequence fragment assembly without the necessity of making reference to an overlying physical map. An html form-based interface is available at http:/(/)www.ibc.wustl.edu/services/validate. html.

Algorithms↗

DNA sequence analysis of a mouse pro alpha 1 (I) procollagen gene: evidence for a mouse B1 element within the gene.

In a 3.8-kilobase mouse DNA sequence encoding amino acid sequences for the pro alpha 1(I) chain of type I procollagen, 14 coding sequences were identified which specify a sequence 95% homologous to amino acid residues 568 to 963 of the bovine alpha 1(I) chain. All of these coding sequences were flanked by appropriate splice junctions following the GT/AG rule. These observations suggest, but do not prove, that this pro alpha 1(I) gene is transcriptionally active. Of the 14 coding sequences, 7 were 54 base pairs in length, whereas the remainder were higher multiples of 54 base pairs. Nonrandom utilization of codons pertained throughout all of the coding sequences showing a preference (56%) for U in the wobble position. Two of the intervening sequences encoded imperfect vestiges of coding sequences which exhibited a codon preference different from that of the pro alpha 1(I) gene proper and were not flanked by splice junctions. One intervening sequence encoded a member of the mouse B1 family of middle repetitive sequences. It was flanked by 8-base-pair direct repeats and had a truncated A-rich region, suggesting that it may be a mobile element. Within this element were sequences which could function as a RNA polymerase III split promoter.

Animals↗

DNA sequence analysis of the dnaK gene of Escherichia coli B and of two dnaK genes carrying the temperature-sensitive mutations dnaK7(Ts) and dnaK756(Ts).

The DNA sequence of the dnaK gene of Escherichia coli was analyzed. The nucleotide sequence of the wild-type dnaK gene of E. coli B differed from that of E. coli K-12 in 15 bp, none of which altered the amino acid sequence. Two temperature-sensitive dnaK mutations were examined by cloning and sequence analyses. Results showed that one dnaK mutation, dnaK7(Ts), was a one-base substitution of T for C at nucleotide position 448 in the open reading frame yielding an amber nonsense codon. The other mutation, dnaK756(Ts), consisted of base substitutions (A for G) at three nucleotide positions, 95, 1364, and 1403, in the open reading frame resulting in an aspartic acid codon in place of a glycine codon.

Amino Acid Sequence↗

DNA sequence analysis of ARS elements from chromosome III of Saccharomyces cerevisiae: identification of a new conserved sequence.

Four fragments of Saccharomyces cerevisiae chromosome III DNA which carry ARS elements have been sequenced. Each fragment contains multiple copies of sequences that have at least 10 out of 11 bases of homology to a previously reported 11 bp core consensus sequence. A survey of these new ARS sequences and previously reported sequences revealed the presence of an additional 11 bp conserved element located on the 3' side of the T-rich strand of the core consensus. Subcloning analysis as well as deletion and transposon insertion mutagenesis of ARS fragments support a role for 3' conserved sequence in promoting ARS activity.

Base Sequence↗

Antigenic determinant in human coagulation factor IX: immunological screening and DNA sequence analysis of recombinant phage map a monoclonal antibody to residues 111 through 132 of the zymogen.

As an approach to the study of structure-function relationships in the normal and defective forms of human coagulation factor IX, we have begun to develop a series of monoclonal antibodies against specific sites on the protein. Zymogen and activated forms of normal factor IX were used initially as antigen for the preparation of monoclonal antibodies. Recombinant phage were prepared by cloning small (50- to 500-nucleotide) random DNA fragments from the coding region of a factor IX cDNA clone into the expression vector lambda gt11. Immunological screening of these recombinants with mixtures of monoclonal antibodies identified several immunoreactive phage. Further analysis showed that the monoclonal antibody designated IX-30 was generating the positive signals at a frequency of approximately 1/2,500 recombinants. Subcloning and sequence analysis of the inserted DNA in the immunoreactive phage revealed overlapping in-frame insertions, from which it could be inferred that the site in factor IX recognized by IX-30 is confined to residues 111 through 132 in the light chain. Similar mapping with other monoclonal antibodies should provide additional probes for the protein structure of human factor IX.

Amino Acid Sequence↗

Spontaneous mutagenesis in Escherichia coli harbouring plasmid pKM101: DNA sequence analysis of forward lacI- mutations.

To investigate the influence of plasmid pKM101 on spontaneous mutagenesis, 198 lacI- mutations generated in Escherichia coli harbouring pKM101 were characterized at the DNA sequence level. pKM101 by itself did not enhance the lacI- forward mutation frequency. In general, the resultant distribution of mutation highlights particular sequences at which a variety of mutational events repeatedly occur, including 5'-(G/C)TGG-3', 5'-CCAGG-3', 5'-GATC-3' and 5'-TCGCG-3' sequences. Specifically, the distribution of mutation within base substitution and deletion classes distinguishes the pKM101 spectrum from the wild-type distribution (i.e. absence of pKM101). An even distribution amongst base substitutions was observed which corresponds to a 2.9- to 6.3-fold increase in occurrence of low frequency events (A-->G,T,C; G-->C); high frequency events in the wild-type distribution (G-->A,T) were not influenced by the presence of pKM101. One complex event was recovered which was comprised of two base substitutions separated by 4 bp. An 11-fold increase in small deletion events (3-6 bp) was also observed. The observed pKM101 spectrum does not closely resemble the mutational consequences of SOS induction in the absence of mutagenic treatment (recA441 spectrum) but does return a distribution like that obtained in E. coli deficient in polymerase I activity (polA1 spectrum).

Amino Acid Sequence↗

Cloning and DNA sequence analysis of a Lactococcus bacteriophage lysin gene.

A gene for the lysin of Lactococcus lactis bacteriphage phi vML3 was cloned using an Escherichia coli/bacteriophage lambda host-vector system. The gene was detected by its expression of antimicrobial activity against L. lactis cells in a bioassay. The cloned fragment was analysed by sub-cloning on to E. coli plasmid vectors and by restriction endonuclease and deletion mapping. Its entire DNA sequence was determined and an open reading frame for the lysin structural gene was identified. The sequenced lysin gene would express a protein of 187 amino acids with a molecular weight of 21,090, which is in good agreement with that of a protein detected after in vitro transcription and translation of DNA encoding the gene. Expression of the lysin gene in E. coli and B. subtilis from an adjacent bacteriophage promoter was readily detected but in L. lactis expression of lysin was found to be lethal. The bacteriophage phi vML3 lysin had sequence homology with protein 15 of B. subtilis bacteriophage PZA. This protein is involved in DNA packaging during bacteriophage maturation rather than in host cell lysis. The cloning and analysis of the phi vML3 lysin gene is of importance in further understanding lactic streptococcal bacteriophages, for the development of positive selection vectors and for biotechnological applications of relevance to the dairy industry.

Amino Acid Sequence↗

DNA sequence analysis on the IBM-PC.

We have developed, for the IBM-PC microcomputer, a menu driven, interactive set of programs which provide the functions routinely used for DNA sequence data analyses.

Amino Acid Sequence↗

16S ribosomal DNA sequence analysis of a large collection of environmental and clinical unidentifiable bacterial isolates.

Some bacteria are difficult to identify with phenotypic identification schemes commonly used outside reference laboratories. 16S ribosomal DNA (rDNA)-based identification of bacteria potentially offers a useful alternative when phenotypic characterization methods fail. However, as yet, the usefulness of 16S rDNA sequence analysis in the identification of conventionally unidentifiable isolates has not been evaluated with a large collection of isolates. In this study, we evaluated the utility of 16S rDNA sequencing as a means to identify a collection of 177 such isolates obtained from environmental, veterinary, and clinical sources. For 159 isolates (89.8%) there was at least one sequence in GenBank that yielded a similarity score of > or =97%, and for 139 isolates (78.5%) there was at least one sequence in GenBank that yielded a similarity score of > or =99%. These similarity score values were used to defined identification at the genus and species levels, respectively. For isolates identified to the species level, conventional identification failed to produce accurate results because of inappropriate biochemical profile determination in 76 isolates (58.7%), Gram staining in 16 isolates (11.6%), oxidase and catalase activity determination in 5 isolates (3.6%) and growth requirement determination in 2 isolates (1.5%). Eighteen isolates (10.2%) remained unidentifiable by 16S rDNA sequence analysis but were probably prototype isolates of new species. These isolates originated mainly from environmental sources (P = 0.07). The 16S rDNA approach failed to identify Enterobacter and Pantoea isolates to the species level (P = 0.04; odds ratio = 0.32 [95% confidence interval, 0.10 to 1.14]). Elsewhere, the usefulness of 16S rDNA sequencing was compromised by the presence of 16S rDNA sequences with >1% undetermined positions in the databases. Unlike phenotypic identification, which can be modified by the variability of expression of characters, 16S rDNA sequencing provides unambiguous data even for rare isolates, which are reproducible in and between laboratories. The increase in accurate new 16S rDNA sequences and the development of alternative genes for molecular identification of certain taxa should further improve the usefulness of molecular identification of bacteria.

Animals↗

DNA sequence analysis of a Drosophila foldback transposable element rearrangement.

The complete nucleotide sequence of a DNA rearrangement associated with the foldback 4 (FB 4) transposable element is presented. The results demonstrate that the entire loop sequence and almost all of one of the inverted terminal repeats is absent. Moreover, the sequence of the remaining inverted repeat suggests that the FB elements might undergo inversions via recombinations between the two inverted repeats of a single element.

Base Sequence↗

Plasmid vectors for selecting IS1-promoted deletions in cloned DNA: sequence analysis of the omega interposon.

We have constructed two plasmid vectors which allow selection for in vivo deletions within cloned DNA fragments. The plasmids are derivatives of pBR322 which carry the Escherichia coli rpsL (strA) gene, known to confer a dominant streptomycin (Sm)-sensitivity phenotype to the host cell, and a copy of the IS1 transposable element. Sm-resistant strains that harbor these plasmids display sensitivity to Sm. Spontaneous IS1-promoted deletions across the rpsL gene can be isolated simply by selection for Sm resistance. Hence, nested sets of deletions of a cloned DNA can be obtained and sequenced with an IS1-specific primer. Using this approach, we have determined the complete nucleotide sequence of the omega interposon [Prentki and Krisch, Gene 29 (1984) 303-313].

Amino Acid Sequence↗

Cloning and DNA sequence analysis of pepQ, a prolidase gene from Lactobacillus delbrueckii subsp. lactis DSM7290 and partial characterization of its product.

From a genomic library of Lactobacillus delbrueckii subsp. lactis (DSM7290) DNA, in the low-copy-number vector pLG339, a recombinant clone was selected, which complemented a mutation in the prolidase gene (pepQ) of Escherichia coli UK173. Nucleotide sequence analysis revealed an open reading frame of 1104 nucleotides corresponding to a protein of 368 amino acids with a calculated pI of 4.64 and a molecular mass of 41,087 Da. The start site of pepQ transcription was determined by primer extension analysis with mRNA prepared from L. delbrueckii. Based on homology of the gene product to various peptidases and on the substrate specificity determined, the peptidase was designated PepQ. The influence of various protease inhibitors and cations on peptidase activity indicated that PepQ is a metalloprotease. The absence of a membrane-spanning domain and a signal peptide sequence argues for a cytoplasmic localization of the enzyme.

Amino Acid Sequence↗

Molecular cloning, expression, and DNA sequence analysis of the gene that encodes the 16-kilodalton outer membrane lipoprotein of Serpulina hyodysenteriae.

The gene (smpA) that encodes the 16-kDa outer membrane lipoprotein of Serpulina hyodysenteriae was cloned in Escherichia coli, and its primary structure was determined by nucleotide sequencing. The putative open reading frame encodes a prolipoprotein of 16.8 kDa which in its fully acylated and cleaved form is 15.1 kDa. Analysis of the N-terminal amino acid sequence derived from the DNA sequence revealed the presence of a signal sequence and a putative acylation and signal peptidase II cleavage site (Phe-Ala-Val-Ser-Cys). In E. coli, processing of the prolipoprotein was less efficient than that observed in S. hyodysenteriae, and globomycin, an inhibitor of signal peptidase II, inhibited cleavage of the lipoprotein expressed in E. coli but did not inhibit cleavage in S. hyodysenteriae.

Amino Acid Sequence↗

Statistical methods of DNA sequence analysis: detection of intragenic recombination or gene conversion.

Simple but exact statistical tests for detecting a cluster of associated nucleotide changes in DNA are presented. The tests are based on the linear distribution of a set of s sites among a total of n sites, where the s sites may be the variable sites, sites of insertion/deletion, or categorized in some other way. These tests are especially useful for detecting gene conversion and intragenic recombination in a sample of DNA sequences. In this case, the sites of interest are those that correspond to particular ways of splitting the sequences into two groups (e.g., sequences A and D vs. sequences B, C, and E-J). Each such split is termed a phylogenetic partition. Application of these methods to a well-documented case of gene conversion in human gamma-globin genes shows that sites corresponding to two of the three observed partitions are significantly clustered, whereas application to hominoid mitochondrial DNA sequences--among which no recombination is expected to occur--shows no evidence of such clustering. This indicates that clustering of partition-specific sites is largely due to intragenic recombination or gene conversion. Alternative hypotheses explaining the observed clustering of sites, such as biased selection or mutation, are discussed.

Animals↗

Molecular cloning and DNA sequence analysis of pepL, a leucyl aminopeptidase gene from Lactobacillus delbrueckii subsp. lactis DSM7290.

A genomic library of Lactobacillus delbrueckii subsp. lactis DSM7290 DNA fragments from a Sau3A partial digestion in the low-copy-number vector pLG339, was used to screen Escherichia coli for the presence of peptidases. Using the chromogenic substrate leucine-beta-naphthylamide (Leu-NH-Nap) and E. coli strain CM89 lacking the corresponding enzyme activity in an enzymic plate assay, allowed the isolation of two peptidase genes; the newly described pepL and the recently cloned and sequenced pepN. Clones could be distinguished not only by the restriction pattern of isolated plasmids but also by the rate and intensity of their colour reaction with Leu-NH-Nap. Three out of five clones were identified to express the Lactobacillus pepN gene; the others were shown to express a second aminopeptidase gene, designated pepL. This gene, together with 200 bp upstream of the proposed AUG initiation codon, was further subcloned and sequenced. The corresponding open reading frame of 897 nucleotides is predicted to encode a protein of 299 amino acids (34,541 Da). Searching the EMBL database revealed similarity to the prolinase of Lactobacillus helveticus (45.8% identity), to the iminopeptidases of Lb. delbrueckii subsp. lactis and Lb. delbrueckii subsp. bulgaricus (25.5%), and to the Bacillus coagulans prolinase (21.5%). Minor similarities were detected for hydrolytic enzymes with serine active sites. The product encoded by the pepL gene was functional but could not be visualized on Coomassie-blue-stained polyacrylamide gels. High level expression of peptidase L in E. coli was achieved by placing the gene under the control of the T7 promoter.

Amino Acid Sequence↗