PubMed HealthSearch

PubMed · 9144219

Multiple-complete-digest restriction fragment mapping: generating sequence-ready maps for large-scale DNA sequencing.

Abstract

Multiple-complete-digest mapping is a DNA mapping technique based on complete-restriction-digest fingerprints of a set of clones that provides highly redundant coverage of the mapping target. The maps assembled from these fingerprints order both the clones and the restriction fragments. Maps are coordinated across three enzymes in the examples presented. Starting with yeast artificial chromosome contigs from the 7q31.3 and 7p14 regions of the human genome, we have produced cosmid-based maps spanning more than one million base pairs. Each yeast artificial chromosome is first subcloned into cosmids at a redundancy of x15-30. Complete-digest fragments are electrophoresed on agarose gels, poststained, and imaged on a fluorescent scanner. Aberrant clones that are not representative of the underlying genome are rejected in the map construction process. Almost every restriction fragment is ordered, allowing selection of minimal tiling paths with clone-to-clone overlaps of only a few thousand base pairs. These maps demonstrate the practicality of applying the experimental and software-based steps in multiple-complete-digest mapping to a target of significant size and complexity. We present evidence that the maps are sufficiently accurate to validate both the clones selected for sequencing and the sequence assemblies obtained once these clones have been sequenced by a "shotgun" method.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

G K Wong, J Yu, E C Thayer, M V Olson. 1997-05-13. Multiple-complete-digest restriction fragment mapping: generating sequence-ready maps for large-scale DNA sequencing.. https://doi.org/10.1073/pnas.94.10.5225

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Variant effects of non-native kissing-loop hairpin palindromes on HIV replication and HIV RNA dimerization: role of stem-loop B in HIV replication and HIV RNA dimerization.

The genome of all retroviruses consists of two identical RNAs noncovalently linked near their 5' end. In vitro synthesized RNAs from human immunodeficiency virus type 1 (HIV-1) can form loose or tight dimers depending on whether their respective kissing-loop hairpins (nts 248-270 in HIV-1Lai) bond via their hexameric autocomplementary sequences (ACS), also called palindromes, or via the ACS and stem sequences [Laughrea, M., and Jetté, L. (1996) Biochemistry 35, 1589-1598]. To understand the role of the ACS in HIV-1 replication and in the formation and stability of HIV-1 RNA dimers, we replaced the central CGCG261(or tetramer) of the HIV-1Lai ACS by two other HIV-1 tetramers (UGCA/UGCG), four non-HIV-1 tetramers [GUAC, UUAA (respectively found in HIV-2Rod and SIVmnd), GGCC and AGCU (absent from HIV and SIV viruses)], or GGCG, a nonpalindromic tetramer. The infectivity of GGCC, GUAC, and UGCA viruses was unchanged or insignificantly decreased; the infectivity of AGCU and UGCG viruses was decreased by 80%; the infectivity of UUAA and GGCG viruses was decreased by 92-98%. Thus, the four non-HIV-1 palindromes yielded phenotypes ranging from wild-type to as defective as a virus bearing a nonpalindrome. Studies of in vitro synthesized HIV-1 RNAs were generally consistent with in vivo results, specifically: (i) loose dimerization of GGCC and GUAC RNAs, but not of UUAA and AGCU RNAs, was influenced by the 3' DLS (a sequence located downstream of the 5' splice junction) in a way expected for a wild-type ACS; (ii) the 3' DLS strongly reduced tight dimerization of UUAA and AGCU RNAs, but not of GGCC and GUAC RNAs. We conclude that HIV-1 is sensitive to the ACS sequence without discriminating against all nonnative ACS: GGCC/GUAC, but not AGCU/UUAA, are good substitutes for the prevalent CGCG/UGCA native tetramers and better substitutes than the very rare UGCG native tetramer. The correlation between in vivo and in vitro results suggests that in vitro assays measure parameters of in vivo relevance. Deletion of CUCGG247 (the 5' strand of stem-loop B) decreased the replicative capacity by more than 99.9% and metamorphosed the 3' DLS into an inhibitor of the loose dimerization of HIV-1 RNA.

Base Composition

Biased nucleotide composition of the genome of HERV-K related endogenous retroviruses and its evolutionary implications.

The human genome contains a large number of sequences that belong to the HERV-K family of human endogenous retroviruses. Most of these elements are likely remnants of ancient infections by ancestral exogenous retroviruses. To obtain further insight into the evolutionary history and molecular mechanisms responsible for the diversity of the human HERV-K elements, we analyzed several aspects of their genome structure. The nucleotide composition of the HERV-K genome was found to be highly biased and asymmetric, with an abundance of the A nucleotide in the viral (+) strand. A similar trend has been reported for the genomes of several exogenous retroviruses, with different nucleotides as the preferred building block. Other genome characteristics that were reported previously for actively replicating retroviruses are also apparent for the endogenous HERV-K virus. In particular, we observed suppression of the dinucleotide CpG, which represents potential methylation sites, and a strong preference for synonymous substitutions within the open reading frame of the reverse transcriptase (RT) enzyme. Furthermore, the mutational spectrum of the HERV-K RT enzyme was evaluated by nucleotide sequence comparison of 34 available elements. Interestingly, this analysis revealed a striking similarity with the mutational pattern of the HIV-1 RT enzyme, with a preference for G-to-A and C-to-T transitions. It is proposed that the mutational bias of the HERV-K RT enzyme played a role in the shaping of this retroviral genome, which was actively replicating more than 30 million years ago. This effect can still be observed in the contemporary endogenous HERV-K elements.

Base Composition

Two distinct mechanisms cause heterogeneity of 16S rRNA.

To investigate the frequency of heterogeneity among the multiple 16S rRNA genes within a single microorganism, we determined directly the 120-bp nucleotide sequences containing the hypervariable alpha region of the 16S rRNA gene from 475 Streptomyces strains. Display of the direct sequencing patterns revealed the existence of 136 heterogeneous loci among a total of 33 strains. The heterogeneous loci were detected only in the stem region designated helix 10. All of the substitutions conserved the relevant secondary structure. The 33 strains were divided into two groups: one group, including 22 strains, had less than two heterogeneous bases; the other group, including 11 strains, had five or more heterogeneous bases. The two groups were different in their combinations of heterogeneous bases. The former mainly contained transitional substitutions, and the latter was mainly composed of transversional substitutions, suggesting that at least two mechanisms, possibly misincorporation during DNA replication and horizontal gene transfer, cause rRNA heterogeneity.

Base Composition