PubMed Health⌕ Search

PubMed · 12176825

Why are complementary DNA strands symmetric?

Abstract

MOTIVATION: Over sufficiently long windows, complementary strands of DNA tend to have the same base composition. A few reports have indicated that this first-order parity rule extends at higher orders to oligonucleotide composition, at least in some organisms or taxa. However, the scientific literature falls short of providing a comprehensive study of reverse-complement symmetry at multiple orders and across the kingdom of life. It also lacks a characterization of this symmetry and a convincing explanation or clarification of its origin. RESULTS: We develop methods to measure and characterize symmetry at multiple orders, and analyze a wide set of genomes, encompassing single- and double-stranded RNA and DNA viruses, bacteria, archae, mitochondria, and eukaryota. We quantify symmetry at orders 1 to 9 for contiguous sequences and pools of coding and non-coding upstream regions, compare the observed symmetry levels to those predicted by simple statistical models, and factor out the effect of lower-order distributions. We establish the universality and variability range of first-order strand symmetry, as well as of its higher-order extensions, and demonstrate the existence of genuine high-order symmetric constraints. We show that ubiquitous reverse-complement symmetry does not result from a single cause, such as point mutation or recombination, but rather emerges from the combined effects of a wide spectrum of mechanisms operating at multiple orders and length scales.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pierre-François Baisnée, Steve Hampson, Pierre Baldi. 2002. Why are complementary DNA strands symmetric?. https://doi.org/10.1093/bioinformatics%2F18.8.1021

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Sp1 is essential for p16 expression in human diploid fibroblasts during senescence.

BACKGROUND: p16(INK4a) tumor suppressor protein has been widely proposed to mediate entrance of the cells into the senescent stage. Promoter of p16(INK4a) gene contains at least five putative GC boxes, named GC-I to V, respectively. Our previous data showed that a potential Sp1 binding site, within the promoter region from -466 to -451, acts as a positive transcription regulatory element. These results led us to examine how Sp1 and/or Sp3 act on these GC boxes during aging in cultured human diploid fibroblasts. METHODOLOGY/PRINCIPAL FINDINGS: Mutagenesis studies revealed that GC-I, II and IV, especially GC-II, are essential for p16(INK4a) gene expression in senescent cells. Electrophoretic mobility shift assays (EMSA) and ChIP assays demonstrated that both Sp1 and Sp3 bind to these elements and the binding activity is enhanced in senescent cells. Ectopic overexpression of Sp1, but not Sp3, induced the transcription of p16(INK4a). Both Sp1 RNAi and Mithramycin, a DNA intercalating agent that interferes with Sp1 and Sp3 binding activities, reduced p16(INK4a) gene expression. In addition, the enhanced binding of Sp1 to p16(INK4a) promoter during cellular senescence appeared to be the result of increased Sp1 binding affinity, not an alteration in Sp1 protein level. CONCLUSIONS/SIGNIFICANCE: All these results suggest that GC- II is the key site for Sp1 binding and increase of Sp1 binding activity rather than protein levels contributes to the induction of p16(INK4a) expression during cell aging.

Base Composition↗

Application of CE for determination of DNA base composition.

DNA base composition expressed as mol% of guanine plus cytosine (% GC) or GC content is a key parameter of bacterial taxonomy and genomic analyses. Direct chemical determination methods such as HPLC as well as indirect methods based on physical properties of deoxyribonucleic acid (DNA), melting point (T(m)), and buoyant density (B(d)) have been conventionally applied to determine the GC content. However, these methods require relatively large amounts of sample DNA, time, and labor. We have developed a protocol to determine the GC content by fine separation of nucleosides with CZE. Genomic DNAs with known GC content from 23 bacterial strains were determined by CE at the optimized conditions of 27 degrees C, 20 kV in 50 mM of NaHCO(3) (pH 9.0) and 70 mM SDS added. Nucleosides from <1 microg of DNA hydrolyzed with nuclease-P1 and bacterial alkaline phosphatase were separated in a 75 microm wide and 80 cm long silica capillary. The nucleoside peak areas were determined at 254 nm in less than 12 min. The CE-based determination of GC content requires only small amounts of DNA, and thus should be applicable to environmental genomics (metagenomics), as >90% of environmental micro-organisms are nonculturable and produce only small amounts of genomic DNA.

Base Composition↗