PubMed Health⌕ Search

PubMed · 16960799

A new method for detecting human recombination hotspots and its applications to the HapMap ENCODE data.

Abstract

Computational detection of recombination hotspots from population polymorphism data is important both for understanding the nature of recombination and for applications such as association studies. We propose a new method for this task based on a multiple-hotspot model and an (approximate) log-likelihood ratio test. A truncated, weighted pairwise log-likelihood is introduced and applied to the calculation of the log-likelihood ratio, and a forward-selection procedure is adopted to search for the optimal hotspot predictions. The method shows a relatively high power with a low false-positive rate in detecting multiple hotspots in simulation data and has a performance comparable to the best results of leading computational methods in experimental data for which recombination hotspots have been characterized by sperm-typing experiments. The method can be applied to both phased and unphased data directly, with a very fast computational speed. We applied the method to the 10 500-kb regions of the HapMap ENCODE data and found 172 hotspots among the three populations, with average hotspot width of 2.4 kb. By comparisons with the simulation data, we found some evidence that hotspots are not all identical across populations. The correlations between detected hotspots and several genomic characteristics were examined. In particular, we observed that DNaseI-hypersensitive sites are enriched in hotspots, suggesting the existence of human beta hotspots similar to those found in yeast.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jun Li, Michael Q Zhang, Xuegong Zhang. 2006-08-30. A new method for detecting human recombination hotspots and its applications to the HapMap ENCODE data.. https://doi.org/10.1086/508066

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Sp1 is essential for p16 expression in human diploid fibroblasts during senescence.

BACKGROUND: p16(INK4a) tumor suppressor protein has been widely proposed to mediate entrance of the cells into the senescent stage. Promoter of p16(INK4a) gene contains at least five putative GC boxes, named GC-I to V, respectively. Our previous data showed that a potential Sp1 binding site, within the promoter region from -466 to -451, acts as a positive transcription regulatory element. These results led us to examine how Sp1 and/or Sp3 act on these GC boxes during aging in cultured human diploid fibroblasts. METHODOLOGY/PRINCIPAL FINDINGS: Mutagenesis studies revealed that GC-I, II and IV, especially GC-II, are essential for p16(INK4a) gene expression in senescent cells. Electrophoretic mobility shift assays (EMSA) and ChIP assays demonstrated that both Sp1 and Sp3 bind to these elements and the binding activity is enhanced in senescent cells. Ectopic overexpression of Sp1, but not Sp3, induced the transcription of p16(INK4a). Both Sp1 RNAi and Mithramycin, a DNA intercalating agent that interferes with Sp1 and Sp3 binding activities, reduced p16(INK4a) gene expression. In addition, the enhanced binding of Sp1 to p16(INK4a) promoter during cellular senescence appeared to be the result of increased Sp1 binding affinity, not an alteration in Sp1 protein level. CONCLUSIONS/SIGNIFICANCE: All these results suggest that GC- II is the key site for Sp1 binding and increase of Sp1 binding activity rather than protein levels contributes to the induction of p16(INK4a) expression during cell aging.

Base Composition↗

Application of CE for determination of DNA base composition.

DNA base composition expressed as mol% of guanine plus cytosine (% GC) or GC content is a key parameter of bacterial taxonomy and genomic analyses. Direct chemical determination methods such as HPLC as well as indirect methods based on physical properties of deoxyribonucleic acid (DNA), melting point (T(m)), and buoyant density (B(d)) have been conventionally applied to determine the GC content. However, these methods require relatively large amounts of sample DNA, time, and labor. We have developed a protocol to determine the GC content by fine separation of nucleosides with CZE. Genomic DNAs with known GC content from 23 bacterial strains were determined by CE at the optimized conditions of 27 degrees C, 20 kV in 50 mM of NaHCO(3) (pH 9.0) and 70 mM SDS added. Nucleosides from <1 microg of DNA hydrolyzed with nuclease-P1 and bacterial alkaline phosphatase were separated in a 75 microm wide and 80 cm long silica capillary. The nucleoside peak areas were determined at 254 nm in less than 12 min. The CE-based determination of GC content requires only small amounts of DNA, and thus should be applicable to environmental genomics (metagenomics), as >90% of environmental micro-organisms are nonculturable and produce only small amounts of genomic DNA.

Base Composition↗