PubMed Health⌕ Search

PubMed · 16380714

Conserved noncoding sequences are selectively constrained and not mutation cold spots.

Abstract

Noncoding genetic variants are likely to influence human biology and disease, but recognizing functional noncoding variants is difficult. Approximately 3% of noncoding sequence is conserved among distantly related mammals, suggesting that these evolutionarily conserved noncoding regions (CNCs) are selectively constrained and contain functional variation. However, CNCs could also merely represent regions with lower local mutation rates. Here we address this issue and show that CNCs are selectively constrained in humans by analyzing HapMap genotype data. Specifically, new (derived) alleles of SNPs within CNCs are rarer than new alleles in nonconserved regions (P = 3 x 10(-18)), indicating that evolutionary pressure has suppressed CNC-derived allele frequencies. Intronic CNCs and CNCs near genes show greater allele frequency shifts, with magnitudes comparable to those for missense variants. Thus, conserved noncoding variants are more likely to be functional. Allele frequency distributions highlight selectively constrained genomic regions that should be intensively surveyed for functionally important variation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jared A Drake, Christine Bird, James Nemesh, Daryl J Thomas, Christopher Newton-Cheh, Alexandre Reymond, Laurent Excoffier, Homa Attar, Stylianos E Antonarakis, Emmanouil T Dermitzakis, Joel N Hirschhorn. 2005-12-25. Conserved noncoding sequences are selectively constrained and not mutation cold spots.. https://doi.org/10.1038/ng1710

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Base-pair resolution conservation data improves cell type specific sequence-to-expression prediction.

MOTIVATION: Genomic sequence-to-activity models can decipher gene regulatory mechanisms and predict the functional impact of regulatory variants. However, current models struggle to integrate information from sequences outside promoters, especially information from cell type specific regulatory elements. RESULTS: Here, we propose incorporating base-pair resolution evolutionary conservation data into genomic sequence-to-expression predictors. We explore two training strategies-training from scratch or fine-tuning an existing sequence-only model with additional conservation input. We find that in both cases, base-pair resolution conservation data improves cell type specific sequence-to-expression prediction, with training from scratch yielding the greatest benefit. The improvement in cell type specific expression prediction can be attributed in part to the fact that models trained on sequence and conservation data learn to better recognize cell type specific regulatory elements than models trained on sequence alone. AVAILABILITY: Code is available at https://github.com/ni-lab/basenji-phyloP.

Conserved Sequence↗

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence↗

New functional identity for the DNA uptake sequence in transformation and its presence in transcriptional terminators.

The frequently occurring DNA uptake sequence (DUS), recognized as a 10-bp repeat, is required for efficient genetic transformation in the human pathogens Neisseria meningitidis and Neisseria gonorrhoeae. Genome scanning for DUS occurrences in three different species of Neisseria demonstrated that 76% of the nearly 2,000 neisserial DUS were found to have two semiconserved base pairs extending from the 5' end of DUS to constitute a 12-mer repeat. Plasmids containing sequential variants of the neisserial DUS were tested for their ability to transform N. meningitidis and N. gonorrhoeae, and the 12-mer was found to outperform the 10-mer DUS in transformation efficiency. Assessment of meningococcal uptake of DNA confirmed the enhanced performance of the 12-mer compared to the 10-mer DUS. An inverted repeat DUS was not more efficient in transformation than DNA species containing a single or direct repeat DUS. Genome-wide analysis revealed that half of the nearly 1,500 12-mer DUS are arranged as inverted repeats predicted to be involved in rho-independent transcriptional termination or attenuation. The distribution of the uptake signal sequence required for transformation in the Pasteurellaceae was also biased towards transcriptional terminators, although to a lesser extent. In addition to assessing the intergenic location of DUS, we propose that the 10-mer identity of DUS should be extended and recognized as a 12-mer DUS. The dual role of DUS in transformation and as a structural component on RNA affecting transcription makes this a relevant model system for assessing significant roles of repeat sequences in biology.

Conserved Sequence↗