PubMed Health⌕ Search

PubMed · 11435394

Identifying functional elements by comparative DNA sequence analysis.

Abstract

The source did not provide an abstract. Follow the original record for more information.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M Tompa. 2001. Identifying functional elements by comparative DNA sequence analysis.. https://doi.org/10.1101/gr.197101

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Base-pair resolution conservation data improves cell type specific sequence-to-expression prediction.

MOTIVATION: Genomic sequence-to-activity models can decipher gene regulatory mechanisms and predict the functional impact of regulatory variants. However, current models struggle to integrate information from sequences outside promoters, especially information from cell type specific regulatory elements. RESULTS: Here, we propose incorporating base-pair resolution evolutionary conservation data into genomic sequence-to-expression predictors. We explore two training strategies-training from scratch or fine-tuning an existing sequence-only model with additional conservation input. We find that in both cases, base-pair resolution conservation data improves cell type specific sequence-to-expression prediction, with training from scratch yielding the greatest benefit. The improvement in cell type specific expression prediction can be attributed in part to the fact that models trained on sequence and conservation data learn to better recognize cell type specific regulatory elements than models trained on sequence alone. AVAILABILITY: Code is available at https://github.com/ni-lab/basenji-phyloP.

Conserved Sequence↗

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence↗

Functional analysis of the four DNA binding domains of replication protein A. The role of RPA2 in ssDNA binding.

Replication Protein A (RPA), the heterotrimeric single-stranded DNA (ssDNA)-binding protein of eukaryotes, contains four ssDNA binding domains (DBDs) within its two largest subunits, RPA1 and RPA2. We analyzed the contribution of the four DBDs to ssDNA binding affinity by assaying recombinant yeast RPA in which a single DBD (A, B, C, or D) was inactive. Inactivation was accomplished by mutating the two conserved aromatic stacking residues present in each DBD. Mutation of domain A had the most severe effect and eliminated binding to a short substrate such as (dT)12. RPA containing mutations in DBDs B and C bound to substrates (dT)12, 17, and 23 but with reduced affinity compared with wild type RPA. Mutation of DBD-D had little or no effect on the binding of RPA to these substrates. However, mutations in domain D did affect the binding to oligonucleotides larger than 23 nucleotides (nt). Protein-DNA cross-linking indicated that DBD-A (in RPA1) is essential for RPA1 to interact efficiently with substrates of 12 nt or less and that DBD-D (RPA2) interacts efficiently with oligonucleotides of 27 nt or larger. The data support a sequential model of binding in which DBD-A is responsible for the initial interaction with ssDNA, that domains A, B, and C (RPA1) contact 12-23 nt of ssDNA, and that DBD-D (RPA2) is needed for RPA to interact with substrates that are 23-27 nt in length.

Conserved Sequence↗