PubMed Health⌕ Search

PubMed · 15673715

Quantitative evaluation of protein-DNA interactions using an optimized knowledge-based potential.

Abstract

Computational evaluation of protein-DNA interaction is important for the identification of DNA-binding sites and genome annotation. It could validate the predicted binding motifs by sequence-based approaches through the calculation of the binding affinity between a protein and DNA. Such an evaluation should take into account structural information to deal with the complicated effects from DNA structural deformation, distance-dependent multi-body interactions and solvation contributions. In this paper, we present a knowledge-based potential built on interactions between protein residues and DNA tri-nucleotides. The potential, which explicitly considers the distance-dependent two-body, three-body and four-body interactions between protein residues and DNA nucleotides, has been optimized in terms of a Z-score. We have applied this knowledge-based potential to evaluate the binding affinities of zinc-finger protein-DNA complexes. The predicted binding affinities are in good agreement with the experimental data (with a correlation coefficient of 0.950). On a larger test set containing 48 protein-DNA complexes with known experimental binding free energies, our potential has achieved a high correlation coefficient of 0.800, when compared with the experimental data. We have also used this potential to identify binding motifs in DNA sequences of transcription factors (TF). The TFs in 79.4% of the known TF-DNA complexes have accurately found their native binding sequences from a large pool of DNA sequences. When tested in a genome-scale search for TF-binding motifs of the cyclic AMP regulatory protein (CRP) of Escherichia coli, this potential ranks all known binding motifs of CRP in the top 15% of all candidate sequences.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhijie Liu, Fenglou Mao, Jun-tao Guo, Bo Yan, Peng Wang, Youxing Qu, Ying Xu. 2005-01-26. Quantitative evaluation of protein-DNA interactions using an optimized knowledge-based potential.. https://doi.org/10.1093/nar%2Fgki204

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence↗

Highly sensitive detection of cytotoxicity using a modified HSP70B' promoter.

We have previously found that the DNA fragment from nucleotides (nts) -287 to +110 in the HSP70B' gene is a functional promoter responding to Cadmium Chloride-induced cytotoxicity (Wada et al., Biotechnol Bioeng, 92, 410-415, 2005). In order to increase the cytotoxic response of this promoter, we first determined the location of the cytotoxic responding element (CRE) and then constructed tandem repeats of the CRE in front of the HSP70B' promoter. 5'- and 3'-deletion analysis revealed that the DNA fragment from nts -192 to -56 in the HSP70B' gene induces a significant response to cytotoxicity. When the AP-1 binding site in this region was mutated, the basal activity of HSP70B' gene promoter decreased but the cytotoxic response was unchanged. Thus, the CRE is located in nts -192 to -56 in the HSP70B' promoter, and the AP-1 binding site is not essential for the cytotoxic response. In addition, cells transfected with a luciferase construct carrying three tandem repeats of the CRE upstream of the HSP70B' promoter and containing AP-1 binding site mutation, showed a 2.28-fold higher response than that of no repeats. Moreover, the detection limit of Cadmium Chloride in the cells was 382 pmol/mL. Thus, highly sensitive sensor cells for Cadmium Chloride can be constructed using a HSP70B' promoter construct containing upstream tandem repeats of the CRE and mutation of the AP-1 binding site.

Base Sequence↗

Synaptotagmin XI as a candidate gene for susceptibility to schizophrenia.

Synaptotagmin XI (Syt11) is a member of the synaptotagmin family, which is localized in cells either in synaptic vesicles or the cellular membrane, and is known to act as a calcium sensor. The Syt11 gene is located on chromosome locus 1q21-q22, which was previously reported as a major susceptibility locus of familial schizophrenia. Here, we present evidence for an association between the number of 33-bp repeats in the promoter region of the Syt11 gene and schizophrenia. We found that the transcriptional activity of the gene is affected by the number of 33-bp repeats, which include an Sp1 binding site, suggesting that the excessive expression of Syt11 can be associated with schizophrenia. Another (single nucleotide) polymorphism in the Syt11 5'UTR region, where the potent transcription factor YY1 can bind, also affects the transcriptional activity of Syt11.

Base Sequence↗