PubMed HealthSearch

Biomedical subjects

A Sarai

Publications and source records attributed to A Sarai.

At least 19 recordsLinked to original sources

Construction of an artificial tandem protein of the c-Myb DNA-binding domain and analysis of its DNA binding specificity.

An artificial tandem protein was generated using the third repeat of the c-Myb DNA-binding domain, and its DNA binding affinity and specificity were analyzed by a filter binding assay, isothermal titration calorimetry, and surface plasmon resonance. Although this artificial protein had the proper secondary structure, which is similar to the third repeat by itself, it could not bind to the expected base sequences specifically. Compared with the successful results of the zinc finger fusion proteins with novel sequence specificities, the cooperativity between the adjacent repeats, observed in the c-Myb-DNA complex, should also be required for the DNA recognition by the artificial tandem protein. Using the previous analyses of the DNA binding specificities by Myb homologous proteins, the differences in the DNA recognition mechanisms between the animal and plant Myb domains are also discussed.

Amino Acid Sequence

Kinetic analysis of DNA binding by the c-Myb DNA-binding domain using surface plasmon resonance.

Kinetics of the interaction of the c-Myb DNA-binding domain (R2R3) with its target DNA have been analyzed by surface plasmon resonance measurements. The association and dissociation rate constants between the standard R2R3, the Cys130 mutant substituted with Ile, and the cognate DNA are 2.3x10(5) M(-1) s(-1) and 2.6x10(-3) s(-1) at pH 7.5 and 20 degrees C, respectively. Kinetic analyses of the binding of the standard R2R3 to the non-cognate DNAs and those of the R2R3 mutant proteins to the cognate DNA showed that the reduction of the binding affinity was mainly due to an increase in the dissociation rate.

Animals

The mutant RecA proteins, RecAR243Q and RecAK245N, exhibit defective DNA binding in homologous pairing.

In homologous pairing, the RecA protein sequentially binds to single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA), aligning the two DNA molecules within the helical nucleoprotein filament. To identify the DNA binding region, which stretches from the outside to the inside of the filament, we constructed two mutant RecA proteins, RecAR243Q and RecAK245N, with the amino acid substitutions of Arg243 to Gln and Lys245 to Asn, respectively. These amino acids are exposed to the solvent in the crystal structure of the RecA protein and are located in the central domain, which is believed to be the catalytic center of the homologous pairing activity. The mutations of Arg243 to Gln (RecAR243Q) and Lys245 to Asn (RecAK245N) impair the repair of UV-damaged DNA in vivo and cause defective homologous pairing of ssDNA and dsDNA in vitro. Although RecAR243Q is only slightly defective and RecAK245N is completely proficient in ssDNA binding to form the presynaptic filament, both mutant RecA proteins are defective in the formation of the three-component complex including ssDNA, dsDNA, and RecA protein. The ability to form dsDNA from complementary single strands is also defective in both RecAR243Q and RecAK245N. These results suggest that the region including Arg243 and Lys245 may be involved in the path of secondary DNA binding to the presynaptic filament.

Base Pairing

Structure-based prediction of DNA target sites by regulatory proteins.

Regulatory proteins play a critical role in controlling complex spatial and temporal patterns of gene expression in higher organism, by recognizing multiple DNA sequences and regulating multiple target genes. Increasing amounts of structural data on the protein-DNA complex provides clues for the mechanism of target recognition by regulatory proteins. The analyses of the propensities of base-amino acid interactions observed in those structural data show that there is no one-to-one correspondence in the interaction, but clear preferences exist. On the other hand, the analysis of spatial distribution of amino acids around bases shows that even those amino acids with strong base preference such as Arg with G are distributed in a wide space around bases. Thus, amino acids with many different geometries can form a similar type of interaction with bases. The redundancy and structural flexibility in the interaction suggest that there are no simple rules in the sequence recognition, and its prediction is not straightforward. However, the spatial distributions of amino acids around bases indicate a possibility that the structural data can be used to derive empirical interaction potentials between amino acids and bases. Such information extracted from structural databases has been successfully used to predict amino acid sequences that fold into particular protein structures. We surmised that the structures of protein-DNA complexes could be used to predict DNA target sites for regulatory proteins, because determining DNA sequences that bind to a particular protein structure should be similar to finding amino acid sequences that fold into a particular structure. Here we demonstrate that the structural data can be used to predict DNA target sequences for regulatory proteins. Pairwise potentials that determine the interaction between bases and amino acids were empirically derived from the structural data. These potentials were then used to examine the compatibility between DNA sequences and the protein-DNA complex structure in a combinatorial "threading" procedure. We applied this strategy to the structures of protein-DNA complexes to predict DNA binding sites recognized by regulatory proteins. To test the applicability of this method in target-site prediction, we examined the effects of cognate and noncognate binding, cooperative binding, and DNA deformation on the binding specificity, and predicted binding sites in real promoters and compared with experimental data. These results show that target binding sites for several regulatory proteins are successfully predicted, and our data suggest that this method can serve as a powerful tool for predicting multiple target sites and target genes for regulatory proteins.

Base Sequence

ProTherm: Thermodynamic Database for Proteins and Mutants.

The first release of the Thermodynamic Database for Proteins and Mutants (ProTherm) contains more than 3300 data of several thermodynamic parameters for wild type and mutant proteins. Each entry includes numerical data for unfolding Gibbs free energy change, enthalpy change, heat capacity change, transition temperature, activity etc., which are important for understanding the mechanism of protein stability. ProTherm also includes structural information such as secondary structure and solvent accessibility of wild type residues, and experimental methods and other conditions. A WWW interface enables users to search data based on various conditions with different sorting options for outputs. Further, ProTherm is cross-linked with NCBI PUBMED literature database, Protein Mutant Database, Enzyme Code and Protein Data Bank structural database. Moreover, all the mutation sites associated with each PDB structure are automatically mapped and can be directly viewed through 3DinSight developed in our laboratory. The database is available at the URL, http://www.rtc.riken.go.jp/protherm.htm l

Calorimetry

Role of structural and sequence information in the prediction of protein stability changes: comparison between buried and partially buried mutations.

Predicting mutation-induced changes in protein stability is one of the greatest challenges in molecular biology. In this work, we analyzed the correlation between stability changes caused by buried and partially buried mutations and changes in 48 physicochemical, energetic and conformational properties. We found that properties reflecting hydrophobicity strongly correlated with stability of buried mutations, and there was a direct relation between the property values and the number of carbon atoms. Classification of mutations based on their location within helix, strand, turn or coil segments improved the correlation of mutations with stability. Buried mutations within beta-strand segments correlated better than did those in alpha-helical segments, suggesting stronger hydrophobicity of the beta-strands. The stability changes caused by partially buried mutations in ordered structures (helix, strand and turn) correlated most strongly and were mainly governed by hydrophobicity. Due to the disordered nature of coils, the mechanism underlying their stability differed from that of the other secondary structures: the stability changes due to mutations within the coil were mainly influenced by the effects of entropy. Further classification of mutations within coils, based on their hydrogen-bond forming capability, led to much stronger correlations. Hydrophobicity was the major factor in determining the stability of buried mutations, whereas hydrogen bonds, other polar interactions and hydrophobic interactions were all important determinants of the stability of partially buried mutations. Information about local sequence and structural effects were more important for the prediction of stability changes caused by partially buried mutations than for buried mutations; they strengthened correlations by an average of 27% among all data sets.

Amino Acids

Crystallization and preliminary X-ray analysis of wild-type and V103L mutant Myb R2 DNA-binding domain.

The R2 subdomain of the mouse c-Myb DNA-binding domain and its V103L mutant have been crystallized by the vapour-diffusion method using highly concentrated sodium citrate at pH 6.8 as a precipitant. Using ammonium sulfate as precipitant in MES buffer only produced crystals for the mutant R2. All crystals are isomorphous and belong to space group P212121. The unit-cell dimensions for wild-type R2 crystals grown from sodium citrate precipitant are a = 28.83, b = 40.18, c = 49.23 A. Crystals contain one R2 molecule per asymmetric unit. They are stable during 3 d exposure to X-rays and diffract to 1.37-1.45 A resolution.

Animals

Unique mode of GCC box recognition by the DNA-binding domain of ethylene-responsive element-binding factor (ERF domain) in plant.

Ethylene-responsive element-binding proteins (EREBPs)have novel DNA-binding domains (ERF domains), which are widely conserved in plants, and interact specifically with sequences containing AGCCGCC motifs (GCC box). Deletion experiments show that some flanking region at the N terminus of the conserved 59-amino acid ERF domain is required for stable binding to the GCC box. Three ERF domain-containing fragments of EREBP2, EREBP4, and AtERF1 from tobacco and Arabidopsis, bind to the sequence containing the GCC box with a high binding affinity in the pM range. The high affinity binding is conferred by a monomeric ERF domain fragment, and DNA truncation experiments show that only 11-base pair DNA containing the GCC box is sufficient for stable ERF domain interaction. Systematic DNA mutation analyses demonstrate that the specific amino acid contacts are confined within the 6-base pair GCCGCC region of the GCC box, and the first G, the fourth G, and the sixth C exhibit highest binding specificity common in all three ERF domain-containing fragments studied. Other bases within the GCC box exhibit modulated binding specificity varying from protein to protein, implying that these positions are important for differential binding by different EREBPs. The conserved N-terminal half is likely responsible for formation of a stable complex with the GCC box and the divergent C-terminal half for modulating the specificity.

Amino Acid Sequence

Thermodynamics of specific and non-specific DNA binding by the c-Myb DNA-binding domain.

The thermodynamics of the c-Myb DNA-binding domain (R2R3) interaction with its target DNA have been analyzed using isothermal titration calorimetry and amino acid mutagenesis. The enthalpy of association between the standard R2R3, the Cys130 mutant substituted with Ile, and the cognate DNA is -12.5 (+/- 0.1) kcal mol-1 at pH 7.5 and at 20 degrees C, and this interaction is enthalpically driven throughout the physiological temperature range. In order to understand the DNA recognition mechanism, several pairs of interactions were investigated using single and multiple-base alterations with single and multiple-amino acid substituted mutants. The interactions between the standard R2R3 and many non-cognate DNAs were accompanied by binding enthalpy changes and heat capacity changes, although their affinities were reduced. The roles of the electrostatic interactions in binding to the cognate and the non-cognate DNAs were also analyzed from the dependency of the thermodynamic parameters on the salt concentration. The heat capacity change was found to be significantly dependent upon the salt concentration. Several mutant proteins bound to the multiple-base altered DNA with very small enthalpy changes, although they bound to the cognate and the single-base altered DNAs with detectable enthalpy and heat capacity changes. From the thermodynamic cycles derived from the DNA binding of the amino acid substituted R2R3 to the base substituted DNA duplexes, the individual thermodynamic mechanisms of the specific DNA recognition of R2R3 were dissected. The local folding mechanism was highlighted by the substitution of Pro with either Gly or Ala at the linker between R2 and R3. The characteristic thermodynamic features of specific and non-specific DNA binding are discussed.

Amino Acid Substitution

3DinSight: an integrated relational database and search tool for the structure, function and properties of biomolecules.

MOTIVATION: Although a large amount of information on the structure, function and properties of biomolecules is becoming available, it is difficult to understand the relationship between them. Thus, we have attempted to create an integrated relational database, search and visualization tool, 3DinSight, to help researchers to gain insight into their relationship. RESULTS: We have gathered data on the structure, function and properties of biomolecules, and implemented them into a relational database system. The structural data contain several subset data such as protein homologues, protein-DNA complex, in order to enable searching within a specific class of data. The functional data include motif sequence and mutation data of proteins. Also, various amino acid properties are implemented as a relational table. The World Wide Web (WWW) interfaces enable users to carry out various kinds of searches among these data. The locations of motif sequences and mutations are automatically mapped on the structure, and visualized in three-dimensional (3D) space by interactive viewers, VRML (Virtual Reality Modeling Language) and RasMol. In the case of VRML, the mapped 3D objects are hyper-linked to the corresponding document data. Also, amino acid properties, linked with structure, functional and mutation sites, can be displayed as graph plots. AVAILABILITY: 3DinSight is freely accessible through the Internet (http://www.rtc.riken.go.jp/3DinSight.h tml). CONTACT: sarai@rtc.riken.go.jp

Computer Communication Networks

Investigation of the pyrimidine preference by the c-Myb DNA-binding domain at the initial base of the consensus sequence.

The principal determinant of the pyrimidine preference by the c-Myb DNA-binding domain at the initial base of the consensus sequence was investigated by mutation of both the protein and the DNA base pairs, with analysis by a filter binding assay. Amino acid residue 187 was revealed to interact with the pyrimidine base position, as estimated from our previous complex structure. Unexpectedly, since the pyrimidine preference is retained even in the Gly187 mutant, the principal origin of the base specificity should not occur via the direct-readout mechanism, but by an indirect-readout mechanism, namely in the intrinsic "bendability" of the pyrimidine-purine step of the DNA duplex. A significant but rather small positive base pair roll is detectable in the conformation of DNA in complex with the c-Myb DNA-binding domain. Following the conventional chemical rules of the direct-readout mechanism, amino acid mutagenesis at position 187 yielded several new base preferences for the protein.

Amino Acid Sequence

Identification of indispensable residues for specific DNA-binding in the imperfect tandem repeats of c-Myb R2R3.

The individual repeats, R2 and R3, of the minimum specific DNA-binding domain (R2R3) of c-Myb have very similar structures, with a helix-turn-helix variation motif, although their sequence identity in the tandem repeats is only 31%. From previous mutational and structural studies, the third helices in both repeats were shown to directly recognize the specific base sequence, PyAACG/TG. In order to elucidate the reason for the imperfection of the tandem repeats at amino acid positions other than the recognition helices, a series of R2R3 mutants was generated by swapping the helices and the N-terminus in R2 to those in R3. Consequently, the sequence composing the first helix of R2 was found to be essential for specific DNA-binding, in addition to the third recognition helix of R2. Further mutational studies revealed that the only indispensable residues in the first helix are Val103 and Val1O7, which are involved in the hydrophobic core of R2. These residues do not directly interact with the DNA, but they contribute to the correct formation of helix 1 and the characteristic packing of R2, which is slightly different from that of R3, and are required for specific base recognition through strong cooperativity with R3.

Binding Sites

Effect of chemical modification of oligohomopyrimidine on triplex formation: thermodynamic and kinetic studies.

To investigate the effect of chemical modification of the third strand on the stability of triplex DNA, we have examined the thermodynamic properties of the triplex formation between a 23-mer double-stranded homopurine-homopyrimidine and each of five kinds of 15-mer chemically modified single-stranded homopyrimidines by isothermal titration calorimetry, and the kinetic properties by interaction analysis system. The modifications of the third strand included two base modifications, two sugar moiety modifications, and one phosphate backbone modification. The thermodynamic and kinetic parameters for the triplex formation were similar in magnitude among the two base-modified and two sugar-modified single strands. By contrast, the binding constant for the triplex formation with the single strand with phosphorothioate backbone was more than ten times as small as that for the other triplex formation. On the basis of the kinetic analyses, the single strand with phosphorothioate backbone was more difficult to associate with and easier to dissociate from the target double strand than the other single strands, which resulted in the much smaller binding constant.

Base Sequence

A possible role of the C-terminal domain of the RecA protein. A gateway model for double-stranded DNA binding.

According to the crystal structure, the RecA protein has a domain near the C terminus consisting of amino acid residues 270-328 (from the N terminus). Our model building pointed out the possibility that this domain is a part of "gateway" through which double-stranded DNA finds a path for direct contact with single-stranded DNA within a presynaptic RecA filament in the search for homology. To test this possible function of the domain, we made mutant RecA proteins by site-directed single (or double, in one case) replacement of 2 conserved basic amino acid residues and 5 among 9 nonconserved basic amino acid residues in the domain. Replacement of either of the 2 conserved amino acid residues caused deficiencies in repair of UV-damaged DNA, an in vivo function of RecA protein, whereas the replacement of most (except one) of the tested nonconserved ones gave little or no effect. Purified mutant RecA proteins showed no (or only slight) deficiencies in the formation of presynaptic filaments as assessed by various assays. However, presynaptic filaments of both proteins that had replacement of a conserved amino acid residue had significant defects in binding to and pairing with duplex DNA (secondary binding). These results are consistent with our model that the conserved amino acid residues in the C-terminal domain have a direct role in double-stranded DNA binding and that they constitute a part of a gateway for homologous recognition.

Adenosine Triphosphate

Binding site analysis of c-Myb: screening of potential binding sites by using the mutation matrix derived from systematic binding affinity measurements.

The c-Myb oncoprotein is known to bind to multiple sites in the promoters of target genes. We have developed a protocol to screen the binding site of c-Myb by using the systematic binding data derived form measurements of binding affinity for oligonucleotide containing a known Myb-binding site and its complete single mutants. We first applied the method to predict the binding affinity for the known binding sites and compared with available experimental data. The predicated binding sites agree with many putative binding sites of known target promoters. However, there are some binding sites not predicated by the analysis. These sequences deviate from the consensus sequence derived from the binding analyses. In the light of the structure of Myb-DNA complex, these results indicate that different DNA-binding modes may be used by c-Myb to recognize different classes of binding sites. We also screened the sequence database for potential Myb-binding sites, and found sequences of several promoters that have not been identified experimentally but could be the target for c-Myb.

Base Sequence