PubMed Health⌕ Search

Biomedical subjects

Marcos J Araúzo-Bravo

Publications and source records attributed to Marcos J Araúzo-Bravo.

5 recordsLinked to original sources

ReadOut: structure-based calculation of direct and indirect readout energies and specificities for protein-DNA recognition.

Protein-DNA interactions play a central role in regulatory processes at the genetic level. DNA-binding proteins recognize their targets by direct base-amino acid interactions and indirect conformational energy contribution from DNA deformations and elasticity. Knowledge-based approach based on the statistical analysis of protein-DNA complex structures has been successfully used to calculate interaction energies and specificities of direct and indirect readouts in protein-DNA recognition. Here, we have implemented the method as a webserver, which calculates direct and indirect readout energies and Z-scores, as a measure of specificity, using atomic coordinates of protein-DNA complexes. This server is freely available at http://gibk26.bse.kyutech.ac.jp/jouhou/readout/. The only input to this webserver is the Protein Data Bank (PDB) style coordinate data of atoms or the PDB code itself. The server returns total energy Z-scores, which estimate the degree of sequence specificity of the protein-DNA complex. This webserver is expected to be useful for estimating interaction energy and DNA conformation energy, and relative contributions to the specificity from direct and indirect readout. It may also be useful for checking the quality of protein-DNA complex structures, and for engineering proteins and target DNAs.

DNA↗

Dimensionality of amino acid space and solvent accessibility prediction with neural networks.

Solvent accessibility prediction from amino acid sequences has been pursued by several researchers. Such a prediction typically starts by transforming the amino acid category (or type) information into numerical representations. All twenty amino acids can be completely and uniquely represented by 20-dimensional vectors. Here, we investigate if the amino acid space defined in this way really requires twenty dimensions. We tried to develop corresponding representations in fewer dimensions. A method for searching optimal codification schema in an arbitrary space using neural networks was developed. The method is used to obtain optimal encoding of amino acids at various levels of dimensionality, and applied to optimize the amino acid codifications for the prediction of the solvent accessibility values of the proteins using feed-forward neural networks. The traditional 20-dimensional codification seems to be redundant in solving the solvent accessibility prediction problem, since a 1-dimensional codification is able to achieve almost the same degree of accuracy as the 20-dimensional codification. Optimal coding in much fewer dimensions could be used to make the predictions of accessible surface area with almost the same degree of accuracy as that obtained by a fully unique 20-dimensional coding. The 1-dimensional amino acid codification for solvent accessibility prediction obtained by a purely mathematical way based on neural networks is highly correlated with a physical property of the amino acids, namely their average solvent accessibility. The method developed to find the optimal codification is general, although the codification thus produced is dependent on the type of estimated property.

Amino Acid Sequence↗

Sequence-dependent conformational energy of DNA derived from molecular dynamics simulations: toward understanding the indirect readout mechanism in protein-DNA recognition.

Sequence dependence of DNA conformation plays a crucial role in its recognition by proteins and ligands. To clarify the relationship between sequence and conformation, it is necessary to quantify the conformational energy and specificity of DNA. Here, we make a systematic analysis of dodecamer DNA structures including all the 136 unique tetranucleotide sequences at the center by molecular dynamics simulations. Using a simplified conformational model with six parameters to describe the geometry of adjacent base pairs and harmonic potentials along these coordinates, we estimated the equilibrium conformational parameters and the harmonic potentials of mean force for the central base-pair steps from many trajectories of the simulations. This enabled us to estimate the conformational energy and the specificity for any given DNA sequence and structure. We tested our method by using sequence-structure threading to estimate the conformational energy and the Z-score as a measure of specificity for many B-DNA and A-DNA crystal structures. The average Z-scores were negative for both kinds of structures, indicating that the potential of mean force from the simulation is capable of predicting sequence specificity for the crystal structures and that it may be used to study the sequence specificity of both types of DNA. We also estimated the positional distribution of conformational energy and Z-score within DNA and showed that they are strongly position dependent. This analysis enabled us to identify particular conformations responsible for the specificity. The presented results will provide an insight into the mechanisms of DNA sequence recognition by proteins and ligands.

Base Sequence↗

Knowledge-based prediction of DNA atomic structure from nucleic sequence.

A simple knowledge-based method for DNA atomic structure prediction from nucleic sequence is presented. We used free B-DNA crystal structures to estimate the distribution of trinucleotide base pairs and tetranucleotide base-pair steps conformational coordinates. We used these distributions as a basis to predict the 3D position of the non-hydrogen atoms of the nucleic bases of any arbitrary DNA sequence of any length. The only constraint imposed was that the structure is a B-DNA one with Watson-Crick complementary base pairs. The method was tested on not seen DNA structures with sequence lengths varying from 6bp to 12bp. The obtained predictions have RMSE around 0.5 A for the translational conformational coordinates, and around 5 degrees for the rotational. For the estimation of the nucleic base non-hydrogen atom coordinates the RMSE is around 1.1 A. The knowledge-based method outperformed a technique based on genetic algorithms in the prediction of B-DNA structures.

Base Sequence↗

An improved method for statistical analysis of metabolic flux analysis using isotopomer mapping matrices with analytical expressions.

Analytical expressions were derived for calculating the sensitivities of isotopomer distribution vectors, the weighted output matrix with respect to the fluxes, and the covariance matrix for the metabolic flux analysis based on isotopomers mapping matrices (IMM). These expressions allow us to implement efficient statistical analysis, avoiding the time-consuming Monte Carlo techniques for estimating the confidence interval of the fluxes. The analytical expressions are also useful in implementing a faster design of experiment, which requires repetitive computation of the covariance matrix that is not straightforward to make in practice with the numerical techniques based on the conventional IMM. The proposed method was applied for analyzing the central carbon metabolism of the mixotrophically cultivated Synechocystis sp. PCC6803, and the confidence intervals of all its fluxes were computed based on the isotopomer distribution measured using NMR and GC-MS. It was found that the best feasible mixture for labeling experiment is 70% unlabeled, 10% [U-13C] and 20% [1,2-13C2] labeled glucose to obtain the most reliable metabolic fluxes.

Biomass↗