PubMed Health⌕ Search

Biomedical subjects

H Eduardo Roman

Publications and source records attributed to H Eduardo Roman.

11 recordsLinked to original sources

A protein evolution model with independent sites that reproduces site-specific amino acid distributions from the Protein Data Bank.

BACKGROUND: Since thermodynamic stability is a global property of proteins that has to be conserved during evolution, the selective pressure at a given site of a protein sequence depends on the amino acids present at other sites. However, models of molecular evolution that aim at reconstructing the evolutionary history of macromolecules become computationally intractable if such correlations between sites are explicitly taken into account. RESULTS: We introduce an evolutionary model with sites evolving independently under a global constraint on the conservation of structural stability. This model consists of a selection process, which depends on two hydrophobicity parameters that can be computed from protein sequences without any fit, and a mutation process for which we consider various models. It reproduces quantitatively the results of Structurally Constrained Neutral (SCN) simulations of protein evolution in which the stability of the native state is explicitly computed and conserved. We then compare the predicted site-specific amino acid distributions with those sampled from the Protein Data Bank (PDB). The parameters of the mutation model, whose number varies between zero and five, are fitted from the data. The mean correlation coefficient between predicted and observed site-specific amino acid distributions is larger than = 0.70 for a mutation model with no free parameters and no genetic code. In contrast, considering only the mutation process with no selection yields a mean correlation coefficient of = 0.56 with three fitted parameters. The mutation model that best fits the data takes into account increased mutation rate at CpG dinucleotides, yielding = 0.90 with five parameters. CONCLUSION: The effective selection process that we propose reproduces well amino acid distributions as observed in the protein sequences in the PDB. Its simplicity makes it very promising for likelihood calculations in phylogenetic studies. Interestingly, in this approach the mutation process influences the effective selection process, i.e. selection and mutation must be entangled in order to obtain effectively independent sites. This interdependence between mutation and selection reflects the deep influence that mutation has on the evolutionary process: The bias in the mutation influences the thermodynamic properties of the evolving proteins, in agreement with comparative studies of bacterial proteomes, and it also influences the rate of accepted mutations.

Algorithms↗

Looking at structure, stability, and evolution of proteins through the principal eigenvector of contact matrices and hydrophobicity profiles.

We review and further develop an analytical model that describes how thermodynamic constraints on the stability of the native state influence protein evolution in a site-specific manner. To this end, we represent both protein sequences and protein structures as vectors: structures are represented by the principal eigenvector (PE) of the protein contact matrix, a quantity that resembles closely the effective connectivity of each site; sequences are represented through the "interactivity" of each amino acid type, using novel parameters that are correlated with hydropathy scales. These interactivity parameters are more strongly correlated than the other hydropathy scales that we examine with: (1) the change upon mutations of the unfolding free energy of proteins with two-states thermodynamics; (2) genomic properties as the genome-size and the genome-wide GC content; (3) the main eigenvectors of the substitution matrices. The evolutionary average of the interactivity vector correlates very strongly with the PE of a protein structure. Using this result, we derive an analytic expression for site-specific distributions of amino acids across protein families in the form of Boltzmann distributions whose "inverse temperature" is a function of the PE component. We show that our predictions are in agreement with site-specific amino acid distributions obtained from the Protein Data Bank, and we determine the mutational model that best fits the observed site-specific amino acid distributions. Interestingly, the optimal model almost minimizes the rate at which deleterious mutations are eliminated by natural selection.

Amino Acid Motifs↗

Principal eigenvector of contact matrices and hydrophobicity profiles in proteins.

With the aim of studying the relationship between protein sequences and their native structures, we adopted vectorial representations for both sequence and structure. The structural representation was based on the principal eigenvector of the fold's contact matrix (PE). As has been recently shown, the latter encodes sufficient information for reconstructing the whole contact matrix. The sequence was represented through a hydrophobicity profile (HP), using a generalized hydrophobicity scale that we obtained from the principal eigenvector of a residue-residue interaction matrix, and denoted as interactivity scale. Using this novel scale, we defined the optimal HP of a protein fold, and, by means of stability arguments, predicted to be strongly correlated with the PE of the fold's contact matrix. This prediction was confirmed through an evolutionary analysis, which showed that the PE correlates with the HP of each individual sequence adopting the same fold and, even more strongly, with the average HP of this set of sequences. Thus, protein sequences evolve in such a way that their average HP is close to the optimal one, implying that neutral evolution can be viewed as a kind of motion in sequence space around the optimal HP. Our results indicate that the correlation coefficient between N-dimensional vectors constitutes a natural metric in the vectorial space in which we represent both protein sequences and protein structures, which we call vectorial protein space. In this way, we define a unified framework for sequence-to-sequence, sequence-to-structure and structure-to-structure alignments. We show that the interactivity scale is nearly optimal both for the comparison of sequences to sequences and sequences to structures.

Hydrophobic and Hydrophilic Interactions↗

Prediction of site-specific amino acid distributions and limits of divergent evolutionary changes in protein sequences.

We derive an analytic expression for site-specific stationary distributions of amino acids from the structurally constrained neutral (SCN) model of protein evolution with conservation of folding stability. The stationary distributions that we obtain have a Boltzmann-like shape, and their effective temperature parameter, measuring the limit of divergent evolutionary changes at a given site, can be predicted from a site-specific topological property, the principal eigenvector of the contact matrix of the native conformation of the protein. These analytic results, obtained without free parameters, are compared with simulations of the SCN model and with the site-specific amino acid distributions obtained from the Protein Data Bank. These results also provide new insights into how the topology of a protein fold influences its designability, i.e., the number of sequences compatible with that fold. The dependence of the effective temperature on the principal eigenvector decreases for longer proteins, as a possible consequence of the fact that selection for thermodynamic stability becomes weaker in this case.

Amino Acids↗

Reconstruction of protein structures from a vectorial representation.

We show that the contact map of the native structure of globular proteins can be reconstructed starting from the sole knowledge of the contact map's principal eigenvector, and present an exact algorithm for this purpose. Our algorithm yields a unique contact map for all 221 globular structures of PDBselect25 of length N</=120. We also show that the reconstructed contact maps allow in turn for the accurate reconstruction of the three-dimensional structure. These results indicate that the reduced vectorial representation provided by the principal eigenvector of the contact map is equivalent to the protein structure itself. This representation is expected to provide a useful tool in bioinformatics algorithms for protein structure comparison and alignment, as well as a promising intermediate step towards protein structure prediction.

Algorithms↗

Autoregressive processes with anomalous scaling behavior: applications to high-frequency variations of a stock market index.

We employ autoregressive conditional heteroskedasticity processes to model the probability distribution function (PDF) of high-frequency relative variations of the Standard & Poors 500 market index data, obtained at the time horizon of 1 min. The model reproduces quantitatively the shape of the PDF, characterized by a Lévy-type power-law decay around its center, followed by a crossover to a faster decay at the tails. Furthermore, it is able to reproduce accurately the anomalous decay of the central part of the PDF at larger time horizons and, by the introduction of a short-range memory, also the crossover behavior of the corresponding standard deviations and the time scale of the exponentially decaying autocorrelation function of returns displayed by the empirical data.

Journal Article↗

Statistical properties of neutral evolution.

Neutral evolution is the simplest model of molecular evolution and thus it is most amenable to a comprehensive theoretical investigation. In this paper, we characterize the statistical properties of neutral evolution of proteins under the requirement that the native state remains thermodynamically stable, and compare them to the ones of Kimura's model of neutral evolution. Our study is based on the Structurally Constrained Neutral (SCN) model which we recently proposed. We show that, in the SCN model, the substitution rate decreases as longer time intervals are considered. Fluctuations from one branch of the evolutionary tree to another are strong, leading to a non-Poissonian statistics for the substitution process. Such strong fluctuations are in part due to the fact that neutral substitution rates for individual residues are strongly correlated for most residue pairs. Interestingly, structurally conserved residues, characterized by a much below average substitution rate, are also much less correlated to other residues and evolve in a much more regular way. Our results can improve methods aimed at distinguishing between neutral and adaptive substitutions as well as methods for computing the expected number of substitutions occurred since the divergence of two protein sequences. In particular, we compute the minimal sequence similarity below which no information about the evolutionary divergence of the compared sequences can be obtained.

Amino Acid Substitution↗

Lack of self-averaging in neutral evolution of proteins.

We simulate neutral evolution of proteins imposing conservation of the thermodynamic stability of the native state in the framework of an effective model of folding thermodynamics. This procedure generates evolutionary trajectories in sequence space which share two universal features for all of the examined proteins. First, the number of neutral mutations fluctuates broadly from one sequence to another, leading to a non-Poissonian substitution process. Second, the number of neutral mutations displays strong correlations along the trajectory, thus causing the breakdown of self-averaging of the resulting evolutionary substitution process.

Cytochrome c Group↗

Autoregressive processes with exponentially decaying probability distribution functions: applications to daily variations of a stock market index.

We consider autoregressive conditional heteroskedasticity (ARCH) processes in which the variance sigma(2)(y) depends linearly on the absolute value of the random variable y as sigma(2)(y) = a+b absolute value of y. While for the standard model, where sigma(2)(y) = a + b y(2), the corresponding probability distribution function (PDF) P(y) decays as a power law for absolute value of y-->infinity, in the linear case it decays exponentially as P(y) approximately exp(-alpha absolute value of y), with alpha = 2/b. We extend these results to the more general case sigma(2)(y) = a+b absolute value of y(q), with 0 < q < 2. We find stretched exponential decay for 1 < q < 2 and stretched Gaussian behavior for 0 < q < 1. As an application, we consider the case q=1 as our starting scheme for modeling the PDF of daily (logarithmic) variations in the Dow Jones stock market index. When the history of the ARCH process is taken into account, the resulting PDF becomes a stretched exponential even for q = 1, with a stretched exponent beta = 2/3, in a much better agreement with the empirical data.

Journal Article↗

Self-avoiding walks on Sierpinski lattices in two and three dimensions.

The scaling properties of linear polymers on deterministic fractal structures, modeled by self-avoiding random walks (SAW) on Sierpinski lattices in two and three dimensions, are studied. To this end, all possible SAW configurations of N steps are enumerated exactly and averages over suitable sets of starting lattice points for the walks are performed to extract the mean quantities of interest reliably. We determine the critical exponent describing the mean end-to-end chemical distance (-)l(N) after N steps and the corresponding distribution function, P(S)(l,N). A des Cloizeaux-type relation between the exponent characterizing the asymptotic shape of the distribution, for l-->0 and N--> infinity, and the one describing the total number of SAW of N steps is suggested and supported by numerical results. These results are confronted with those obtained recently on the backbone of the incipient percolation cluster, where the corresponding exponents are very well described by a generalized des Cloizeaux relation valid for statistically self-similar structures.

Journal Article↗

"Generalized des Cloizeaux" exponent for self-avoiding walks on the incipient percolation cluster.

We study the asymptotic shape of self-avoiding random walks (SAW) on the backbone of the incipient percolation cluster in d-dimensional lattices analytically. It is generally accepted that the configurational averaged probability distribution function for the end-to-end distance r of an N step SAW behaves as a power law for r-->0. In this work, we determine the corresponding exponent using scaling arguments, and show that our suggested "generalized des Cloizeaux" expression for the exponent is in excellent agreement with exact enumeration results in two and three dimensions.

Journal Article↗