PubMed Health⌕ Search

Biomedical subjects

Ugo Bastolla

Publications and source records attributed to Ugo Bastolla.

14 recordsLinked to original sources

A protein evolution model with independent sites that reproduces site-specific amino acid distributions from the Protein Data Bank.

BACKGROUND: Since thermodynamic stability is a global property of proteins that has to be conserved during evolution, the selective pressure at a given site of a protein sequence depends on the amino acids present at other sites. However, models of molecular evolution that aim at reconstructing the evolutionary history of macromolecules become computationally intractable if such correlations between sites are explicitly taken into account. RESULTS: We introduce an evolutionary model with sites evolving independently under a global constraint on the conservation of structural stability. This model consists of a selection process, which depends on two hydrophobicity parameters that can be computed from protein sequences without any fit, and a mutation process for which we consider various models. It reproduces quantitatively the results of Structurally Constrained Neutral (SCN) simulations of protein evolution in which the stability of the native state is explicitly computed and conserved. We then compare the predicted site-specific amino acid distributions with those sampled from the Protein Data Bank (PDB). The parameters of the mutation model, whose number varies between zero and five, are fitted from the data. The mean correlation coefficient between predicted and observed site-specific amino acid distributions is larger than = 0.70 for a mutation model with no free parameters and no genetic code. In contrast, considering only the mutation process with no selection yields a mean correlation coefficient of = 0.56 with three fitted parameters. The mutation model that best fits the data takes into account increased mutation rate at CpG dinucleotides, yielding = 0.90 with five parameters. CONCLUSION: The effective selection process that we propose reproduces well amino acid distributions as observed in the protein sequences in the PDB. Its simplicity makes it very promising for likelihood calculations in phylogenetic studies. Interestingly, in this approach the mutation process influences the effective selection process, i.e. selection and mutation must be entangled in order to obtain effectively independent sites. This interdependence between mutation and selection reflects the deep influence that mutation has on the evolutionary process: The bias in the mutation influences the thermodynamic properties of the evolving proteins, in agreement with comparative studies of bacterial proteomes, and it also influences the rate of accepted mutations.

Algorithms↗

Biodiversity in model ecosystems, I: coexistence conditions for competing species.

This is the first of two papers where we discuss the limits imposed by competition to the biodiversity of species communities. In this first paper, we study the coexistence of competing species at the fixed point of population dynamic equations. For many simple models, this imposes a limit on the width of the productivity distribution, which is more severe the more diverse the ecosystem is (1994, Theor. Popul. Biol. 45, 227-276). Here we review and generalize this analysis, beyond the "mean-field"-like approximation of the competition matrix used in previous works, and extend it to structured food webs. In all cases analysed, we obtain qualitatively similar relations between biodiversity and competition: the narrower the productivity distribution is, the more species can stably coexist. We discuss how this result, considered together with environmental fluctuations, limits the maximal biodiversity that a trophic level can host.

Animals↗

Biodiversity in model ecosystems, II: species assembly and food web structure.

This is the second of two papers dedicated to the relationship between population models of competition and biodiversity. Here, we consider species assembly models where the population dynamics is kept far from fixed points through the continuous introduction of new species, and generalize to such models the coexistence condition derived for systems at the fixed point. The ecological overlap between species and shared preys, that we define here, provides a quantitative measure of the effective interspecies competition and of the trophic network topology. We obtain distributions of the overlap from simulations of a new model based both on immigration and speciation, and show that they are in good agreement with those measured for three large natural food webs. As discussed in the first paper, rapid environmental fluctuations, interacting with the condition for coexistence of competing species, limit the maximal biodiversity that a trophic level can host. This horizontal limitation to biodiversity is here combined with either dissipation of energy or growth of fluctuations, which in our model limit the length of food webs in the vertical direction. These ingredients yield an effective model of food webs that produce a biodiversity profile with a maximum at an intermediate trophic level, in agreement with field studies.

Animals↗

Looking at structure, stability, and evolution of proteins through the principal eigenvector of contact matrices and hydrophobicity profiles.

We review and further develop an analytical model that describes how thermodynamic constraints on the stability of the native state influence protein evolution in a site-specific manner. To this end, we represent both protein sequences and protein structures as vectors: structures are represented by the principal eigenvector (PE) of the protein contact matrix, a quantity that resembles closely the effective connectivity of each site; sequences are represented through the "interactivity" of each amino acid type, using novel parameters that are correlated with hydropathy scales. These interactivity parameters are more strongly correlated than the other hydropathy scales that we examine with: (1) the change upon mutations of the unfolding free energy of proteins with two-states thermodynamics; (2) genomic properties as the genome-size and the genome-wide GC content; (3) the main eigenvectors of the substitution matrices. The evolutionary average of the interactivity vector correlates very strongly with the PE of a protein structure. Using this result, we derive an analytic expression for site-specific distributions of amino acids across protein families in the form of Boltzmann distributions whose "inverse temperature" is a function of the PE component. We show that our predictions are in agreement with site-specific amino acid distributions obtained from the Protein Data Bank, and we determine the mutational model that best fits the observed site-specific amino acid distributions. Interestingly, the optimal model almost minimizes the rate at which deleterious mutations are eliminated by natural selection.

Amino Acid Motifs↗

Protein evolution in viral quasispecies under selective pressure: a thermodynamic and phylogenetic analysis.

The evolution of RNA viruses under antiviral pressure is characterized by high mutation rates and strong selective forces that induce extremely rapid changes of protein sequences. This makes the course of molecular evolution directly observable on time scales of months. Here we study the interplay between selection for drug resistance and selection for thermodynamic stability in the protease (PR) and the reverse transcriptase (RT) of human immunodeficiency virus type 1 (HIV-1) clones extracted from two patients with complex treatment histories. This analysis shows that folding thermodynamic properties may fluctuate very strongly in the course of quasispecies evolution under selective pressure. For the first case, our data suggest that folding efficiency of the RT is sacrificed at the advantage of drug resistance, while the corresponding PR seems to undergo selection for thermodynamic stability in the absence of substitutions associated to resistance. The PR of the second case is not submitted to antiviral pressure during the period analyzed and seems to initiate random fluctuations that lead to the accidental increase of its folding efficiency. In summary, joint consideration of sequence evolution and thermodynamic parameters can represent a more comprehensive approach for the study of the evolution of RNA viruses.

Anti-HIV Agents↗

Principal eigenvector of contact matrices and hydrophobicity profiles in proteins.

With the aim of studying the relationship between protein sequences and their native structures, we adopted vectorial representations for both sequence and structure. The structural representation was based on the principal eigenvector of the fold's contact matrix (PE). As has been recently shown, the latter encodes sufficient information for reconstructing the whole contact matrix. The sequence was represented through a hydrophobicity profile (HP), using a generalized hydrophobicity scale that we obtained from the principal eigenvector of a residue-residue interaction matrix, and denoted as interactivity scale. Using this novel scale, we defined the optimal HP of a protein fold, and, by means of stability arguments, predicted to be strongly correlated with the PE of the fold's contact matrix. This prediction was confirmed through an evolutionary analysis, which showed that the PE correlates with the HP of each individual sequence adopting the same fold and, even more strongly, with the average HP of this set of sequences. Thus, protein sequences evolve in such a way that their average HP is close to the optimal one, implying that neutral evolution can be viewed as a kind of motion in sequence space around the optimal HP. Our results indicate that the correlation coefficient between N-dimensional vectors constitutes a natural metric in the vectorial space in which we represent both protein sequences and protein structures, which we call vectorial protein space. In this way, we define a unified framework for sequence-to-sequence, sequence-to-structure and structure-to-structure alignments. We show that the interactivity scale is nearly optimal both for the comparison of sequences to sequences and sequences to structures.

Hydrophobic and Hydrophilic Interactions↗

Prediction of site-specific amino acid distributions and limits of divergent evolutionary changes in protein sequences.

We derive an analytic expression for site-specific stationary distributions of amino acids from the structurally constrained neutral (SCN) model of protein evolution with conservation of folding stability. The stationary distributions that we obtain have a Boltzmann-like shape, and their effective temperature parameter, measuring the limit of divergent evolutionary changes at a given site, can be predicted from a site-specific topological property, the principal eigenvector of the contact matrix of the native conformation of the protein. These analytic results, obtained without free parameters, are compared with simulations of the SCN model and with the site-specific amino acid distributions obtained from the Protein Data Bank. These results also provide new insights into how the topology of a protein fold influences its designability, i.e., the number of sequences compatible with that fold. The dependence of the effective temperature on the principal eigenvector decreases for longer proteins, as a possible consequence of the fact that selection for thermodynamic stability becomes weaker in this case.

Amino Acids↗

Genomic determinants of protein folding thermodynamics in prokaryotic organisms.

Here we investigate how thermodynamic properties of orthologous proteins are influenced by the genomic environment in which they evolve. We performed a comparative computational study of 21 protein families in 73 prokaryotic species and obtained the following main results. (i) Protein stability with respect to the unfolded state and with respect to misfolding are anticorrelated. There appears to be a trade-off between these two properties, which cannot be optimized simultaneously. (ii) Folding thermodynamic parameters are strongly correlated with two genomic features, genome size and G+C composition. In particular, the normalized energy gap, an indicator of folding efficiency in statistical mechanical models of protein folding, is smaller in proteins of organisms with a small genome size and a compositional bias towards A+T. Such genomic features are characteristic for bacteria with an intracellular lifestyle. We interpret these correlations in light of mutation pressure and natural selection. A mutational bias toward A+T at the DNA level translates into a mutational bias toward more hydrophobic (and in general more interactive) proteins, a consequence of the structure of the genetic code. Increased hydrophobicity renders proteins more stable against unfolding but less stable against misfolding. Proteins with high hydrophobicity and low stability against misfolding occur in organisms with reduced genomes, like obligate intracellular bacteria. We argue that they are fixed because these organisms experience weaker purifying selection due to their small effective population sizes. This interpretation is supported by the observation of a high expression level of chaperones in these bacteria. Our results indicate that the mutational spectrum of a genome and the strength of selection significantly influence protein folding thermodynamics.

Archaea↗

Reconstruction of protein structures from a vectorial representation.

We show that the contact map of the native structure of globular proteins can be reconstructed starting from the sole knowledge of the contact map's principal eigenvector, and present an exact algorithm for this purpose. Our algorithm yields a unique contact map for all 221 globular structures of PDBselect25 of length N</=120. We also show that the reconstructed contact maps allow in turn for the accurate reconstruction of the three-dimensional structure. These results indicate that the reduced vectorial representation provided by the principal eigenvector of the contact map is equivalent to the protein structure itself. This representation is expected to provide a useful tool in bioinformatics algorithms for protein structure comparison and alignment, as well as a promising intermediate step towards protein structure prediction.

Algorithms↗

Reductive genome evolution in Buchnera aphidicola.

We have sequenced the genome of the intracellular symbiont Buchnera aphidicola from the aphid Baizongia pistacea. This strain diverged 80-150 million years ago from the common ancestor of two previously sequenced Buchnera strains. Here, a field-collected, nonclonal sample of insects was used as source material for laboratory procedures. As a consequence, the genome assembly unveiled intrapopulational variation, consisting of approximately 1,200 polymorphic sites. Comparison of the 618-kb (kbp) genome with the two other Buchnera genomes revealed a nearly perfect gene-order conservation, indicating that the onset of genomic stasis coincided closely with establishment of the symbiosis with aphids, approximately 200 million years ago. Extensive genome reduction also predates the synchronous diversification of Buchnera and its host; but, at a slower rate, gene loss continues among the extant lineages. A computational study of protein folding predicts that proteins in Buchnera, as well as proteins of other intracellular bacteria, are generally characterized by smaller folding efficiency compared with proteins of free living bacteria. These and other degenerative genomic features are discussed in light of compensatory processes and theoretical predictions on the long-term evolutionary fate of symbionts like Buchnera.

Base Sequence↗

Testing similarity measures with continuous and discrete protein models.

There are many ways to define the distance between two protein structures, thus assessing their similarity. Here, we investigate and compare the properties of five different distance measures, including the standard root-mean-square deviation (cRMSD). The performance of these measures is studied from different perspectives with two different protein models, one continuous and the other discrete. Using the continuous model, we examine the correlation between energy and native distance, and the ability of the different measures to discriminate between the two possible topologies of a three-helix bundle. Using the discrete model, we perform fits to real protein structures by minimizing different distance measures. The properties of the fitted structures are found to depend strongly on the distance measure used and the scale considered. We find that the cRMSD measure very effectively describes long-range features but is less effective with short-range features, and it correlates weakly with energy. A stronger correlation with energy and a better description of short-range properties is obtained when we use measures based on intramolecular distances.

Models, Molecular↗

Connectivity of neutral networks, overdispersion, and structural conservation in protein evolution.

Protein structures are much more conserved than sequences during evolution. Based on this observation, we investigate the consequences of structural conservation on protein evolution. We study seven of the most studied protein folds, determining that an extended neutral network in sequence space is associated with each of them. Within our model, neutral evolution leads to a non-Poissonian substitution process, due to the broad distribution of connectivities in neutral networks. The observation that the substitution process has non-Poissonian statistics has been used to argue against the original Kimura neutral theory, while our model shows that this is a generic property of neutral evolution with structural conservation. Our model also predicts that the substitution rate can strongly fluctuate from one branch to another of the evolutionary tree. The average sequence similarity within a neutral network is close to the threshold of randomness, as observed for families of sequences sharing the same fold. Nevertheless, some positions are more difficult to mutate than others. We compare such structurally conserved positions to positions conserved in protein evolution, suggesting that our model can be a valuable tool to distinguish structural from functional conservation in databases of protein families. These results indicate that a synergy between database analysis and structurally based computational studies can increase our understanding of protein evolution.

Amino Acid Sequence↗

Statistical properties of neutral evolution.

Neutral evolution is the simplest model of molecular evolution and thus it is most amenable to a comprehensive theoretical investigation. In this paper, we characterize the statistical properties of neutral evolution of proteins under the requirement that the native state remains thermodynamically stable, and compare them to the ones of Kimura's model of neutral evolution. Our study is based on the Structurally Constrained Neutral (SCN) model which we recently proposed. We show that, in the SCN model, the substitution rate decreases as longer time intervals are considered. Fluctuations from one branch of the evolutionary tree to another are strong, leading to a non-Poissonian statistics for the substitution process. Such strong fluctuations are in part due to the fact that neutral substitution rates for individual residues are strongly correlated for most residue pairs. Interestingly, structurally conserved residues, characterized by a much below average substitution rate, are also much less correlated to other residues and evolve in a much more regular way. Our results can improve methods aimed at distinguishing between neutral and adaptive substitutions as well as methods for computing the expected number of substitutions occurred since the divergence of two protein sequences. In particular, we compute the minimal sequence similarity below which no information about the evolutionary divergence of the compared sequences can be obtained.

Amino Acid Substitution↗

Lack of self-averaging in neutral evolution of proteins.

We simulate neutral evolution of proteins imposing conservation of the thermodynamic stability of the native state in the framework of an effective model of folding thermodynamics. This procedure generates evolutionary trajectories in sequence space which share two universal features for all of the examined proteins. First, the number of neutral mutations fluctuates broadly from one sequence to another, leading to a non-Poissonian substitution process. Second, the number of neutral mutations displays strong correlations along the trajectory, thus causing the breakdown of self-averaging of the resulting evolutionary substitution process.

Cytochrome c Group↗