PubMed Health⌕ Search

Biomedical subjects

Nicolas Lartillot

Publications and source records attributed to Nicolas Lartillot.

12 recordsLinked to original sources

A Bayesian compound stochastic process for modeling nonstationary and nonhomogeneous sequence evolution.

Variations of nucleotidic composition affect phylogenetic inference conducted under stationary models of evolution. In particular, they may cause unrelated taxa sharing similar base composition to be grouped together in the resulting phylogeny. To address this problem, we developed a nonstationary and nonhomogeneous model accounting for compositional biases. Unlike previous nonstationary models, which are branchwise, that is, assume that base composition only changes at the nodes of the tree, in our model, the process of compositional drift is totally uncoupled from the speciation events. In addition, the total number of events of compositional drift distributed across the tree is directly inferred from the data. We implemented the method in a Bayesian framework, relying on Markov Chain Monte Carlo algorithms, and applied it to several nucleotidic data sets. In most cases, the stationarity assumption was rejected in favor of our nonstationary model. In addition, we show that our method is able to resolve a well-known artifact. By Bayes factor evaluation, we compared our model with 2 previously developed nonstationary models. We show that the coupling between speciations and compositional shifts inherent to branchwise models may lead to an overparameterization, resulting in a lesser fit. In some cases, this leads to incorrect conclusions, concerning the nature of the compositional biases. In contrast, our compound model more flexibly adapts its effective number of parameters to the data sets under investigation. Altogether, our results show that accounting for nonstationary sequence evolution may require more elaborate and more flexible models than those currently used.

Animals↗

A maximum likelihood framework for protein design.

BACKGROUND: The aim of protein design is to predict amino-acid sequences compatible with a given target structure. Traditionally envisioned as a purely thermodynamic question, this problem can also be understood in a wider context, where additional constraints are captured by learning the sequence patterns displayed by natural proteins of known conformation. In this latter perspective, however, we still need a theoretical formalization of the question, leading to general and efficient learning methods, and allowing for the selection of fast and accurate objective functions quantifying sequence/structure compatibility. RESULTS: We propose a formulation of the protein design problem in terms of model-based statistical inference. Our framework uses the maximum likelihood principle to optimize the unknown parameters of a statistical potential, which we call an inverse potential to contrast with classical potentials used for structure prediction. We propose an implementation based on Markov chain Monte Carlo, in which the likelihood is maximized by gradient descent and is numerically estimated by thermodynamic integration. The fit of the models is evaluated by cross-validation. We apply this to a simple pairwise contact potential, supplemented with a solvent-accessibility term, and show that the resulting models have a better predictive power than currently available pairwise potentials. Furthermore, the model comparison method presented here allows one to measure the relative contribution of each component of the potential, and to choose the optimal number of accessibility classes, which turns out to be much higher than classically considered. CONCLUSION: Altogether, this reformulation makes it possible to test a wide diversity of models, using different forms of potentials, or accounting for other factors than just the constraint of thermodynamic stability. Ultimately, such model-based statistical analyses may help to understand the forces shaping protein sequences, and driving their evolution.

Amino Acid Sequence↗

Assessing site-interdependent phylogenetic models of sequence evolution.

In recent works, methods have been proposed for applying phylogenetic models that allow for a general interdependence between the amino acid positions of a protein. As of yet, such models have focused on site interdependencies resulting from sequence-structure compatibility constraints, using simplified structural representations in combination with a set of statistical potentials. This structural compatibility criterion is meant as a proxy for sequence fitness, and the methods developed thus far can incorporate different site-interdependent fitness proxies based on other measurements. However, no methods have been proposed for comparing and evaluating the adequacy of alternative fitness proxies in this context, or for more general comparisons with canonical models of protein evolution. In the present work, we apply Bayesian methods of model selection-based on numerical calculations of marginal likelihoods and posterior predictive checks-to evaluate models encompassing the site-interdependent framework. Our application of these methods indicates that considering site-interdependencies, as done here, leads to an improved model fit for all data sets studied. Yet, we find that the use of pairwise contact potentials alone does not suitably account for across-site rate heterogeneity or amino acid exchange propensities; for such complexities, site-independent treatments are still called for. The most favored models combine the use of statistical potentials with a suitably rich site-independent model. Altogether, the methodology employed here should allow for a more rigorous and systematic exploration of different ways of modeling explicit structural constraints, or any other site-interdependent criterion, while best exploiting the richness of previously proposed models.

Bayes Theorem↗

Computing Bayes factors using thermodynamic integration.

In the Bayesian paradigm, a common method for comparing two models is to compute the Bayes factor, defined as the ratio of their respective marginal likelihoods. In recent phylogenetic works, the numerical evaluation of marginal likelihoods has often been performed using the harmonic mean estimation procedure. In the present article, we propose to employ another method, based on an analogy with statistical physics, called thermodynamic integration. We describe the method, propose an implementation, and show on two analytical examples that this numerical method yields reliable estimates. In contrast, the harmonic mean estimator leads to a strong overestimation of the marginal likelihood, which is all the more pronounced as the model is higher dimensional. As a result, the harmonic mean estimator systematically favors more parameter-rich models, an artefact that might explain some recent puzzling observations, based on harmonic mean estimates, suggesting that Bayes factors tend to overscore complex models. Finally, we apply our method to the comparison of several alternative models of amino-acid replacement. We confirm our previous observations, indicating that modeling pattern heterogeneity across sites tends to yield better models than standard empirical matrices.

Amino Acid Sequence↗

Multipolar consensus for phylogenetic trees.

Collections of phylogenetic trees are usually summarized using consensus methods. These methods build a single tree, supposed to be representative of the collection. However, in the case of heterogeneous collections of trees, the resulting consensus may be poorly resolved (strict consensus, majority-rule consensus, ...), or may perform arbitrary choices among mutually incompatible clades, or splits (greedy consensus). Here, we propose an alternative method, which we call the multipolar consensus (MPC). Its aim is to display all the splits having a support above a predefined threshold, in a minimum number of consensus trees, or poles. We show that the problem is equivalent to a graph-coloring problem, and propose an implementation of the method. Finally, we apply the MPC to real data sets. Our results indicate that, typically, all the splits down to a weight of 10% can be displayed in no more than 4 trees. In addition, in some cases, biologically relevant secondary signals, which would not have been present in any of the classical consensus trees, are indeed captured by our method, indicating that the MPC provides a convenient exploratory method for phylogenetic analysis. The method was implemented in a package freely available at http://www.lirmm.fr/~cbonnard/MPC.html

Algorithms↗

Site interdependence attributed to tertiary structure in amino acid sequence evolution.

Standard likelihood-based frameworks in phylogenetics consider the process of evolution of a sequence site by site. Assuming that sites evolve independently greatly simplifies the required calculations. However, this simplification is known to be incorrect in many cases. Here, a computational method that allows for general dependence between sites of a sequence is investigated. Using this method, measures acting as sequence fitness proxies can be considered over a phylogenetic tree. In this work, a set of statistically derived amino acid pairwise potentials, developed in the context of protein threading, is used to account for what we call the structural fitness of a sequence. We describe a model combining statistical potentials with an empirical amino acid substitution matrix. We propose such a combination as a useful way of capturing the complexity of protein evolution. Finally, we outline features of the model using three datasets and show the approach's sensitivity to different tree topologies.

Amino Acid Sequence↗

Multigene analyses of bilaterian animals corroborate the monophyly of Ecdysozoa, Lophotrochozoa, and Protostomia.

Almost a decade ago, a new phylogeny of bilaterian animals was inferred from small-subunit ribosomal RNA (rRNA) that claimed the monophyly of two major groups of protostome animals: Ecdysozoa (e.g., arthropods, nematodes, onychophorans, and tardigrades) and Lophotrochozoa (e.g., annelids, molluscs, platyhelminths, brachiopods, and rotifers). However, it received little additional support. In fact, several multigene analyses strongly argued against this new phylogeny. These latter studies were based on a large amount of sequence data and therefore showed an apparently strong statistical support. Yet, they covered only a few taxa (those for which complete genomes were available), making systematic artifacts of tree reconstruction more probable. Here we expand this sparse taxonomic sampling and analyze a large data set (146 genes, 35,371 positions) from a diverse sample of animals (35 species). Our study demonstrates that the incongruences observed between rRNA and multigene analyses were indeed due to long-branch attraction artifacts, illustrating the enormous impact of systematic biases on phylogenomic studies. A refined analysis of our data set excluding the most biased genes provides strong support in favor of the new animal phylogeny and in addition suggests that urochordates are more closely related to vertebrates than are cephalochordates. These findings have important implications for the interpretation of morphological and genomic data.

Animals↗

A Bayesian mixture model for across-site heterogeneities in the amino-acid replacement process.

Most current models of sequence evolution assume that all sites of a protein evolve under the same substitution process, characterized by a 20 x 20 substitution matrix. Here, we propose to relax this assumption by developing a Bayesian mixture model that allows the amino-acid replacement pattern at different sites of a protein alignment to be described by distinct substitution processes. Our model, named CAT, assumes the existence of distinct processes (or classes) differing by their equilibrium frequencies over the 20 residues. Through the use of a Dirichlet process prior, the total number of classes and their respective amino-acid profiles, as well as the affiliations of each site to a given class, are all free variables of the model. In this way, the CAT model is able to adapt to the complexity actually present in the data, and it yields an estimate of the substitutional heterogeneity through the posterior mean number of classes. We show that a significant level of heterogeneity is present in the substitution patterns of proteins, and that the standard one-matrix model fails to account for this heterogeneity. By evaluating the Bayes factor, we demonstrate that the standard model is outperformed by CAT on all of the data sets which we analyzed. Altogether, these results suggest that the complexity of the pattern of substitution of real sequences is better captured by the CAT model, offering the possibility of studying its impact on phylogenetic reconstruction and its connections with structure-function determinants.

Amino Acid Sequence↗

The expression of a caudal homologue in a mollusc, Patella vulgata.

We cloned and analyzed the expression of a caudal homologue (PvuCdx) during the early development of the marine gastropod, Patella vulgata. PvuCdx is expressed at the onset of gastrulation in the ectodermal cells that constitute the posterior edge of the blastopore, as well as in the paired mesentoblasts, the stem cells that generate the posterior mesoderm of the trochophore larva. During larval stages, PvuCdx is expressed in the posterior neurectoderm of the larva, as well as in part of the mesoderm. This is the first report of the expression of a caudal gene in a lophotrochozoan species. The striking similarities with the expression of caudal in other organisms, such as chordates, suggest that a posterior expression of caudal is ancestral to Bilateria.

Amino Acid Sequence↗

Expression patterns of fork head and goosecoid homologues in the mollusc Patella vulgata supports the ancestry of the anterior mesendoderm across Bilateria.

We have characterised orthologues of the genes fork head and goosecoid in the gastropod Patella vulgata. In this species, the anterior-posterior (AP) axis is determined just before gastrulation, and leads to the specification of two mesodermal components on each side of the presumptive endoderm, one anterior (ectomesoderm), and one posterior (endomesoderm). Both fork head and goosecoid are expressed from the time the AP axis is specified, up to the end of gastrulation. fork head mRNA is detected in the whole endoderm, as well as in the anterior mesoderm, whereas goosecoid is only expressed anteriorly, in the three germ layers. The two genes are thus coexpressed in the anterior mesoderm, suggesting the latter's homology with vertebrate prechordal mesoderm. In addition, since prechordal plate is known to belong to an anterior, so called "head organiser", and since its inductive role is dependent on the function of the vertebrate fork head and goosecoid orthologues, we further suggest that the anterior mesoderm may also have a role in anterior inductive patterning in Spiralia. Finally, we propose that a mode of axial development involving two organisers, one anterior and one posterior, is ancestral to the Bilateria, and that both organisers evolved from the single head organiser of a putative hydra-like ancestor.

Animals↗

Phylogenetic analysis of the Wnt gene family. Insights from lophotrochozoan members.

The Wnt gene family encodes secreted signaling molecules that control cell fate specification, proliferation, polarity, and movements during animal development. We investigate here the evolutionary history of this large multigenic family. Wnt genes have been almost exclusively isolated from two of the three main subdivisions of bilaterian animals, the deuterostomes (which include chordates and echinoderms) and the ecdysozoans (e.g., arthropods and nematodes). However, orthology relationships between deuterostome and ecdysozoan Wnt genes, and, more generally, the phylogeny of the Wnt family, are not yet clear. We report here the isolation of several Wnt genes from two species, the annelid Platynereis dumerilii and the mollusc Patella vulgata, which both belong to the third large bilaterian clade, the lophotrochozoans (which constitute, together with ecdysozoans, the protostomes). Multiple phylogenetic analyses of these sequences with a large set of other Wnt gene sequences, in particular, the complete set of Wnt genes of human, nematode, and fly, allow us to subdivide the Wnt family into 12 subfamilies. At least nine of them were already present in the last common ancestor of all bilaterian animals, and this further highlights the genetic complexity of this ancestor. The orthology relationships we present here open new perspectives for future developmental comparisons.

Animals↗

Expression pattern of Brachyury in the mollusc Patella vulgata suggests a conserved role in the establishment of the AP axis in Bilateria.

We report the characterisation of a Brachyury ortholog (PvuBra) in the marine gastropod Patella vulgata. In this mollusc, the embryo displays an equal cleavage pattern until the 32-cell stage. There, an inductive event takes place that sets up the bilateral symmetry, by specifying one of the four initially equipotent vegetal macromeres as the posterior pole of all subsequent morphogenesis. This macromere, usually designated as 3D, will subsequently act as an organiser. We show that 3D expresses PvuBra as soon as its fate is determined. As reported for another mollusc (J. D. Lambert and L. M. Nagy (2001) Development 128, 45-56), we found that 3D determination and activity also involve the activation of the MAP kinase ERK, and we further show that PvuBra expression in 3D requires ERK activity. PvuBra expression then rapidly spreads to neighbouring cells that cleave in a bilateral fashion and whose progeny will constitute the posterior edge of the blastopore during gastrulation, suggesting a role for PvuBra in regulating cell movements and cleavage morphology in Patella. Until the completion of gastrulation, PvuBra expression is maintained at the posterior pole, and along the developing anterior-posterior axis. Comparing this expression pattern with what is known in other Bilateria, we advocate that Brachyury might have a conserved role in the regulation of anterior-posterior patterning among Bilateria, through the maintenance of a posterior growth zone, suggesting that a teloblastic mode of axis formation might be ancestral to the Bilateria.

Amino Acid Sequence↗