PubMed Health⌕ Search

Biomedical subjects

Peter J Waddell

Publications and source records attributed to Peter J Waddell.

7 recordsLinked to original sources

Evolutionary history of LINE-1 in the major clades of placental mammals.

BACKGROUND: LINE-1 constitutes an important component of mammalian genomes. It has a dynamic evolutionary history characterized by the rise, fall and replacement of subfamilies. Most data concerning LINE-1 biology and evolution are derived from the human and mouse genomes and are often assumed to hold for all placentals. METHODOLOGY: To examine LINE-1 relationships, sequences from the 3' region of the reverse transcriptase from 21 species (representing 13 orders across Afrotheria, Xenarthra, Supraprimates and Laurasiatheria) were obtained from whole genome sequence assemblies, or by PCR with degenerate primers. These sequences were aligned and analysed. PRINCIPAL FINDINGS: Our analysis reflects accepted placental relationships suggesting mostly lineage-specific LINE-1 families. The data provide clear support for several clades including Glires, Supraprimates, Laurasiatheria, Boreoeutheria, Xenarthra and Afrotheria. Within the afrotherian LINE-1 (AfroLINE) clade, our tree supports Paenungulata, Afroinsectivora and Afroinsectiphillia. Xenarthran LINE-1 (XenaLINE) falls sister to AfroLINE, providing some support for the Atlantogenata (Xenarthra+Afrotheria) hypothesis. SIGNIFICANCE: LINEs and SINEs make up approximately half of all placental genomes, so understanding their dynamics is an essential aspect of comparative genomics. Importantly, a tree of LINE-1 offers a different view of the root, as long edges (branches) such as that to marsupials are shortened and/or broken up. Additionally, a robust phylogeny of diverse LINE-1 is essential in testing that site-specific LINE-1 insertions, often regarded as homoplasy-free phylogenetic markers, are indeed unique and not convergent.

Animals↗

Phylogenetic methodology for detecting protein interactions.

Detecting protein-protein interactions and assigning proteins to functional complexes are key challenges of modern biology. The rise of genomics has lead to evidence that correlated patterns of presence/absence and/or fusing of proteins in any organism suggest these proteins interact. Unfortunately, methods based on such data work best with divergent genomes, whereas major sequencing efforts in vertebrates, for example, are yielding alignments of the same set of proteins sampled from the same set of taxa (species). Using vertebrate mitochondrial genomes to illustrate a novel method, we associate proteins based on vectors of their evolutionary tree edge (branch or internode) lengths. This approach is based on the expectation that molecular coevolution is greatest between proteins that interact in some way. Mitochondrial DNA-encoded proteins are associated into groups largely consistent with the complexes they come from. This association is apparently not due to the tree structure or mutation processes, leaving coevolution as the best explanation. We show that it is important that the tree used to derive the edge-length vector is estimated accurately in terms of both topology and edge lengths. Although more complex substitution models reduce systematic error, they also inflate stochastic error. This makes the use of less complex substitution models preferable in some circumstances. We describe a method to estimate correlations of pairwise evolutionary distances, which adjusts for non-independent correlations due to shared evolutionary history. Associations of proteins based on their edge-length vectors are visualized and assessed using a variety of hierarchical clustering and multidimensional scaling methods. New formula for estimating the fit of data to model, including the average percent standard deviation of distances on least squares trees, are presented. Use of edge-length vectors is compared and contrasted with correlated distance methods, correlated rates methods, and site-specific evidence of coevolution.

Cluster Analysis↗

Measuring the fit of sequence data to phylogenetic model: allowing for missing data.

It is fundamentally important to assess the fit of data to model in phylogenetic and evolutionary studies. Phylogenetic methods using molecular sequences typically start with a multiple alignment. It is possible to measure the fit of data to model expectations of data, for example, via the likelihood-ratio (G) test or the X(2) test, if all sites in all sequences have an unambiguous residue. However, nearly all alignments of interest contain sites (columns of the alignment) with missing data, that is, ambiguous nucleotides, gaps, or unsequenced regions, which must presently be removed before using the above tests. Unfortunately, this is often either undesirable or impractical, as it will discard much of the data. Here, we show how iterative ML estimators may directly estimate the site-pattern probabilities for columns with missing data, given only standard i.i.d. assumptions. The optimization may use an EM or Newton algorithm, or any other hill-climbing approach. The resulting optimal likelihood under the unconstrained or multinomial model may be compared directly with the likelihood of the data coming from the model (a G statistic). Alternatively the modified observed and the expected frequencies of site patterns may be compared using a X(2) test. The distribution of such statistics is best assessed using appropriate simulations. The new method is applicable to models using codons or paired sites. The methods are also useful with Hadamard conjugations (spectral analysis) and are illustrated with these and with ML evolutionary models that allow site-rate variability.

Amino Acid Sequence↗

Molecular systematics of primary reptilian lineages and the tuatara mitochondrial genome.

We provide phylogenetic analyses for primary Reptilia lineages including, for the first time, Sphenodon punctatus (tuatara) using data from whole mitochondrial genomes. Our analyses firmly support a sister relationship between Sphenodon and Squamata, which includes lizards and snakes. Using Sphenodon as an outgroup for select squamates, we found evidence indicating a sister relationship, among our study taxa, between Serpentes (represented by Dinodon) and Varanidae. Our analyses support monophyly of Archosauria, and a sister relationship between turtles and archosaurs. This latter relationship is congruent with a growing set of morphological and molecular analyses placing turtles within crown Diapsida and recognizing them as secondarily anapsid (lacking a skull fenestration). Inclusion of Sphenodon, as the only surviving member of Sphenodontia (with fossils from the mid-Triassic), helps to fill a sampling gap within previous analyses of reptilian phylogeny. We also report a unique configuration for the mitochondrial genome of Sphenodon, including two tRNA(Lys) copies and an absence of ND5, tRNA(His), and tRNA(Thr) genes.

Animals↗

Evaluating placental inter-ordinal phylogenies with novel sequences including RAG1, gamma-fibrinogen, ND6, and mt-tRNA, plus MCMC-driven nucleotide, amino acid, and codon models.

It is essential to test a priori scientific hypotheses with independent data, not least to partly negate factors such as gene-specific base composition biases misleading our models. Seven new gene segments and sequences plus Bayesian likelihood phylogenetic methods were used to compare and test five recent placental phylogenies. These five phylogenies are similar to each other, yet quite different from Fthose of previously proposed trees, and span Waddell et al. [Syst. Biol. 48 (1999) 1] to Murphy et al. [Science 294 (2001b) 2348]. Trees for RAG1, gamma-fibrinogen, ND6, mt-tRNA, mt-RNA, c-MYC, epsilon -globin, and GHR are significantly congruent with the four main groups of mammals common to the five phylogenies, i.e., Afrotheria, Laurasiatheria, Euarchontoglires, Xenarthra plus Boreoeutheria (Laurasiatheria plus Euarchontoglires). Where these five a priori phylogenies differ, remain areas generally hard to resolve with the new sequences. The root remains ambiguous and does not reject a basal Afrotheria (the Exafroplacentalia hypothesis), Afrotheria plus Xenarthra together with basal (Atlantogenata), or Epitheria (Xenarthra basal) convincingly. Good evidence is found that Eulipotyphla is monophyletic and is located at the base of Laurasiatheria. The shrew mole, Uropsilus, is found to cluster consistently with other moles, while Solenodon may be the sister taxa to all other eulipotyphlans. Support is found for a probable sister pairing of just hedgehogs/gymnures and shrews. Relationships within Afrotheria, except the Paenungulata clade, remain hard to resolve, although there is congruent support for Afroinsectiphillia (aardvark, elephant shrews, golden moles, and tenrecs). A first-time use is made of MCMC enacted general time-reversible (GTR) amino acid and codon-based models for general tree selection. Even with ND6, a GTR amino acid model provided resolution of fine features, such as the sister group relationship of walrus to Otatriidae, and with BRCA a more reasonable rooting. An extensive analysis of GHR sequences reveals strong congruence with prior phylogenies, including strong support for Eulipotyphla, and good resolution within Rodentia. A codon model gives a worse likelihood than a nucleotide model and sometimes switches support, e.g., with RAG1+gamma-fibrinogen from a hyrax-sirenian association to support for Tethytheria. An analysis of the concatenated data is in accordance with well-resolved features of the gene trees. Taken all together, this work suggests that we are on the right path finding strong confirmation of prior phylogenies. However, with the use of robust criteria for assessing trees (i.e., not Bayesian posteriors), it is apparent parts of the tree remain hard to resolve. Since our current models are far from fitting the sequence data, we should continue with our exploratory analyses to arrive at a refined set of hypotheses for future testing using more model independent characters (e.g., rare indels, gene rearrangement, and SINE data).

Animals↗

Pika and vole mitochondrial genomes increase support for both rodent monophyly and glires.

Complete mitochondrial genomes are reported for a pika (Ochotona collaris) and a vole (Volemys kikuchii) then analysed together with 35 other mitochondrial genomes from mammals. With standard phylogenetic methods the pika joins with the other lagomorph (rabbit) and the vole with the other murid rodents (rat and mouse). In addition, with hedgehog excluded, the seven rodent genomes consistently form a homogeneous group in the unrooted placental tree. Except for uncertainty of the position of tree shrew, the clade Glires (monophyletic rodents plus lagomorphs) is consistently found. The unrooted tree obtained by ProtML (Protein Maximum Likelihood, a program in MOLPHY) is compatible with a reclassification of mammals [Syst. Biol. 48, 1-5 (1999)] which is also supported by other recent studies. However, when this tree is rooted with marsupials plus platypus, the outgroup often joins the lineage leading to the three murid rodents, so the rodents are no longer monophyletic. Apart from misplacing the root, the presence of the outgroups also distorts other parts of the unrooted tree. Either constraining the tree to maintain rodents monophyletic, or omitting murids, maintains the ingroup tree and sees the outgroup join on the edge to Xenarthra, to Afrotheria, or to these two groups together. This emphasises the importance of carrying out both an unrooted and a rooted analysis. It is known from cancer research that murid rodents have reduced activity in some DNA repair mechanisms and this alters their substitution pattern - this may be the case for mitochnodrial DNA as well. Comparing nucleotide compositions may identify taxa that differ in aspects of their DNA repair mechanisms.

Animals↗

Very fast algorithms for evaluating the stability of ML and Bayesian phylogenetic trees from sequence data.

Evolutionary trees sit at the core of all realistic models describing a set of related sequences, including alignment, homology search, ancestral protein reconstruction and 2D/3D structural change. It is important to assess the stochastic error when estimating a tree, including models using the most realistic likelihood-based optimizations, yet computation times may be many days or weeks. If so, the bootstrap is computationally prohibitive. Here we show that the extremely fast "resampling of estimated log likelihoods" or RELL method behaves well under more general circumstances than previously examined. RELL approximates the bootstrap (BP) proportions of trees better that some bootstrap methods that rely on fast heuristics to search the tree space. The BIC approximation of the Bayesian posterior probability (BPP) of trees is made more accurate by including an additional term related to the determinant of the information matrix (which may also be obtained as a product of gradient or score vectors). Such estimates are shown to be very close to MCMC chain values. Our analysis of mammalian mitochondrial amino acid sequences suggest that when model breakdown occurs, as it typically does for sequences separated by more than a few million years, the BPP values are far too peaked and the real fluctuations in the likelihood of the data are many times larger than expected. Accordingly, several ways to incorporate the bootstrap and other types of direct resampling with MCMC procedures are outlined. Genes evolve by a process which involves some sites following a tree close to, but not identical with, the species tree. It is seen that under such a likelihood model BP (bootstrap proportions) and BPP estimates may still be reasonable estimates of the species tree. Since many of the methods studied are very fast computationally, there is no reason to ignore stochastic error even with the slowest ML or likelihood based methods.

Algorithms↗