PubMed Health⌕ Search

Biomedical subjects

Andrew J Roger

Publications and source records attributed to Andrew J Roger.

At least 19 recordsLinked to original sources

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

Invasion and persistence of a selfish gene in the Cnidaria.

BACKGROUND: Homing endonuclease genes (HEGs) are superfluous, but are capable of invading populations that mix alleles by biasing their inheritance patterns through gene conversion. One model suggests that their long-term persistence is achieved through recurrent invasion. This circumvents evolutionary degeneration, but requires reasonable rates of transfer between species to maintain purifying selection. Although HEGs are found in a variety of microbes, we found the previous discovery of this type of selfish genetic element in the mitochondria of a sea anemone surprising. METHODS/PRINCIPAL FINDINGS: We surveyed 29 species of Cnidaria for the presence of the COXI HEG. Statistical analyses provided evidence for HEG invasion. We also found that 96 individuals of Metridium senile, from five different locations in the UK, had identical HEG sequences. This lack of sequence divergence illustrates the stable nature of Anthozoan mitochondria. Our data suggests this HEG conforms to the recurrent invasion model of evolution. CONCLUSIONS: Ordinarily such low rates of HEG transfer would likely be insufficient to enable major invasion. However, the slow rate of Anthozoan mitochondrial change lengthens greatly the time to HEG degeneration: this significantly extends the periodicity of the HEG life-cycle. We suggest that a combination of very low substitution rates and rare transfers facilitated metazoan HEG invasion.

Animals↗

Using confidence set heuristics during topology search improves the robustness of phylogenetic inference.

We examine the impact of likelihood surface characteristics on phylogenetic inference. Amino acid data sets simulated from topologies with branch length features chosen to represent varying degrees of difficulty for likelihood maximization are analyzed. We present situations where the tree found to achieve the global maximum in likelihood is often not equal to the true tree. We use the program covSEARCH to demonstrate how the use of adaptively sized pools of candidate trees that are updated using confidence tests results in solution sets that are highly likely to contain the true tree. This approach requires more computation than traditional maximum likelihood methods, hence covSEARCH is best suited to small to medium-sized alignments or large alignments with some constrained nodes. The majority rule consensus tree computed from the confidence sets also proves to be different from the generating topology. Although low phylogenetic signal in the input alignment can result in large confidence sets of trees, some biological information can still be obtained based on nodes that exhibit high support within the confidence set. Two real data examples are analyzed: mammal mitochondrial proteins and a small tubulin alignment. We conclude that the technique of confidence set optimization can significantly improve the robustness of phylogenetic inference at a reasonable computational cost. Additionally, when either very short internal branches or very long terminal branches are present, confident resolution of specific bipartitions or subtrees, rather than whole-tree phylogenies, may be the most realistic goal for phylogenetic methods.

Algorithms↗

The glycolytic pathway of Trimastix pyriformis is an evolutionary mosaic.

BACKGROUND: Glycolysis and subsequent fermentation is the main energy source for many anaerobic organisms. The glycolytic pathway consists of ten enzymatic steps which appear to be universal amongst eukaryotes. However, it has been shown that the origins of these enzymes in specific eukaryote lineages can differ, and sometimes involve lateral gene transfer events. We have conducted an expressed sequence tag (EST) survey of the anaerobic flagellate Trimastix pyriformis to investigate the nature of the evolutionary origins of the glycolytic enzymes in this relatively unstudied organism. RESULTS: We have found genes in the Trimastix EST data that encode enzymes potentially catalyzing nine of the ten steps of the glycolytic conversion of glucose to pyruvate. Furthermore, we have found two different enzymes that in principle could catalyze the conversion of phosphoenol pyruvate (PEP) to pyruvate (or the reverse reaction) as part of the last step in glycolysis. Our phylogenetic analyses of all of these enzymes revealed at least four cases where the relationship of the Trimastix genes to homologs from other species is at odds with accepted organismal relationships. Although lateral gene transfer events likely account for these anomalies, with the data at hand we were not able to establish with confidence the bacterial donor lineage that gave rise to the respective Trimastix enzymes. CONCLUSION: A number of the glycolytic enzymes of Trimastix have been transferred laterally from bacteria instead of being inherited from the last common eukaryotic ancestor. Thus, despite widespread conservation of the glycolytic biochemical pathway across eukaryote diversity, in a number of protist lineages the enzymatic components of the pathway have been replaced by lateral gene transfer from disparate evolutionary sources. It remains unclear if these replacements result from selectively advantageous properties of the introduced enzymes or if they are neutral outcomes of a gene transfer 'ratchet' from food or endosymbiotic organisms or a combination of both processes.

Amino Acid Sequence↗

Testing for covarion-like evolution in protein sequences.

The covarion hypothesis of molecular evolution proposes that selective pressures on an amino acid or nucleotide site change through time, thus causing changes of evolutionary rate along the edges of a phylogenetic tree. Several kinds of Markov models for the covarion process have been proposed. One model, proposed by Huelsenbeck (2002), has 2 substitution rate classes: the substitution process at a site can switch between a single variable rate, drawn from a discrete gamma distribution, and a zero invariable rate. A second model, suggested by Galtier (2001), assumes rate switches among an arbitrary number of rate classes but switching to and from the invariable rate class is not allowed. The latter model allows for some sites that do not participate in the rate-switching process. Here we propose a general covarion model that combines features of both models, allowing evolutionary rates not only to switch between variable and invariable classes but also to switch among different rates when they are in a variable state. We have implemented all 3 covarion models in a maximum likelihood framework for amino acid sequences and tested them on 23 protein data sets. We found significant likelihood increases for all data sets for the 3 models, compared with a model that does not allow site-specific rate switches along the tree. Furthermore, we found that the general model fit the data better than the simpler covarion models in the majority of the cases, highlighting the complexity in modeling the covarion process. The general covarion model can be used for comparing tree topologies, molecular dating studies, and the investigation of protein adaptation.

Algorithms↗

The origin and diversification of eukaryotes: problems with molecular phylogenetics and molecular clock estimation.

Determining the relationships among and divergence times for the major eukaryotic lineages remains one of the most important and controversial outstanding problems in evolutionary biology. The sequencing and phylogenetic analyses of ribosomal RNA (rRNA) genes led to the first nearly comprehensive phylogenies of eukaryotes in the late 1980s, and supported a view where cellular complexity was acquired during the divergence of extant unicellular eukaryote lineages. More recently, however, refinements in analytical methods coupled with the availability of many additional genes for phylogenetic analysis showed that much of the deep structure of early rRNA trees was artefactual. Recent phylogenetic analyses of a multiple genes and the discovery of important molecular and ultrastructural phylogenetic characters have resolved eukaryotic diversity into six major hypothetical groups. Yet relationships among these groups remain poorly understood because of saturation of sequence changes on the billion-year time-scale, possible rapid radiations of major lineages, phylogenetic artefacts and endosymbiotic or lateral gene transfer among eukaryotes. Estimating the divergence dates between the major eukaryote lineages using molecular analyses is even more difficult than phylogenetic estimation. Error in such analyses comes from a myriad of sources including: (i) calibration fossil dates, (ii) the assumed phylogenetic tree, (iii) the nucleotide or amino acid substitution model, (iv) substitution number (branch length) estimates, (v) the model of how rates of evolution change over the tree, (vi) error inherent in the time estimates for a given model and (vii) how multiple gene data are treated. By reanalysing datasets from recently published molecular clock studies, we show that when errors from these various sources are properly accounted for, the confidence intervals on inferred dates can be very large. Furthermore, estimated dates of divergence vary hugely depending on the methods used and their assumptions. Accurate dating of divergence times among the major eukaryote lineages will require a robust tree of eukaryotes, a much richer Proterozoic fossil record of microbial eukaryotes assignable to extant groups for calibration, more sophisticated relaxed molecular clock methods and many more genes sampled from the full diversity of microbial eukaryotes.

Eukaryotic Cells↗

Phylogenetic estimation under codon models can be biased by codon usage heterogeneity.

In theory, codon models that account for the dependence of nucleotide substitutions between codon positions as well as differences between synonymous and non-synonymous changes best describe the sequence evolution in protein coding genes. However, in practice we know little about the degree to which violations of the assumptions of codon model-based estimates occur, and how significant these artifacts may be. In nucleotide-based phylogenies from first and second codon positions in a concatenated plastid gene data set, two distantly related taxa--dinoflagellate and haptophyte plastids--were robustly grouped together. This artifactual grouping is attributed to the parallel heterogeneity in leucine (Leu) and serine (Ser) codon usages in the data set. Here, by using this data set, we demonstrated that codon-based phylogenetic estimations are seriously biased, robustly uniting the dinoflagellate and haptophyte plastids into a monophyletic clade, when the model assumption of homogeneity of codon composition was violated. Our results suggest that similar phylogenetic artifacts may occur via codon usage heterogeneity in any amino acids in codon model-based estimations. We advise that homogeneity in codon usage across taxa in a data set be confirmed before codon model-based phylogenetic estimation is attempted.

Codon↗

Evolution of four gene families with patchy phylogenetic distributions: influx of genes into protist genomes.

BACKGROUND: Lateral gene transfer (LGT) in eukaryotes from non-organellar sources is a controversial subject in need of further study. Here we present gene distribution and phylogenetic analyses of the genes encoding the hybrid-cluster protein, A-type flavoprotein, glucosamine-6-phosphate isomerase, and alcohol dehydrogenase E. These four genes have a limited distribution among sequenced prokaryotic and eukaryotic genomes and were previously implicated in gene transfer events affecting eukaryotes. If our previous contention that these genes were introduced by LGT independently into the diplomonad and Entamoeba lineages were true, we expect that the number of putative transfers and the phylogenetic signal supporting LGT should be stable or increase, rather than decrease, when novel eukaryotic and prokaryotic homologs are added to the analyses. RESULTS: The addition of homologs from phagotrophic protists, including several Entamoeba species, the pelobiont Mastigamoeba balamuthi, and the parabasalid Trichomonas vaginalis, and a large quantity of sequences from genome projects resulted in an apparent increase in the number of putative transfer events affecting all three domains of life. Some of the eukaryotic transfers affect a wide range of protists, such as three divergent lineages of Amoebozoa, represented by Entamoeba, Mastigamoeba, and Dictyostelium, while other transfers only affect a limited diversity, for example only the Entamoeba lineage. These observations are consistent with a model where these genes have been introduced into protist genomes independently from various sources over a long evolutionary time. CONCLUSION: Phylogenetic analyses of the updated datasets using more sophisticated phylogenetic methods, in combination with the gene distribution analyses, strengthened, rather than weakened, the support for LGT as an important mechanism affecting the evolution of these gene families. Thus, gene transfer seems to be an on-going evolutionary mechanism by which genes are spread between unrelated lineages of all three domains of life, further indicating the importance of LGT from non-organellar sources into eukaryotic genomes.

Alcohol Dehydrogenase↗

Recombination between elongation factor 1alpha genes from distantly related archaeal lineages.

Homologous recombination (HR) and lateral gene transfer are major processes in genome evolution. The combination of the two processes, HR between genes in different species, has been documented but is thought to be restricted to very similar sequences in relatively closely related organisms. Here we report two cases of interspecific HR in the gene encoding the core translational protein translation elongation factor 1alpha (EF-1alpha) between distantly related archaeal groups. Maximum-likelihood sliding window analyses indicate that a fragment of the EF-1alpha gene from the archaeal lineage represented by Methanopyrus kandleri was recombined into the orthologous gene in a common ancestor of the Thermococcales. A second recombination event appears to have occurred between the EF-1alpha gene of the genus Methanothermobacter and its ortholog in a common ancestor of the Methanosarcinales, a distantly related euryarchaeal lineage. These findings suggest that HR occurs across a much larger evolutionary distance than generally accepted and affects highly conserved essential "informational" genes. Although difficult to detect by standard whole-gene phylogenetic analyses, interspecific HR in highly conserved genes may occur at an appreciable frequency, potentially confounding deep phylogenetic inference and hypothesis testing.

Archaea↗

On the correlation between genomic G+C content and optimal growth temperature in prokaryotes: data quality and confounding factors.

The correlation between genomic G+C content and optimal growth temperature in prokaryotes has gained renewed interest after Musto et al. [H. Musto, H. Naya, A. Zavala, H. Romero, F. Alvarex-Valin, G. Bernardi, Correlations between genomic GC levels and optimal growth temperatures in prokaryotes, FEBS Lett. 573 (2004) 73-77], reported that positive correlations exist in 15 families studied. We have reanalyzed their data and found that when genome size and data quality were adjusted for, there was no significant evidence of relationship between optimal temperature and GC content for two of the families that had previously shown strongly significant correlations. Using updated temperature optima for Halobacteriaceae species we found the correlation is insignificant in this family. For the family Enterobacteriaceae when genome size and optimal temperature are included in a multiple linear regression, only genome size is significant as a predictor of GC content. We showed that more profound statistical methods than simple two factor correlation analysis should be used for analyzing complex intrinsic and extrinsic factors that affect genomic GC content. We further found that a positive correlation between temperature and genomic GC is only evident in free-living species of low optimal growth temperatures.

Base Composition↗

Comprehensive multigene phylogenies of excavate protists reveal the evolutionary positions of "primitive" eukaryotes.

Many of the protists thought to represent the deepest branches on the eukaryotic tree are assigned to a loose assemblage called the "excavates." This includes the mitochondrion-lacking diplomonads and parabasalids (e.g., Giardia and Trichomonas) and the jakobids (e.g., Reclinomonas). We report the first multigene phylogenetic analyses to include a comprehensive sampling of excavate groups (six nuclear-encoded protein-coding genes, nine of the 10 recognized excavate groups). Excavates coalesce into three clades with relatively strong maximum likelihood bootstrap support. Only the phylogenetic position of Malawimonas is uncertain. Diplomonads, parabasalids, and the free-living amitochondriate protist Carpediemonas are closely related to each other. Two other amitochondriate excavates, oxymonads and Trimastix, form the second monophyletic group. The third group is comprised of Euglenozoa (e.g., trypanosomes), Heterolobosea, and jakobids. Unexpectedly, jakobids appear to be specifically related to Heterolobosea. This tree topology calls into question the concept of Discicristata as a supergroup of eukaryotes united by discoidal mitochondrial cristae and makes it implausible that jakobids represent an independent early-diverging eukaryotic lineage. The close jakobids-Heterolobosea-Euglenozoa connection demands complex evolutionary scenarios to explain the transition between the presumed ancestral bacterial-type mitochondrial RNA polymerase found in jakobids and the phage-type protein in other eukaryotic lineages, including Euglenozoa and Heterolobosea.

Animals↗

The tree of eukaryotes.

Recent advances in resolving the tree of eukaryotes are converging on a model composed of a few large hypothetical 'supergroups', each comprising a diversity of primarily microbial eukaryotes (protists, or protozoa and algae). The process of resolving the tree involves the synthesis of many kinds of data, including single-gene trees, multigene analyses, and other kinds of molecular and structural characters. Here, we review the recent progress in assembling the tree of eukaryotes, describing the major evidence for each supergroup, and where gaps in our knowledge remain. We also consider other factors emerging from phylogenetic analyses and comparative genomics, in particular lateral gene transfer, and whether such factors confound our understanding of the eukaryotic tree.

Journal Article↗

Biases in phylogenetic estimation can be caused by random sequence segments.

We consider the effects of fully or partially random sequences on the estimation of four-taxon phylogenies. Fully or partially random sequences occur when whole subsets of sequences or some sites for subsets of sequences are independent of sequence data for the other taxa. Random sequences can be a consequence of misalignment or because sites evolve at very fast rates in some portions of a tree, a situation that occurs especially in analyses involving deep divergence times. One might reasonably speculate that random sites will only add noise to the estimation of a phylogeny. We show that in the case that a random sequence is added to a three-taxa alignment, it is more likely to be a neighbor of the sequence corresponding to the longest branch in the three-taxon tree. Surprisingly, when only about half of the sites show randomness, a long-branch-repels form of small sample bias occurs, and when a minority of sites show randomness this becomes a long-branch-attraction bias again. The most serious bias, one that does not vanish with increasing sequence length, occurs when more than one sequence is partially random. If there is a large amount of overlap in the random sites for two sequences, those two sequences will be attracted to each other; otherwise, they will repel each other. Random sequences or sites can, therefore, cause complicated biases in phylogenetic inference. We suggest performing analyses with and without potentially saturated sequences and/or misaligned sites, to check that these biases are not affecting the inferred branching pattern.

Animals↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗

Likelihood, parsimony, and heterogeneous evolution.

Evolutionary rates vary among sites and across the phylogenetic tree (heterotachy). A recent analysis suggested that parsimony can be better than standard likelihood at recovering the true tree given heterotachy. The authors recommended that results from parsimony, which they consider to be nonparametric, be reported alongside likelihood results. They also proposed a mixture model, which was inconsistent but better than either parsimony or standard likelihood under heterotachy. We show that their main conclusion is limited to a special case for the type of model they study. Their mixture model was inconsistent because it was incorrectly implemented. A useful nonparametric model should perform well over a wide range of possible evolutionary models, but parsimony does not have this property. Likelihood-based methods are therefore the best way to deal with heterotachy.

Animals↗

Gene transfers from nanoarchaeota to an ancestor of diplomonads and parabasalids.

Rare evolutionary events, such as lateral gene transfers and gene fusions, may be useful to pinpoint, and correlate the timing of, key branches across the tree of life. For example, the shared possession of a transferred gene indicates a phylogenetic relationship among organismal lineages by virtue of their shared common ancestral recipient. Here, we present phylogenetic analyses of prolyl-tRNA and alanyl-tRNA synthetase genes that indicate lateral gene transfer events to an ancestor of the diplomonads and parabasalids from lineages more closely related to the newly discovered archaeal hyperthermophile Nanoarchaeum equitans (Nanoarchaeota) than to Crenarchaeota or Euryarchaeota. The support for this scenario is strong from all applied phylogenetic methods for the alanyl-tRNA sequences, whereas the phylogenetic analyses of the prolyl-tRNA sequences show some disagreements between methods, indicating that the donor lineage cannot be identified with a high degree of certainty. However, in both trees, the diplomonads and parabasalids branch together within the Archaea, strongly suggesting that these two groups of unicellular eukaryotes, often regarded as the two earliest independent offshoots of the eukaryotic lineage, share a common ancestor to the exclusion of the eukaryotic root. Unfortunately, the phylogenetic analyses of these two aminoacyl-tRNA synthetase genes are inconclusive regarding the position of the diplomonad/parabasalid group within the eukaryotes. Our results also show that the lineage leading to Nanoarchaeota branched off from Euryarchaeota and Crenarchaeota before the divergence of diplomonads and parabasalids, that this unexplored archaeal diversity, currently only represented by the hyperthermophilic organism Nanoarchaeum equitans, may include members living in close proximity to mesophilic eukaryotes, and that the presence of split genes in the Nanoarchaeum genome is a derived feature.

Alanine-tRNA Ligase↗