PubMed HealthSearch

Biomedical subjects

D Sankoff

Publications and source records attributed to D Sankoff.

13 recordsLinked to original sources

Gene order comparisons for phylogenetic inference: evolution of the mitochondrial genome.

Detailed knowledge of gene maps or even complete nucleotide sequences for small genomes leads to the feasibility of evolutionary inference based on the macrostructure of entire genomes, rather than on the traditional comparison of homologous versions of a single gene in different organisms. The mathematical modeling of evolution at the genomic level, however, and the associated inferential apparatus are qualitatively different from the usual sequence comparison theory developed to study evolution at the level of individual gene sequences. We describe the construction of a database of 16 mitochondrial gene orders from fungi and other eukaryotes by using complete or nearly complete genomic sequences; propose a measure of gene order rearrangement based on the minimal set of chromosomal inversions, transpositions, insertions, and deletions necessary to convert the order in one genome to that of the other; report on algorithm design and the development of the DERANGE software for the calculation of this measure; and present the results of analyzing the mitochondrial data with the aid of this tool.

Biological Evolution

Efficient optimal decomposition of a sequence into disjoint regions, each matched to some template in an inventory.

Given an amino acid sequence, we discuss how to find efficiently an optimal set of disjoint regions (substrings, domains, modules, etc.), each of which can be matched to some element of a predefined inventory containing, for example, consensus sequences, protosequences, or protein family profiles. A two-stage approach to sequence decomposition, consisting of the detection of all acceptable matches followed by the construction of an optimal subset of compatible matches, leads to computational difficulties. When the problem is reformulated in terms of network comparisons, it can be solved in time quadratic in the length of the sequence and linear with the number of templates in the inventory, by a single pass of a dynamic programming algorithm. This method has the advantage that the criterion for acceptable matches can be relaxed without materially affecting computing time. Except under special conditions it is more efficient than previous segmentation methods based on dynamic programming.

Algorithms

Designer invariants for large phylogenies.

The Cavender-Felsenstein edge-length invariants for binary characters on 4-trees provide the starting point for the development of "customized" invariants for evaluating and comparing phylogenetic hypotheses. The binary character invariants may be generalized to k-valued characters without losing the quadratic nature of the invariants as functions of the theoretical frequencies f(UVXY) of observable character configurations (U at organism 1, V at 2, etc.). The key to the approach is that certain sets of these configurations constitute events which are probabilistically independent from other such sets, under the symmetric Markov change models studied. By introducing more complex sets of configurations, we find the quadratic invariants for 5-trees in the binary model and for individual edges in 6-trees or, indeed, in any size tree. The same technique allows us to formulate invariants for entire trees, but these are cubic functions for 6-trees and are higher-degree polynomials for larger trees. With k-valued characters and, especially, with large trees, the types of configuration sets (events) used in the simpler examples are too rare (i.e., their predicted frequencies are too low) to be useful, and the construction of meaningful pairs of independent events becomes an important and nontrivial task in designing invariants suited to testing specific hypotheses. In a very natural way, this approach fits in with well-known statistical methodology for contingency tables. We explore use of events such as "only transitions occur for character i (i.e., position i in a nucleic acid sequence) in subtree a" in analyzing a set of data on ribosomal RNA in the context of the controversy over the origins of archaebacteria, eubacteria, and eukaryotes.

Archaea

Probabilistic models of genome shuffling.

The comparison of entire genomes in evolutionary studies gives rise to alignments characterized by many intersections, or inversions in the order of two fragments in different genomes. To model this, we suggest a random migration process for fragments, and discuss its equilibrium distribution in the case of linear and circular genomes. Simulations are carried out to explore "cut-off" behavior as the process approaches equilibrium. We define a new process to take into account the indistinguishability of two fragments which are adjacent in both genomes being compared. Questions of applicability of these models are discussed.

Biological Evolution

A continuous analog for RNA folding.

A linear segment in which a number of pairs of intervals of equal length are identified as potential stems is the subject of a folding problem analogous to inference of RNA secondary structure. A quantity of free energy (or equivalently, energy per unit length) is associated with each stem, and the various types of loops are assigned energy costs as a function of their lengths. Inference of stable structures can then be carried out in the same way as in RNA folding. More important, perturbation of stem lengths and energy densities (modelling various mutational processes affecting nucleotide sequences) allows the delineation of domains of stability of various foldings, through the explicit calculation of their boundaries, in a low-dimensional parameter space.

Mathematics

Archetypical features in tRNA families.

A compilation of known tRNA, and tRNA gene sequences from archaebacteria, eubacteria, and eukaryotes permits the construction of tRNA cloverleafs which show conserved structural elements for each tRNA family. Positions conserved across the three kingdoms are thought to represent archetypical features of tRNAs which preceded the divergence of these kingdoms.

Archaea

Evolution of methionine initiator and phenylalanine transfer RNAs.

Sequence data from methionine initiator and phenylalanine transfer RNAs were used to construct phylogenetic trees by the maximum parsimony method. Although eukaryotes, prokaryotes and chloroplasts appear related to a common ancestor, no firm conclusion can be drawn at this time about mitochondrial-coded transfer RNAs. tRNA evolution is not appropriately described by random hit models, since the various regions of the molecule differ sharply in their mutational fixation rates. "Hot" mutational spots are identified in the Tpsic, the amino acceptor and the upper anticodon stems; the D arm and the loop areas on the other hand are highly conserved. Crucial tertiary interactions are thus essentially preserved while most of the double helical domain undergoes base pair interchange. Transitions are about half as costly as transversions, suggesting that base pair interchanges proceed mostly through G-U and A-C intermediates. There is a preponderance of replacements starting from G and C but this bias appears to follow the high G + C content of the easily mutated base paired regions.

Animals

Bacteriophage MS2 RNA: a correlation between the stability of the codon: anticodon interaction and the choice of code words.

The non-random distribution of degenerate code words in Bacteriophage MS2 RNA can be explained partially by considerations of the stability of the codon-anticodon complex in prokaryotic systems. Supporting this hypothesis we note that wobble codons are positively selected in codons having G and/or C in the first two positions. In contrast, wobble codons are statistically less likely in codons composed of A and U in the first two positions. Analyses of nucleotides adjacent to 5' and 3' ends of codons indicate a nonrandom distribution as well. It is thus likely that some elements of RNA evolution are independent of the structural needs of the RNA itself and of the translated protein product.

Anticodon

Frequency of insertion-deletion, transversion, and transition in the evolution of 5S ribosomal RNA.

The problem of choosing an alignment of two or more nucleotide sequences is particularly difficult for nucleic acids, such as 5S ribosomal RNA, which do not code for protein and for which secondary structure is unknown. Given a set of 'costs' for the various types of replacement mutations and for base insertion or deletion, we present a dynamic programming algorithm which finds the optimal (least costly) alignment for a set of N sequences simultaneously, where each sequence is associated with one of the N tips of a given evolutionary tree. Concurrently, protosequences are constructed corresponding to the ancestral nodes of the tree. A version of this algorithm, modified to be computationally feasible, is implemented to align the sequences of 5S RNA from nine organisms. Complete sets of alignments and protosequence reconstructions are done for a large number of different configurations of mutation costs. Examination of the family of curbes of total replacements inferred versus the ratio of transitions/transversions inferred, each curve corresponding to a given number of insertions-deletions inferred, provides a method for estimating relative costs and relative frequencies for these different types of mutations.

Base Sequence

The evolutionary relationships among known life forms.

Sequences of small subunit (SSU) and large subunit (LSU) ribosomal RNA genes from archaebacteria, eubacteria, and the nucleus, chloroplasts, and mitochondria of eukaryotes have been compared in order to identify the most conservative positions. Aligned sets of these positions for both SSU and LSU rRNA have been used to generate tree diagrams relating the source organisms/organelles. Branching patterns were evaluated using the statistical bootstrapping technique. The resulting SSU and LSU trees are remarkably congruent and show a high degree of similarity with those based on alternative data sets and/or generated by different techniques. In addition to providing insights into the evolution of prokaryotic and eukaryotic (nuclear) lineages, the analysis reported here provides, for the first time, an extensive phylogeny of the mitochondrial lineage.

Base Sequence