PubMed Health⌕ Search

Biomedical subjects

Mark P Simmons

Publications and source records attributed to Mark P Simmons.

13 recordsLinked to original sources

A penalty of using anonymous dominant markers (AFLPs, ISSRs, and RAPDs) for phylogenetic inference.

AFLPs (and to a lesser extent ISSRs and RAPDs) are increasingly being used for phylogenetic inference among closely related species. Presence/absence characters for each AFLP allele treat all absences as homologous to one another. With three or more alleles, terminals are grouped by their shared absence of alleles in character-based phylogenetic-inference methods in a manner that is not redundant with their shared presence of an alternative allele. We conducted simulations to quantify how severe the negative effect of using presence/absence characters of individual bands is for phylogenetic inference relative to standard multistate characters. We examined alternative tree topologies, relative branch lengths, numbers of characters, rates of evolution, and numbers of alternative alleles, using both parsimony and Nei-and-Li distance analyses. Multistate parsimony generally outperformed presence/absence parsimony, which in turn outperformed Nei-and-Li distance. Increasing the character-state space (i.e., the number of alternative character states available) was found to be advantageous for all three methods of analysis examined, but was most advantageous for multistate parsimony. However, the advantage of multistate parsimony relative to Nei-and-Li distance decreased when applied to more divergent characters. More parsimony-informative variation generally alleviated the problem associated with scoring multistate characters as presence/absence characters. The ensemble consistency index was lower for presence/absence characters relative to multistate characters.

Computer Simulation↗

Comprehensive comparative analysis of kinesins in photosynthetic eukaryotes.

BACKGROUND: Kinesins, a superfamily of molecular motors, use microtubules as tracks and transport diverse cellular cargoes. All kinesins contain a highly conserved approximately 350 amino acid motor domain. Previous analysis of the completed genome sequence of one flowering plant (Arabidopsis) has resulted in identification of 61 kinesins. The recent completion of genome sequencing of several photosynthetic and non-photosynthetic eukaryotes that belong to divergent lineages offers a unique opportunity to conduct a comprehensive comparative analysis of kinesins in plant and non-plant systems and infer their evolutionary relationships. RESULTS: We used the kinesin motor domain to identify kinesins in the completed genome sequences of 19 species, including 13 newly sequenced genomes. Among the newly analyzed genomes, six represent photosynthetic eukaryotes. A total of 529 kinesins was used to perform comprehensive analysis of kinesins and to construct gene trees using the Bayesian and parsimony approaches. The previously recognized 14 families of kinesins are resolved as distinct lineages in our inferred gene tree. At least three of the 14 kinesin families are not represented in flowering plants. Chlamydomonas, a green alga that is part of the lineage that includes land plants, has at least nine of the 14 known kinesin families. Seven of ten families present in flowering plants are represented in Chlamydomonas, indicating that these families were retained in both the flowering-plant and green algae lineages. CONCLUSION: The increase in the number of kinesins in flowering plants is due to vast expansion of the Kinesin-14 and Kinesin-7 families. The Kinesin-14 family, which typically contains a C-terminal motor, has many plant kinesins that have the motor domain at the N terminus, in the middle, or the C terminus. Several domains in kinesins are present exclusively either in plant or animal lineages. Addition of novel domains to kinesins in lineage-specific groups contributed to the functional diversification of kinesins. Results from our gene-tree analyses indicate that there was tremendous lineage-specific duplication and diversification of kinesins in eukaryotes. Since the functions of only a few plant kinesins are reported in the literature, this comprehensive comparative analysis will be useful in designing functional studies with photosynthetic eukaryotes.

Algal Proteins↗

How can third codon positions outperform first and second codon positions in phylogenetic inference? An empirical example from the seed plants.

Greater phylogenetic signal is often found in parsimony-based analyses of third codon positions of protein-coding genes relative to their corresponding first and second codon positions, even for early-derived ("basal") clades. We used the Soltis et al. (2000; Bot. J. Linn. Soc. 133:381-461) data matrix of atpB and rbcL from 567 seed plants to quantify how each of six factors (observed character-state space, frequencies of observed character states, substitution probabilities among nucleotides, rate heterogeneity among sites, overall rate of evolution, and number of parsimony-informative characters) contributed to this phenomenon. Each of these six factors was estimated from the original data matrix for parsimony-informative third codon positions considered separately from first and second codon positions combined. One of the most parsimonious trees found was used as the constraint topology; branch lengths were estimated using likelihood-based distances, and characters were simulated on this tree. Differential frequencies of observed character states were found to be the most limiting of the factors simulated for all three codon positions. Differential frequencies of observed character states and differential substitution probabilities among states were relatively advantageous for first and second codon positions. In contrast, differential numbers of observed character states, differential rate heterogeneity among sites, the greater number of parsimony-informative characters, and the higher overall rate of evolution were relatively advantageous for third codon positions. The amount of possible synapomorphy was predictive of the overall success of resolution.

Base Sequence↗

Phylogeny of the Cucurbitales based on DNA sequences of nine loci from three genomes: implications for morphological and sexual system evolution.

The Cucurbitales are a clade of rosids with a worldwide distribution and a striking heterogeneity in species diversity among its seven family members: the Anisophylleaceae (29-40 species), Begoniaceae (1400 spp.), Coriariaceae (15 spp.), Corynocarpaceae (6 spp.), Cucurbitaceae (800 spp.), Datiscaceae (2 spp.), and Tetramelaceae (2 spp.). Most Cucurbitales have unisexual flowers, and species are monoecious, dioecious, andromonoecious, or androdioecious. To resolve interfamilial relationships within the order and to polarize morphological character evolution, especially of flower sexual systems, we sequenced nine plastids (atpB, matK, ndhF, rbcL, the trnL-F region, and the rpl20-rps12 spacer), nuclear (18S and 26S rDNA), and mitochondrial (nad1 b/c intron) genes (together approximately 12,000 bp) of 26 representatives of the seven families plus eight outgroup taxa from six other orders of the Eurosids I. Cucurbitales are strongly supported as monophyletic and are closest to Fagales, albeit with moderate support; both together are sister to Rosales. The deepest split in the Cucurbitales is that between the Anisophylleaceae and the remaining families; next is a clade of Corynocarpaceae and Coriariaceae, followed by Cucurbitaceae, which are sister to a clade of Begoniaceae, Datiscaceae, and Tetramelaceae. Based on this topology, stipulate leaves, inferior ovaries, parietal placentation, and one-seeded fruits are inferred as ancestral in Cucurbitales; exstipulate leaves, superior ovaries, apical placentation, and many-seeded fruits evolved within the order. Bisexual flowers are reconstructed as ancestral, but dioecy appears to have evolved already in the common ancestor of Begoniaceae, Cucurbitaceae, Datiscaceae, and Tetramelaceae, and then to have been lost repeatedly in Begoniaceae and Cucurbitaceae. Both instances of androdioecy (Datisca glomerata and Schizopepon bryoniifolius) evolved from dioecious ancestors, corroborating recent hypotheses about androdioecy often evolving from dioecy.

Begoniaceae↗

Origin and evolution of Kinesin-like calmodulin-binding protein.

Kinesin-like calmodulin-binding protein (KCBP), a member of the Kinesin-14 family, is a C-terminal microtubule motor with three unique domains including a myosin tail homology region 4 (MyTH4), a talin-like domain, and a calmodulin-binding domain (CBD). The MyTH4 and talin-like domains (found in some myosins) are not found in other reported kinesins. A calmodulin-binding kinesin called kinesin-C (SpKinC) isolated from sea urchin (Strongylocentrotus purpuratus) is the only reported kinesin with a CBD. Analysis of the completed genomes of Homo sapiens, Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and a red alga (Cyanidioschyzon merolae 10D) did not reveal the presence of a KCBP. This prompted us to look at the origin of KCBP and its relationship to SpKinC. To address this, we isolated KCBP from a gymnosperm, Picea abies, and a green alga, Stichococcus bacillaris. In addition, database searches resulted in identification of KCBP in another green alga, Chlamydomonas reinhardtii, and several flowering plants. Gene tree analysis revealed that the motor domain of KCBPs belongs to a clade within the Kinesin-14 (C-terminal motors) family. Only land plants and green algae have a kinesin with the MyTH4 and talin-like domains of KCBP. Further, our analysis indicates that KCBP is highly conserved in green algae and land plants. SpKinC from sea urchin, which has the motor domain similar to KCBP and contains a CBD, lacks the MyTH4 and talin-like regions. Our analysis indicates that the KCBPs, SpKinC, and a subset of the kinesin-like proteins are all more closely related to one another than they are to any other kinesins, but that either KCBP gained the MyTH4 and talin-like domains or SpKinC lost them.

Amino Acid Sequence↗

Efficiently resolving the basal clades of a phylogenetic tree using Bayesian and parsimony approaches: a case study using mitogenomic data from 100 higher teleost fishes.

Many phylogenetic analyses that include numerous terminals but few genes show high resolution and branch support for relatively recently diverged clades, but lack of resolution and/or support for "basal" clades of the tree. The various benefits of increased taxon and character sampling have been widely discussed in the literature, albeit primarily based on simulations rather than empirical data. In this study, we used a well-sampled gene-tree analysis (based on 100 mitochondrial genomes of higher teleost fishes) to test empirically the efficiency of different methods of data sampling and phylogenetic inference to "correctly" resolve the basal clades of a tree (based on congruence with the reference tree constructed using all 100 taxa and 7990 characters). By itself, increased character sampling was an inefficient method by which to decrease the likelihood of "incorrect" resolution (i.e., incongruence with the reference tree) for parsimony analyses. Although increased taxon sampling was a powerful approach to alleviate "incorrect" resolution for parsimony analyses, it had the general effect of increasing the number of, and support for, "incorrectly" resolved clades in the Bayesian analyses. For both the parsimony and Bayesian analyses, increased taxon sampling, by itself, was insufficient to help resolve the basal clades, making this sampling strategy ineffective for that purpose. For this empirical study, the most efficient of the six approaches considered to resolve the basal clades when adding nucleotides to a dataset that consists of a single gene sampled for a small, but representative, number of taxa, is to increase character sampling and analyze the characters using the Bayesian method.

Animals↗

Independence of alignment and tree search.

I assert that similarity is the appropriate homology criterion for sequence alignment, as it is with morphology. Methods that select among alignments using parsimony-based tree lengths, as implemented in MALIGN and POY, arrange the data such that they are consistent with a minimum-evolution model. When combining data sets in phylogenetic analyses, we are not trying to reinforce our earlier hypotheses about relationships, but rather to test them. The severity of this test is compromised when congruence with other characters is favored when selecting among alignment parameters.

Algorithms↗

Relative character-state space, amount of potential phylogenetic information, and heterogeneity of nucleotide and amino acid characters.

We examined a broad selection of protein-coding loci from a diverse array of clades and genomes to quantify three factors that determine whether nucleotide or amino acid characters should be preferred for phylogenetic inference. First, we quantified the difference in observed character-state space between nucleotides and amino acids. Second, we quantified the loss of potential phylogenetic signal from silent substitutions when amino acids are used. Third, we used the disparity index to quantify the relative compositional heterogeneity of nucleotides and amino acids and then determined how commonly convergent (rather than unique) shifts in nucleotide and amino acid composition occur in a phylogenetic context. The greater potential phylogenetic signal for nucleotide characters was found to be enormous (on average 440% that of amino acids), whereas the greater observed character-state space for amino acids was less impressive (on average 150.4% that of nucleotides). While matrices of amino acid sequences had less compositional heterogeneity than their corresponding nucleotide sequences, heterogeneity in amino acid composition may be more homoplasious than heterogeneity in nucleotide composition. Given the ability of increased taxon sampling to better utilize the greater potential phylogenetic signal of nucleotide characters and decrease the potential for artifacts caused by heterogeneous nucleotide composition among taxa, we suggest that increased taxon sampling be performed whenever possible instead of restricting analyses to amino acid characters.

Amino Acids↗

How meaningful are Bayesian support values?

In this study, we used an empirical example based on 100 mitochondrial genomes from higher teleost fishes to compare the accuracy of parsimony-based jackknife values with Bayesian support values. Phylogenetic analyses of 366 partitions, using differential taxon and character sampling from the entire data matrix of 100 taxa and 7,990 characters, were performed for both phylogenetic methods. The tree topology and branch-support values from each partition were compared with the tree inferred from all taxa and characters. Using this approach, we quantified the accuracy of the branch-support values assigned by the jackknife and Bayesian methods, with respect to each of 15 basal clades. In comparing the jackknife and Bayesian methods, we found that (1) both measures of support differ significantly from an ideal support index; (2) the jackknife underestimated support values; (3) the Bayesian method consistently overestimated support; (4) the magnitude by which Bayesian values overestimate support exceeds the magnitude by which the jackknife underestimates support; and (5) both methods performed poorly when taxon sampling was increased and character sampling was not increases. These results indicate that (1) the higher Bayesian support values are inappropriate (in magnitude), and (2) Bayesian support values should not be interpreted as probabilities that clades are correctly resolved. We advocate the continued use of the relatively conservative bootstrap and jackknife approaches to estimating branch support rather than the more extreme overestimates provided by the Markov Chain Monte Carlo-based Bayesian methods.

Animals↗

The effects of increasing genetic distance on alignment of, and tree construction from, rDNA internal transcribed spacer sequences.

We examined how alignment of internal transcribed spacers of rDNA in fungi and plants changes with increasing genetic distance by successive removal of sequences from each data set followed by realignment and phylogenetic analysis. Increasing genetic distance can negatively affect phylogenetic reconstruction in two ways. First, it may cause errors in the alignment and therefore the homology hypotheses of the sequence characters. Second, it may cause errors in the homology assessments of character states because of multiple hits on individual branches. These two causes of error in phylogenetic inference were distinguished from one another in our analysis. The errors in alignment caused by increasing genetic distance were primarily due to inserting too few gaps and inserting gaps at the wrong positions. Errors in tree resolution, topology, and/or branch-support values were more often caused by multiple hits than by misaligned positions. This suggests that increasing genetic distance negatively affects our primary homology assessments of character states more severely than our primary homology assessments of characters. We suggest that increasing taxon sampling with the aim of subdividing long branches is a strategy for obtaining reliable alignments.

DNA, Ribosomal Spacer↗

Uninode coding vs gene tree parsimony for phylogenetic reconstruction using duplicate genes.

Two different methods of using paralogous genes for phylogenetic inference have been proposed: reconciled trees (or gene tree parsimony) and uninode coding. Gene tree parsimony suffers from 10 serious problems, including differential weighting of nucleotide and gap characters, undersampling which can be misinterpreted as synapomorphy, all of the characters not being allowed to interact, and conflict between gene trees being given equal weight, regardless of branch support. These problems are largely avoided by using uninode coding. The uninode coding method is elaborated to address multiple gene duplications within a single gene tree family and handle problems caused by lack of gene tree resolution. An example of vertebrate phylogeny inferred from nine genes is reanalyzed using uninode coding. We suggest that uninode coding be used instead of gene tree parsimony for phylogenetic inference from paralogous genes.

Animals↗

Amino acid vs. nucleotide characters: challenging preconceived notions.

The 567-terminal analysis of atpB, rbcL, and 18S rDNA was used as an empirical example to test the use of amino acid vs. nucleotide characters for protein-coding genes at deeper taxonomic levels. Nucleotides for atpB and rbcL had 6.5 times the amount of possible synapomorphy as amino acids. Based on parsimony analyses with unordered character states, nucleotides outperformed amino acids for all three measures of phylogenetic signal used (resolution, branch support, and congruence with independent evidence). The nucleotide tree was much more resolved than the amino acid tree, for both large and small clades. Nearly twice the percentage of well-supported clades resolved in the 18S rDNA tree were resolved using nucleotides (91.8%) relative to amino acids (49.2%). The well-supported clades resolved by both character types were much better supported by nucleotides (98.7% vs. 83.8% average jackknife support). The faster evolving nucleotides with a smaller average character-state space outperformed the slower evolving amino acids with a larger average character-state space. Nucleotides outperformed amino acids even with 90% of the terminals deleted. The lack of resolution on the amino acid trees appears to be caused by a lack of congruence among the amino acids, not a lack of replacement substitutions.

Amino Acids↗

Limitations of relative apparent synapomorphy analysis (RASA) for measuring phylogenetic signal.

In this paper we use hypothetical and empirical data matrices to evaluate the ability of relative apparent synapomorphy analysis (RASA) to measure phylogenetic signal, select outgroups, and identify terminals subject to long-branch attraction. In all cases, except for equal character-state frequencies, RASA indicated extraordinarily high levels of phylogenetic information for hypothetical data matrices that are uninformative regarding relationships among the terminals. Yet, regardless of the number of characters or character-state frequencies, RASA failed to detect phylogenetic signal for hypothetical matrices with strong phylogenetic signal. In our empirical example, RASA indicated increasing phylogenetic signal for matrices for which the strict consensus of the most parsimonious trees is increasingly poorly resolved, clades are increasingly poorly supported, and for which many relationships are in conflict with more widely sampled analyses. RASA is an ineffective approach to identify outgroup terminal(s) with the most plesiomorphic character states for the ingroup. Our hypothetical example demonstrated that RASA preferred outgroup terminals with increasing numbers of convergent character states with ingroup terminals, and rejected the outgroup terminal with all plesiomorphic character states. Our empirical example demonstrated that RASA, in all three cases examined, selected an ingroup terminal, rather than an outgroup terminal, as the best outgroup. In no case was one of the two outgroup terminals even close to being considered the optimal outgroup by RASA. RASA is an ineffective means of identifying problematic long-branch terminals. In our hypothetical example, RASA indicated a terminal as being a problematic long-branch terminal in spite of the terminal being on a zero-length branch and having no possibility of undergoing long-branch attraction with another terminal. RASA also failed to identify actual problematic long-branch terminals that did undergo long-branch attraction, but only after following Lyons-Weiler and Hoelzer's (1997) three-step process to identify and remove terminals subject to long-branch attraction. We conclude that RASA should not be used for any of these purposes.

Algorithms↗