PubMed Health⌕ Search

Biomedical subjects

Amy C Driskell

Publications and source records attributed to Amy C Driskell.

5 recordsLinked to original sources

Supertree bootstrapping methods for assessing phylogenetic variation among genes in genome-scale data sets.

Nonparamtric bootstrapping methods may be useful for assessing confidence in a supertree inference. We examined the performance of two supertree bootstrapping methods on four published data sets that each include sequence data from more than 100 genes. In "input tree bootstrapping," input gene trees are sampled with replacement and then combined in replicate supertree analyses; in "stratified bootstrapping," trees from each gene's separate (conventional) bootstrap tree set are sampled randomly with replacement and then combined. Generally, support values from both supertree bootstrap methods were similar or slightly lower than corresponding bootstrap values from a total evidence, or supermatrix, analysis. Yet, supertree bootstrap support also exceeded supermatrix bootstrap support for a number of clades. There was little overall difference in support scores between the input tree and stratified bootstrapping methods. Results from supertree bootstrapping methods, when compared to results from corresponding supermatrix bootstrapping, may provide insights into patterns of variation among genes in genome-scale data sets.

Algorithms↗

Prospects for building the tree of life from large sequence databases.

We assess the phylogenetic potential of approximately 300,000 protein sequences sampled from Swiss-Prot and GenBank. Although only a small subset of these data was potentially phylogenetically informative, this subset retained a substantial fraction of the original taxonomic diversity. Sampling biases in the databases necessitate building phylogenetic data sets that have large numbers of missing entries. However, an analysis of two "supermatrices" suggests that even data sets with as much as 92% missing data can provide insights into broad sections of the tree of life.

Animals↗

Phylogeny and evolution of the Australo-Papuan honeyeaters (Passeriformes, Meliphagidae).

We analyzed nucleotide variation at four loci for 75 species to produce a phylogenetic hypothesis for the Meliphagidae, and to examine the evolution and biogeographic history of the Meliphagidae. Both maximum parsimony and Bayesian methods of phylogenetic analysis were employed. The family was found to be monophyletic, though the genera Certhionyx, Anthochaera, and Phylidonyris were not. Four major clades were recovered and the spinebills (Acanthorhynchus) formed the sister clade to the remainder of the family in most analyses. The Australian endemic arid-adapted chats (Epthianura, Ashbyia) were found to be nested deeply within the family Meliphagidae. No evidence was found to support the hypothesis of separate New Guinean and Australian endemic radiations, nor of a close phylogenetic relationship between taxa from the New Guinea highlands and those from Australian northern rainforests.

Animals↗

Obtaining maximal concatenated phylogenetic data sets from large sequence databases.

To improve the accuracy of tree reconstruction, phylogeneticists are extracting increasingly large multigene data sets from sequence databases. Determining whether a database contains at least k genes sampled from at least m species is an NP-complete problem. However, the skewed distribution of sequences in these databases permits all such data sets to be obtained in reasonable computing times even for large numbers of sequences. We developed an exact algorithm for obtaining the largest multigene data sets from a collection of sequences. The algorithm was then tested on a set of 100,000 protein sequences of green plants and used to identify the largest multigene ortholog data sets having at least 3 genes and 6 species. The distribution of sizes of these data sets forms a hollow curve, and the largest are surprisingly small, ranging from 62 genes by 6 species, to 3 genes by 65 species, with more symmetrical data sets of around 15 taxa by 15 genes. These upper bounds to sequence concatenation have important implications for building the tree of life from large sequence databases.

Algorithms↗

The challenge of constructing large phylogenetic trees.

The amount of sequence data available to reconstruct the evolutionary history of genes and species has increased 20-fold in the past decade. Consequently the size of phylogenetic analyses has grown as well, and phylogenetic methods, algorithms and their implementations have struggled to keep pace. Computational and other challenges raised by this burgeoning database emerge at several stages of analysis, from the optimal assembly of large data matrices from sequence databases, to the efficient construction of trees from these large matrices and the piece-wise assembly of 'supertrees' from those trees in turn. A final challenge is posed by the difficulty of visualizing and making inferences from trees that might soon routinely contain thousands of species.

Algorithms↗