PubMed Health⌕ Search

Biomedical subjects

Barbara R Holland

Publications and source records attributed to Barbara R Holland.

9 recordsLinked to original sources

Heterozygosity and functional allelic variation in the Candida albicans efflux pump genes CDR1 and CDR2.

Elevated expression of the plasma membrane drug efflux pump proteins Cdr1p and Cdr2p was shown to accompany decreased azole susceptibility in Candida albicans clinical isolates. DNA sequence analysis revealed extensive allelic heterozygosity, particularly of CDR2. Cdr2p alleles showed different abilities to transport azoles when individually expressed in Saccharomyces cerevisiae. Loss of heterozygosity, however, did not accompany decreased azole sensitivity in isogenic clinical isolates. Two adjacent non-synonymous single nucleotide polymorphisms (NS-SNPs), G1473A and I1474V in the putative transmembrane (TM) helix 12 of CDR2, were found to be present in six strains including two isogenic pairs. Site-directed mutagenesis showed that the TM-12 NS-SNPs, and principally the G1473A NS-SNP, contributed to functional differences between the proteins encoded by the two Cdr2p alleles in a single strain. Allele-specific PCR revealed that both alleles were equally frequent among 69 clinical isolates and that the majority of isolates (81%) were heterozygous at the G1473A/I1474V locus, a significant (P < 0.001) deviation from the Hardy-Weinberg equilibrium. Phylogenetic analysis by maximum likelihood (Paml) identified 33 codons in CDR2 in which amino acid allelic changes showed a high probability of being selectively advantageous. In contrast, all codons in CDR1 were under purifying selection. Collectively, these results indicate that possession of two functionally different CDR2 alleles in individual strains may confer a selective advantage, but that this is not necessarily due to azole resistance.

ATP-Binding Cassette Transporters↗

Deciphering past human population movements in Oceania: provably optimal trees of 127 mtDNA genomes.

The settlement of the many island groups of Remote Oceania occurred relatively late in prehistory, beginning approximately 3,000 years ago when people sailed eastwards into the Pacific from Near Oceania, where evidence of human settlement dates from as early as 40,000 years ago. Archeological and linguistic analyses have suggested the settlers of Remote Oceania had ancestry in Taiwan, as descendants of a proposed Neolithic expansion that began approximately 5,500 years ago. Other researchers have suggested that the settlers were descendants of peoples from Island Southeast Asia or the existing inhabitants of Near Oceania alone. To explore patterns of maternal descent in Oceania, we have assembled and analyzed a data set of 137 mitochondrial DNA (mtDNA) genomes from Oceania, Australia, Island Southeast Asia, and Taiwan that includes 19 sequences generated for this project. Using the MinMax Squeeze Approach (MMS), we report the consensus network of 165 most parsimonious trees for the Oceanic data set, increasing by many orders of magnitude the numbers of trees for which a provable minimal solution has been found. The new mtDNA sequences highlight the limitations of partial sequencing for assigning sequences to haplogroups and dating recent divergence events. The provably optimal trees found for the entire mtDNA sequences using the MMS method provide a reliable and robust framework for the interpretation of evolutionary relationships and confirm that the female settlers of Remote Oceania descended from both the existing inhabitants of Near Oceania and more recent migrants into the region.

DNA, Mitochondrial↗

Genome BLAST distance phylogenies inferred from whole plastid and whole mitochondrion genome sequences.

BACKGROUND: Phylogenetic methods which do not rely on multiple sequence alignments are important tools in inferring trees directly from completely sequenced genomes. Here, we extend the recently described Genome BLAST Distance Phylogeny (GBDP) strategy to compute phylogenetic trees from all completely sequenced plastid genomes currently available and from a selection of mitochondrial genomes representing the major eukaryotic lineages. BLASTN, TBLASTX, or combinations of both are used to locate high-scoring segment pairs (HSPs) between two sequences from which pairwise similarities and distances are computed in different ways resulting in a total of 96 GBDP variants. The suitability of these distance formulae for phylogeny reconstruction is directly estimated by computing a recently described measure of "treelikeness", the so-called delta value, from the respective distance matrices. Additionally, we compare the trees inferred from these matrices using UPGMA, NJ, BIONJ, FastME, or STC, respectively, with the NCBI taxonomy tree of the taxa under study. RESULTS: Our results indicate that, at this taxonomic level, plastid genomes are much more valuable for inferring phylogenies than are mitochondrial genomes, and that distances based on breakpoints are of little use. Distances based on the proportion of "matched" HSP length to average genome length were best for tree estimation. Additionally we found that using TBLASTX instead of BLASTN and, particularly, combining TBLASTX and BLASTN leads to a small but significant increase in accuracy. Other factors do not significantly affect the phylogenetic outcome. The BIONJ algorithm results in phylogenies most in accordance with the current NCBI taxonomy, with NJ and FastME performing insignificantly worse, and STC performing as well if applied to high quality distance matrices. delta values are found to be a reliable predictor of phylogenetic accuracy. CONCLUSION: Using the most treelike distance matrices, as judged by their delta values, distance methods are able to recover all major plant lineages, and are more in accordance with Apicomplexa organelles being derived from "green" plastids than from plastids of the "red" type. GBDP-like methods can be used to reliably infer phylogenies from different kinds of genomic data. A framework is established to further develop and improve such methods. delta values are a topology-independent tool of general use for the development and assessment of distance methods for phylogenetic inference.

Algorithms↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Improved consensus network techniques for genome-scale phylogeny.

Although recent studies indicate that estimating phylogenies from alignments of concatenated genes greatly reduces the stochastic error, the potential for systematic error still remains, heightening the need for reliable methods to analyze multigene data sets. Consensus methods provide an alternative, more inclusive, approach for analyzing collections of trees arising from multiple genes. We extend a previously described consensus network method for genome-scale phylogeny (Holland, B. R., K. T. Huber, V. Moulton, and P. J. Lockhart. 2004. Using consensus networks to visualize contradictory evidence for species phylogeny. Mol. Biol. Evol. 21:1459-1461) to incorporate additional information. This additional information could come from bootstrap analysis, Bayesian analysis, or various methods to find confidence sets of trees. The new methods can be extended to include edge weights representing genetic distance. We use three data sets to illustrate the approach: 61 genes from 14 angiosperm taxa and one gymnosperm, 106 genes from eight yeast taxa, and 46 members of a gene family from 15 vertebrate taxa.

Animals↗

Untangling long branches: identifying conflicting phylogenetic signals using spectral analysis, neighbor-net, and consensus networks.

Long-branch attraction is a well-known source of systematic error that can mislead phylogenetic methods; it is frequently invoked post hoc, upon recovering a different tree from the one expected based on prior evidence. We demonstrate that methods that do not force the data onto a single tree, such as spectral analysis, Neighbor-Net, and consensus networks, can be used to detect conflicting signals within the data, including those caused by long-branch attraction. We illustrate this approach using a set of taxa from three unambiguously monophyletic families within the Pelecaniformes: the darters, the cormorants and shags, and the gannets and boobies. These three families are universally acknowledged as forming a monophyletic group, but the relationship between the families remains contentious. Using sequence data from three mitochondrial genes (12S, ATPase 6, and ATPase 8) we demonstrate that the relationship between these three families is difficult to resolve because they are separated by a short internal branch and there are conflicting signals due to long-branch attraction, which are confounded with nonhomogeneous sequence evolution across the different genes. Spectral analysis, Neighbor-Net, and consensus networks reveal conflicting signals regarding the placement of one of the darters, with support found for darter monophyly, but also support for a conflicting grouping with the outgroup, pelicans. Furthermore, parsimony and maximum-likelihood analyses produced different trees, with one of the two most parsimonious trees not supporting the monophyly of the darters. Monte Carlo simulations, however, were not sensitive enough to reveal long-branch attraction unless the branches are longer than those actually observed. These results indicate that spectral analysis, Neighbor-Net, and consensus networks offer a powerful approach to detecting and understanding the source of conflicting signals within phylogenetic data.

Animals↗

Using consensus networks to visualize contradictory evidence for species phylogeny.

Building species phylogenies from genome data requires the evaluation of phylogenetic evidence from independent gene loci. We propose an approach to do this using consensus networks. We compare gene trees for eight yeast genomes and show that consensus networks have potential for helping to visualize contradictory evidence for species phylogenies.

Genome, Fungal↗

Optimal alphabets for an RNA world.

Experiments have shown that the canonical AUCG genetic alphabet is not the only possible nucleotide alphabet. In this work we address the question 'is the canonical alphabet optimal?' We make the assumption that the genetic alphabet was determined in the RNA world. Computational tools are used to infer the RNA secondary structure (shape) from a given RNA sequence, and statistics from RNA shapes are gathered with respect to alphabet size. Then, simulations based upon the replication and selection of fixed-sized RNA populations are used to investigate the effect of alternative alphabets upon RNA's ability to step through a fitness landscape. These results show that for a low copy fidelity the canonical alphabet is fitter than two-, six- and eight-letter alphabets. In higher copy-fidelity experiments, six-letter alphabets outperform the four-letter alphabets, suggesting that the canonical alphabet is indeed a relic of the RNA world.

Animals↗

Upper bounds on maximum likelihood for phylogenetic trees.

We introduce a mechanism for analytically deriving upper bounds on the maximum likelihood for genetic sequence data on sets of phylogenies. A simple 'partition' bound is introduced for general models. Tighter bounds are developed for the simplest model of evolution, the two state symmetric model of nucleotide substitution under the molecular clock. This follows earlier theoretical work which has been restricted to this model by analytic complexity. A weakness of current numerical computation is that reported 'maximum likelihood' results cannot be guaranteed, both for a specified tree (because of the possibility of multiple maxima) or over the full tree space (as the computation is intractable for large sets of trees). The bounds we develop here can be used to conclusively eliminate large proportions of tree space in the search for the maximum likelihood tree. This is vital in the development of a branch and bound search strategy for identifying the maximum likelihood tree. We report the results from a simulation study of approximately 10(6) data sets generated on clock-like trees of five leaves. In each trial a likelihood value of one specific instance of a parameterised tree is compared to the bound determined for each of the 105 possible rooted binary trees. The proportion of trees that are eliminated from the search for the maximum likelihood tree ranged from 92% to almost 98%, indicating a computational speed-up factor of between 12 and 44.

Algorithms↗

Sixty alleles of the ALS7 open reading frame in Candida albicans: ALS7 is a hypermutable contingency locus.

The ALS (agglutinin-like sequence) gene family encodes proteins that play a role in adherence of the yeast Candida albicans to endothelial and epithelial cells. The proteins are proposed as virulence factors for this important fungal pathogen of humans. We analyzed 66 C. albicans strains, representing a worldwide collection of 266 infection-causing isolates, and discovered 60 alleles of the ALS7 open reading frame (ORF). Differences between alleles were largely caused by rearrangements of repeat elements in the so-called tandem repeat domain (21 different types occurred) and the VASES region (19 different types). C. albicans is diploid, and combinations of ALS7 alleles generated 49 different genotypes. ALS7 expression was detected in samples isolated directly from five oral candidosis patients. ORFs in the opposite direction contained within the ALS7 ORF were also transcribed in all strains tested. Isolates representing a more pathogenic general-purpose genotype (GPG) cluster of strains tended to have more tandem repeats than other strains. Two types of VASES regions were largely exclusive to GPG strains; the remaining types were largely exclusive to noncluster strains. Our results provide evidence that ALS7 is a hypermutable contingency locus and important for the success of C. albicans as an opportunistic pathogen of humans.

Alleles↗