PubMed Health⌕ Search

Biomedical subjects

Tancred Frickey

Publications and source records attributed to Tancred Frickey.

14 recordsLinked to original sources

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho = 0), the beta-binomial distribution (Rho = 0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5 K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus↗

Mclip: motif detection based on cliques of gapped local profile-to-profile alignments.

UNLABELLED: A multitude of motif-finding tools have been published, which can generally be assigned to one of three classes: expectation-maximization, Gibbs-sampling or enumeration. Irrespective of this grouping, most motif detection tools only take into account similarities across ungapped sequence regions, possibly causing short motifs located peripherally and in varying distance to a 'core' motif to be missed. We present a new method, adding to the set of expectation-maximization approaches, that permits the use of gapped alignments for motif elucidation. AVAILABILITY: The program is available for download from: http://bioinfoserver.rsbs.anu.edu.au/downloads/mclip.jar. SUPPLEMENTARY INFORMATION: http://bioinfoserver.rsbs.anu.edu.au/utils/mclip/info.php.

Algorithms↗

Rab14 is part of the early endosomal clathrin-coated TGN microdomain.

Rab14 localizes to the Golgi/TGN and to early endosomes, but its biological function remains unclear. By structural modeling, we identified Rab14-specific residues and established a close relationship between the Rab2/Rab4/Rab14, Rab11/25 and Rab39 sub-groups within the Rab protein family. By quantitative confocal microscopy and by density centrifugation we show that Rab14 is part of the early endosomal AP-1 microdomain. Overexpression of a dominant-negative Rab14 GTP-binding mutant that solely localizes to the Golgi donor compartment accelerated EGF degradation. We suggest that the AP-1 microdomain represents the interconnecting compartment in which Rab14 vesicles cycle between early endosomes and the Golgi cisternae.

Clathrin↗

Classification of AAA+ proteins.

AAA+ proteins form a large superfamily of P-loop ATPases involved in the energy-dependent unfolding and disaggregation of macromolecules. In a clustering study aimed at defining the AAA proteins within this superfamily, we generated a map of AAA+ proteins based on sequence similarity, which suggested higher-order groups. A classification based primarily on morphological characteristics, which was proposed at the same time, differed from the cluster map in several aspects, such as the position of RuvB-like helicases and the inclusion of divergent clades, such as viral SF3 helicases and plant disease resistance proteins (RFL1). Here, we establish the presence of an alpha-helical domain C-terminal to the ATPase domain (the C-domain) as characteristic for AAA+ proteins and re-evaluate all clades proposed to belong to this superfamily, based on this characteristic. We find that RFL1 and its homologs (APAF-1, CED-4, MalT, and AfsR) are AAA+ proteins and SF3 helicases are not. We also present a new and more comprehensive cluster map, which assigns a central position to RuvB and clarifies the relationships between the clades of the AAA+ superfamily.

Amino Acid Sequence↗

AbrB-like transcription factors assume a swapped hairpin fold that is evolutionarily related to double-psi beta barrels.

AbrB is a key transition-state regulator of Bacillus subtilis. Based on the conservation of a betaalphabeta structural unit, we proposed a beta barrel fold for its DNA binding domain, similar to, but topologically distinct from, double-psi beta barrels. However, the NMR structure revealed a novel fold, the "looped-hinge helix." To understand this discrepancy, we undertook a bioinformatics study of AbrB and its homologs; these form a large superfamily, which includes SpoVT, PrlF, MraZ, addiction module antidotes (PemI, MazE), plasmid maintenance proteins (VagC, VapB), and archaeal PhoU homologs. MazE and MraZ form swapped-hairpin beta barrels. We therefore reexamined the fold of AbrB by NMR spectroscopy and found that it also forms a swapped-hairpin barrel. The conservation of the core betaalphabeta element supports a common evolutionary origin for swapped-hairpin and double-psi barrels, which we group into a higher-order class, the cradle-loop barrels, based on the peculiar shape of their ligand binding site.

Amino Acid Sequence↗

WIPI-1alpha (WIPI49), a member of the novel 7-bladed WIPI protein family, is aberrantly expressed in human cancer and is linked to starvation-induced autophagy.

WD-repeat proteins are regulatory beta-propeller platforms that enable the assembly of multiprotein complexes. Here, we report the functional and bioinformatic analysis of human WD-repeat protein Interacting with PhosphoInosides (WIPI)-1alpha (WIPI49/Atg18), a member of a novel WD-repeat protein family with autophagic capacity in Saccharomyces cerevisiae and Caenorhabditis elegans, recently identified as phospholipid-binding effectors. Our phylogenetic analysis divides the WIPI protein family into two paralogous groups that fold into 7-bladed beta-propellers. Structural modeling identified two evolutionary conserved interaction sites in WIPI propellers, one of which may bind phospholipids. Human WIPI-1alpha has LXXLL signature motifs for nuclear receptor interactions and binds androgen and estrogen receptors in vitro. Strikingly, human WIPI genes were found aberrantly expressed in a variety of matched tumor tissues including kidney, pancreatic and skin cancer. We found that endogenous hWIPI-1 protein colocalizes in part with the autophagosomal marker LC3 at punctate cytoplasmic structures in human melanoma cells. In addition, hWIPI-1 accumulated in large vesicular and cup-shaped structures in the cytoplasm when autophagy was induced by amino-acid deprivation. These cytoplasmic formations were blocked by wortmannin, a classic inhibitor of PI-3 kinase-mediated autophagy. Our data suggest that WIPI proteins share an evolutionary conserved function in autophagy and that autophagic capacity may be compromised in human cancers.

Amino Acid Sequence↗

PhyloGenie: automated phylome generation and analysis.

Phylogenetic reconstruction is the method of choice to determine the homologous relationships between sequences. Difficulties in producing high-quality alignments, which are the basis of good trees, and in automating the analysis of trees have unfortunately limited the use of phylogenetic reconstruction methods to individual genes or gene families. Due to the large number of sequences involved, phylogenetic analyses of proteomes preclude manual steps and therefore require a high degree of automation in sequence selection, alignment, phylogenetic inference and analysis of the resulting set of trees. We present a set of programs that automates the steps from seed sequence to phylogeny and a utility to extract all phylogenies that match specific topological constraints from a database of trees. Two example applications that show the type of questions that can be answered by phylome analysis are provided. The generation and analysis of the Thermoplasma acidophilum phylome with regard to lateral gene transfer between Thermoplasmata and Sulfolobus, showed best BLAST hits to be far less reliable indicators of lateral transfer than the corresponding protein phylogenies. The generation and analysis of the Danio rerio phylome provided more than twice as many proteins as described previously, supporting the hypothesis of an additional round of genome duplication in the actinopterygian lineage.

Amino Acid Sequence↗

CLANS: a Java application for visualizing protein families based on pairwise similarity.

SUMMARY: The main source of hypotheses on the structure and function of new proteins is their homology to proteins with known properties. Homologous relationships are typically established through sequence similarity searches, multiple alignments and phylogenetic reconstruction. In cases where the number of potential relationships is large, for example in P-loop NTPases with many thousands of members, alignments and phylogenies become computationally demanding, accumulate errors and lose resolution. In search of a better way to analyze relationships in large sequence datasets we have developed a Java application, CLANS (CLuster ANalysis of Sequences), which uses a version of the Fruchterman-Reingold graph layout algorithm to visualize pairwise sequence similarities in either two-dimensional or three-dimensional space. AVAILABILITY: CLANS can be downloaded at http://protevo.eb.tuebingen.mpg.de/download.

Algorithms↗

Thermoplasma acidophilum TAA43 is an archaeal member of the eukaryotic meiotic branch of AAA ATPases.

Sequencing of the Thermoplasma acidophilum genome revealed a new gene, taa43 , which codes for a 43-kDa protein containing one AAA domain; we therefore termed it Thermoplasma AAA ATPase of 43 kDa (TAA43). Close homologs of TAA43 are found only in related Thermoplasmales, e.g. T. volcanium and Ferroplasma acidarmanus , but not in other Archaea. A detailed phylogenetic analysis showed that TAA43 and its homologs belong to the 'meiotic' branch of the AAA family. Although AAA proteins usually assemble into high-molecular-weight complexes, native TAA43 is predominantly dimeric except for a minor fraction eluting in the void volume of a sizing column. Wild-type and mutant TAA43 proteins were overexpressed in Escherichia coli , purified as dimers and characterized functionally. Since the canonical proteasome activating nucleotidase is not present in Thermoplasmales, TAA43 was tested for stimulation of proteasome activity, which was, however, not detected. Interestingly, immunoprecipitation analysis with TAA43 specific antibodies found a fraction of native TAA43 associated with Thermoplasma ribosomal proteins.

Adenosine Triphosphatases↗

Genome duplication, a trait shared by 22000 species of ray-finned fish.

Through phylogeny reconstruction we identified 49 genes with a single copy in man, mouse, and chicken, one or two copies in the tetraploid frog Xenopus laevis, and two copies in zebrafish (Danio rerio). For 22 of these genes, both zebrafish duplicates had orthologs in the pufferfish (Takifugu rubripes). For another 20 of these genes, we found only one pufferfish ortholog but in each case it was more closely related to one of the zebrafish duplicates than to the other. Forty-three pairs of duplicated genes map to 24 of the 25 zebrafish linkage groups but they are not randomly distributed; we identified 10 duplicated regions of the zebrafish genome that each contain between two and five sets of paralogous genes. These phylogeny and synteny data suggest that the common ancestor of zebrafish and pufferfish, a fish that gave rise to approximately 22000 species, experienced a large-scale gene or complete genome duplication event and that the pufferfish has lost many duplicates that the zebrafish has retained.

Animals↗

Dealing with saturation at the amino acid level: a case study based on anciently duplicated zebrafish genes.

The ray-finned fishes (Actinopterygii) seem to have two copies of many tetrapod (Sarcopterygii) genes. The origin of these duplicate fish genes is the subject of some controversy. One explanation for the existence of these extra fish genes could be an increase in the rate of independent gene duplications in fishes. Alternatively, gene duplicates in fish may have been formed in the ancestor of all or most Actinopterygii during a complete genome duplication event. A third possibility is that tetrapods have lost more genes than fish after gene or genome duplication events in the common ancestor of both lineages. These three hypotheses can be tested by phylogenetic reconstruction. Previously, we found that a large number of anciently duplicated genes of zebrafish are sister sequences in evolutionary trees suggesting that they were produced in Actinopterygii after the divergence of Sarcopterygii [Phil. Trans. R. Soc. Lond. B 356 (2001) 119]. On the other hand, several well-supported trees showed one of the two fish genes as the sister sequence to a monophyletic clade that included the second fish gene and genes from frog, chicken, mouse and human. These so-called outgroup topologies suggest that the origin of many fish duplicates predates the divergence of the Sarcopterygii and Actinopterygii and support the hypothesis that tetrapods have lost duplicates that have been retained in fish. Here we show that many of these 'outgroup' tree topologies are erroneous and can be corrected when mutational saturation is taken into account. To this end, a Java-based application has been developed to visualize the amount of saturation in amino acid sequences. The program graphically displays the number of observed frequent and rare amino acid replacements between pairs of sequences against their overall evolutionary distance. Discrimination between frequent and rare amino acid replacements is based on substitution probability matrices (e.g. PAM and BLOSUM). Evolutionary distances between sequences can be computed from the fraction of unsaturated sites only and evolutionary trees inferred by pairwise distance methods. When trees are computed by omitting the saturated fraction of sites, most fish duplicates are sister sequences.

Amino Acids↗

Phylogenetic analyses suggest lateral gene transfer from the mitochondrion to the apicoplast.

Apicomplexan protozoa contain a single mitochondrion and a multimembranous plastid-like organelle termed apicoplast. The size of the apicomplexan plastid genome is extremely small (35 kb) thus offering a limited number of genes for phylogenetic analysis. Moreover, the sequences of apicoplast genes are highly adenosine+thymidine-rich and rapidly evolving. Due to these facts, phylogenetic analyses based on different genes or the structure of the ribosomal operon show conflicting results and the evolutionary history of this exciting organelle remains unclear. Although it is evident that the apicoplast and its genome is plastid-derived, our detailed phylogenetic analysis of amino acid and nucleotide sequences of selected apicoplast ribosomal protein genes rpl2, rpl14 and rps12 show their possible mitochondrial origin. The affinity of apicoplast ribosomal proteins to their mitochondrial homologs is very stable and well supported. Based on our results we propose that apicoplasts might contain both plastid and mitochondrial genes, thus constituting a hybrid assembly.

Animals↗

Molecular phylogeny and historical biogeography of the Aphanius (Pisces, Cyprinodontiformes) species complex of central Anatolia, Turkey.

Phylogenetic relationships of a subset of Aphanius fish comprising central Anatolia, Turkey, are investigated to test the hypothesis of geographic speciation driven by early Pliocene orogenic events in spite of morphological similarity. We use 3434 aligned base pairs of mitochondrial DNA from 42 samples representing 36 populations of three species and six outgroup species to test this hypothesis. Genes analyzed include those encoding the 12S and 16S ribosomal RNAs; transfer RNAs coding for valine, leucine, isoleucine, glutamine, methionine, tryptophan, alanine, asparagine, cysteine, and tyrosine; and complete NADH dehydrogenase subunits I and II. Distance based minimum evolution and maximum-likelihood analyses identify six well-supported clades consisting of Aphanius danfordii, Aphanius sp. aff danfordii, and four clades of Aphanius anatoliae. Parsimony analysis results in 462 equally parsimonious trees, all of which contain the six well supported clades identified in the other analyses. Our phylogenetic results are supported by hybridization studies (Villwock, 1964), and by the geological history of Anatolia. Phylogenetic relationships among the six clades are only weakly supported, however, and differ among analytical methods. We therefore test and subsequently reject the hypothesis of simultaneous diversification among the six central Anatolian clades. However, our analyses do not identify any internodes that are significantly better supported than expected by chance alone. Therefore, although bifurcating branching order is hypothesized to underlie this radiation, the exact branching order is difficult to estimate with confidence.

Animals↗

Phylogenetic analysis of AAA proteins.

AAA ATPases form a large protein family with manifold cellular roles. They belong to the AAA+ superfamily of ringshaped P-loop NTPases, which exert their activity through the energy-dependent unfolding of macromolecules. Phylogenetic analyses have suggested the existence of five major clades of AAA domains (proteasome subunits, metalloproteases, domains D1 and D2 of ATPases with two AAA domains, and the MSP1/katanin/spastin group), as well as a number of deeply branching minor clades. These analyses however have been characterized by a lack of consistency in defining the boundaries of the AAA family. We have used cluster analysis to delineate unambiguously the group of AAA sequences within the AAA+ superfamily. Phylogenetic and cluster analysis of this sequence set revealed the existence of a sixth major AAA clade, comprising the mitochondrial, membrane-bound protein BCS1 and its homologues. In addition, we identified several deep branches consisting mainly of hypothetical proteins resulting from genomic projects. Analysis of the AAA N-domains provided direct support for the obtained phylogeny for most branches, but revealed some deep splits that had not been apparent from phylogenetic analysis and some unexpected similarities between distant clades. It also revealed highly degenerate D1 domains in plant MSP1 sequences and in at least one deeply branching group of hypothetical proteins (YC46), showing that AAA proteins with two ATPase domains arose at least three times independently.

Adenosine Triphosphatases↗