PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Tree building”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Phylogenetic inference from homologous sequence data: minimum topological assumption, strict mutational compatibility consensus tree as the ultimate solution.

BACKGROUND: For the purposes of phylogenetic inference from molecular data sets many different methods are currently offered as alternatives for researchers in phylogenetic systematics. The vast majority of these methods are based on specific topological assumptions relating to the resultant genealogical tree. Each of these has been shown to perform effectively in special conditions and for specific data sets while yielding less reliable results in other instances. Moreover, the majority of the methods include information from homoplastic characters in spite of a universally accepted agreement in their ineffectiveness for phylogenetic inference, which may often lead to inaccuracy and inconsistency. As an alternative to such methods, a strict mutational compatibility consensus tree building method as a universally applicable and reliable method is reported. RESULTS: The analysis of a data set from a previously published experimental phylogeny demonstrates the accuracy of the strict mutational compatibility consensus tree building method and illustrates its potential for obtaining unambiguous and precise results with full resolution. CONCLUSION: The universal applicability of a simplified compatibility method in its algorithmic form for phylogenetic inference is described. Firstly, dismissal of topological assumptions creates a general potential for agreement of inferred with true phylogeny. Second, exclusion of irregular characters from analysis repeatably enables construction of consistent phylogeny. Third, a direct calculation of bootstrap proportion values for individual nodes of the resulting tree is possible rather than their empirical estimation. Finally, guidance is given for empirical assessment of the sample size necessary for full genealogical resolution and significant bootstrap proportions. REVIEWERS: This article was reviewed by Yuri I. Wolf (nominated by Eugene Koonin), Arcady Mushegian and Martijn Huynen.

Journal Article↗

Effects of nucleotide sequence alignment on phylogeny estimation: a case study of 18S rDNAs of apicomplexa.

The reconstruction of phylogenetic history is predicated on being able to accurately establish hypotheses of character homology, which involves sequence alignment for studies based on molecular sequence data. In an empirical study investigating nucleotide sequence alignment, we inferred phylogenetic trees for 43 species of the Apicomplexa and 3 of Dinozoa based on complete small-subunit rDNA sequences, using six different multiple-alignment procedures: manual alignment based on the secondary structure of the 18S rRNA molecule, and automated similarity-based alignment algorithms using the PileUp, ClustalW, TreeAlign, MALIGN, and SAM computer programs. Trees were constructed using neighboring-joining, weighted-parsimony, and maximum-likelihood methods. All of the multiple sequence alignment procedures yielded the same basic structure for the estimate of the phylogenetic relationship among the taxa, which presumably represents the underlying phylogenetic signal. However, the placement of many of the taxa was sensitive to the alignment procedure used; and the different alignments produced trees that were on average more dissimilar from each other than did the different tree-building methods used. The multiple alignments from the different procedures varied greatly in length, but aligned sequence length was not a good predictor of the similarity of the resulting phylogenetic trees. We also systematically varied the gap weights (the relative cost of inserting a new gap into a sequence or extending an already-existing gap) for the ClustalW program, and this produced alignments that were at least as different from each other as those produced by the different alignment algorithms. Furthermore, there was no combination of gap weights that produced the same tree as that from the structure alignment, in spite of the fact that many of the alignments were similar in length to the structure alignment. We also investigated the phylogenetic information content of the helical and nonhelical regions of the rDNA, and conclude that the helical regions are the most informative. We therefore conclude that many of the literature disagreements concerning the phylogeny of the Apicomplexa are probably based on differences in sequence alignment strategies rather than differences in data or tree-building methods.

Algorithms↗

Prospects for building the tree of life from large sequence databases.

We assess the phylogenetic potential of approximately 300,000 protein sequences sampled from Swiss-Prot and GenBank. Although only a small subset of these data was potentially phylogenetically informative, this subset retained a substantial fraction of the original taxonomic diversity. Sampling biases in the databases necessitate building phylogenetic data sets that have large numbers of missing entries. However, an analysis of two "supermatrices" suggests that even data sets with as much as 92% missing data can provide insights into broad sections of the tree of life.

Animals↗

Phylogenetic relationships within terrestrial mites (Acari: Prostigmata, Parasitengona) inferred from comparative DNA sequence analysis of the mitochondrial cytochrome oxidase subunit I gene.

Partial DNA and amino acid sequences translated from the mitochondrial cytochrome subunit I gene (408 bp) of 17 mite species have been used for analyzing the phylogenetic relationships within the terrestrial Parasitengona (Trombidia). Due to mutational saturation of the third codon position, only first and second codon positions and amino acid sequences were analyzed, applying neighbor-joining, maximum-parsimony, and maximum-likelihood tree-building methods. The reconstructed trees revealed similar topologies of taxa; however, the phylogenetic relationships could be convincingly resolved only within several trombidioid taxa. The proposed basic relationships within the Parasitengona, in particular those of Calyptostomatoidea, Smarididae, and Erythraeidae, were poorly supported in bootstrap tests. A comparison of the presented gene tree with a phylogenetic tree based upon traditional characters revealed only few contradictions in nodes only weakly supported by morphological data. The most astonishing result is the proposed early derivative position of Microtrombidiidae within the terrestrial Parasitengona.

Animals↗

Using models of nucleotide evolution to build phylogenetic trees.

Molecular phylogenetics and its applications are popular and useful tools for making comparative investigations in genetics; however, estimating phylogenetic trees is not always straightforward. Some phylogenetic estimators use an explicit model of nucleotide evolution to estimate evolutionary parameters such as branch lengths and tree topology. There are many models to choose from, and use of the optimal model for a particular data set is important to avoid a loss of power and accuracy in phylogenetic estimations. Here, we review some molecular evolutionary forces and the parameters included in some common models of evolution used to interpret resulting patterns of molecular variation. We present some statistical methods of selecting a particular model of nucleotide evolution, and provide an empirical example of model selection. Statistical model selection strikes a balance between the bias introduced by some models and the increased variance of parameter estimates that results from using other models.

Animals↗

The roots of phylogeny: how did Haeckel build his trees?

Haeckel created much of our current vocabulary in evolutionary biology, such as the term phylogeny, which is currently used to designate trees. Assuming that Haeckel gave the same meaning to this term, one often reproduces Haeckel's trees as the first illustrations of phylogenetic trees. A detailed analysis of Haeckel's own evolutionary vocabulary and theory revealed that Haeckel's trees were genealogical trees and that Haeckel's phylogeny was a morphological concept. However, phylogeny was actually the core of Haeckel's tree reconstruction, and understanding the exact meaning Haeckel gave to phylogeny is crucial to understanding the information Haeckel wanted to convey in his famous trees. Haeckel's phylogeny was a linear series of main morphological stages along the line of descent of a given species. The phylogeny of a single species would provide a trunk around which lateral branches were added as mere ornament; the phylogeny selected for drawing a tree of a given group was considered the most complete line of progress from lower to higher forms of this group, such as the phylogeny of Man for the genealogical tree of Vertebrates. Haeckel's phylogeny was mainly inspired by the idea of the scala naturae, or scale of being. Therefore, Haeckel's genealogical trees, which were only branched on the surface, mainly represented the old idea of scale of being. Even though Haeckel decided to draw genealogical trees after reading On the Origin of Species and was called the German Darwin, he did not draw Darwinian branching diagrams. Although Haeckel always saw Lamarck, Goethe, and Darwin as the three fathers of the theory of evolution, he was mainly influenced by Lamarck and Goethe in his approach to tree reconstruction.

Anatomy, Comparative↗

Building a tree of knowledge: analysis of bitter molecules.

A phylogenetic-like tree of structural fragments has been constructed to extract useful insights from a structural database of bitter molecules. The tree of structural fragments summarizes the substructural groups present in the molecules from the bitter database. These structural fragments are compared with a large number of random molecules to highlight substructures specific to bitter molecules. This organization of the structures enabled the detection of structure-activity relationships for the bitter molecules through the construction of R-tables. Key structural groups, able to distinguish between bitter and random molecules, were identified through an analysis of the tree. This information can be used to further understand which structural components are involved in producing a bitter taste.

Algorithms↗

Evergreen staff: building a tree of knowledge for continuous learning.

Healthcare institutions breathed a collective sigh of relief on January 1, when efforts made for remediation, testing, and contingency planning for the year 2000 finally paid off. Now that the technology has been improved to ensure compatibility, it is important to keep the momentum going to improve efficiency and increase productivity and patient satisfaction levels. One way to incorporate organizational priorities, goals, and strategy at the departmental level is to develop an education plan that stresses mastering fundamental skills. This article explores the components and role of an education plan and identifies the types of efforts that result in the greatest return. It concludes with a case study.

Computer User Training↗

Trio learning: a new strategy for building hybrid neural trees.

Neural trees are constructive algorithms which build decision trees whose nodes are binary neurons. We propose a new learning scheme, "trio-learning," which leads to a significant reduction in the tree complexity. In this strategy, each node of the tree is optimized by taking into account the knowledge that it will be followed by two son nodes. Moreover, trio-learning can be used to build hybrid trees, with internal nodes and terminal nodes of different nature, for solving any standard tasks (e.g. classification, regression, density estimation). Significant results on a handwritten character classification are presented.

Artificial Intelligence↗

Effects of sequence alignment and structural domains of ribosomal DNA on phylogeny reconstruction for the protozoan family sarcocystidae.

Finding correct species relationships using phylogeny reconstruction based on molecular data is dependent on several empirical and technical factors. These include the choice of DNA sequence from which phylogeny is to be inferred, the establishment of character homology within a sequence alignment, and the phylogeny algorithm used. Nevertheless, sequencing and phylogeny tools provide a way of testing certain hypotheses regarding the relationship among the organisms for which phenotypic characters demonstrate conflicting evolutionary information. The protozoan family Sarcocystidae is one such group for which molecular data have been applied phylogenetically to resolve questionable relationships. However, analyses carried out to date, particularly based on small-subunit ribosomal DNA, have not resolved all of the relationships within this family. Analysis of more than one gene is necessary in order to obtain a robust species signal, and some DNA sequences may not be appropriate in terms of their phylogenetic information content. With this in mind, we tested the informativeness of our chosen molecule, the large-subunit ribosomal DNA (lsu rDNA), by using subdivisions of the sequence in phylogenetic analysis through PAUP, fastDNAml, and neighbor joining. The segments of sequence applied correspond to areas of higher nucleotide variation in a secondary-structure alignment involving 21 taxa. We found that subdivision of the entire lsu rDNA is inappropriate for phylogenetic analysis of the Sarcocystidae. There are limited informative nucleotide sites in the lsu rDNA for certain clades, such as the one encompassing the subfamily Toxoplasmatinae. Consequently, the removal of any segment of the alignment compromises the final tree topology. We also tested the effect of using two different alignment procedures (CLUSTAL W and the structure alignment using DCSE) and three different tree-building methods on the final tree topology. This work shows that congruence between different methods in the formation of clades may be a feature of robust topology; however, a sequence alignment based on primary structure may not be comparing homologous nucleotides even though the expected topology is obtained. Our results support previous findings showing the paraphyly of the current genera Sarcocystis and Hammondia and again bring to question the relationships of Sarcocystis muris, Isospora felis, and Neospora caninum. In addition, results based on phylogenetic analysis of the structure alignment suggest that Sarcocystis zamani and Sarcocystis singaporensis, which have reptilian definitive hosts, are monophyletic with Sarcocystis species using mammalian definitive hosts if the genus Frenkelia is synonymized with Sarcocystis.

Animals↗

Radiation and divergence in the Rhagoletis pomonella species group: inferences from allozymes.

The Rhagoletis pomonella species group has for decades been a focal point for debate over the possibility of sympatric speciation via host shift. Here I present the first extensive analysis of genetic (allozyme) divergence in the pomonella group, including all known taxa/populations except the allopatric Mexican population of R. pomonella. The phylogeny is estimated for all four described species (pomonella, mendax, zephyria, and cornivora) plus two undescribed species (the "flowering dogwood fly" and "sparkleberry fly"). Allozyme data for two additional populations of uncertain status (the "plum fly" and "mayhaw fly") are presented for the first time. Two data sets were analyzed, one for 17 loci from 77 populations and one for an additional 12 loci for a subset of 12 of these populations, with more than 4000 flies analyzed in total. Interspecific Nei unbiased genetic distances were generally small, being as low as 0.040. No fixed autapomorphic alleles beyond those already known for R. cornivora and R. zephyria were revealed in the new data, but several loci displaying frequency patterns useful in discriminating the species were discovered. The phylogenetic placement of the flowering dogwood fly differed depending on whether a molecular clock was assumed (UPGMA of Nei distance) or not assumed (frequency parsimony) for tree building. Other than this, however, trees under either assumption were essentially identical. The best tree was used to test the prediction of the sympatric speciation hypothesis that sister taxa should be broadly sympatric. This prediction was not rejected, but the best tree was weakly supported by bootstrap analysis. An unexpected finding was that R. pomonella populations representing ends of its strong latitudinal clines did not cluster together. One possible explanation is that the current R. pomonella is the result of a genetic fusion of two previously isolated, genetically differentiated populations. Such a fusion prior to the origin of the other species in the group could contribute to the poor resolution of the phylogeny.

Animals↗

Resolution of the African hominoid trichotomy by use of a mitochondrial gene sequence.

Mitochondrial DNA sequences encoding the cytochrome oxidase subunit II gene have been determined for five primate species, siamang (Hylobates syndactylus), lowland gorilla (Gorilla gorilla), pygmy chimpanzee (Pan paniscus), crab-eating macaque (Macaca fascicularis), and green monkey (Cercopithecus aethiops), and compared with published sequences of other primate and nonprimate species. Comparisons of cytochrome oxidase subunit II gene sequences provide clear-cut evidence from the mitochondrial genome for the separation of the African ape trichotomy into two evolutionary lineages, one leading to gorillas and the other to humans and chimpanzees. Several different tree-building methods support this same phylogenetic tree topology. The comparisons also yield trees in which a substantial length separates the divergence point of gorillas from that of humans and chimpanzees, suggesting that the lineage most immediately ancestral to humans and chimpanzees may have been in existence for a relatively long time.

Animals↗

Phylogenetic relationships of the genus Frenkelia: a review of its history and new knowledge gained from comparison of large subunit ribosomal ribonucleic acid gene sequences.

The different genera currently classified into the family Sarcocystidae include parasites which are of significant medical, veterinary and economic importance. The genus Sarcocystis is the largest within the family Sarcocystidae and consists of species which infect a broad range of animals including mammals, birds and reptiles. Frenkelia, another genus within this family, consists of parasites that use rodents as intermediate hosts and birds of prey as definitive hosts. Both genera follow an almost identical pattern of life cycle, and their life cycle stages are morphologically very similar. However, the relationship between the two genera remains unresolved because previous analyses of phenotypic characters and of small subunit ribosomal ribonucleic acid gene sequences have questioned the validity of the genus Frenkelia or the monophyly of the genus Sarcocystis if Frenkelia was recognised as a valid genus. We therefore subjected the large subunit ribosomal ribonucleic acid gene sequences of representative taxa in these genera to phylogenetic analyses to ascertain a definitive relationship between the two genera. The full length large subunit ribosomal ribonucleic acid gene sequences obtained were aligned using Clustal W and Dedicated Comparative Sequence Editor secondary structure alignments. The Dedicated Comparative Sequence Editor alignment was then split into two data sets, one including helical regions, and one including non-helical regions, in order to determine the more informative sites. Subsequently, all four alignment data sets were subjected to different tree-building algorithms. All of the analyses produced trees supporting the paraphyly of the genus Sarcocystis if Frenkelia was recognised as a valid genus and, thus, call for a revision of the current definition of these genera. However, an alternative, more parsimonious and more appropriate solution to the Sarcocystis/Frenkelia controversy is to synonymise the genus Frenkelia with the genus Sarcocystis.

Animals↗

POWER: PhylOgenetic WEb Repeater--an integrated and user-optimized framework for biomolecular phylogenetic analysis.

POWER, the PhylOgenetic WEb Repeater, is a web-based service designed to perform user-friendly pipeline phylogenetic analysis. POWER uses an open-source LAMP structure and infers genetic distances and phylogenetic relationships using well-established algorithms (ClustalW and PHYLIP). POWER incorporates a novel tree builder based on the GD library to generate a high-quality tree topology according to the calculated result. POWER accepts either raw sequences in FASTA format or user-uploaded alignment output files. Through a user-friendly web interface, users can sketch a tree effortlessly in multiple steps. After a tree has been generated, users can freely set and modify parameters, select tree building algorithms, refine sequence alignments or edit the tree topology. All the information related to input sequences and the processing history is logged and downloadable for the user's reference. Furthermore, iterative tree construction can be performed by adding sequences to, or removing them from, a previously submitted job. POWER is accessible at http://power.nhri.org.tw.

Algorithms↗

Convergent evolution within the V3 loop domain of human immunodeficiency virus type 1 in association with disease progression.

Phylogenetic analysis was used to study in vivo genetic variation of the V3 region of human immunodeficiency virus type 1 in relation to disease progression in six infants with vertically acquired human immunodeficiency virus type 1 infection. Nucleotide sequences from each infant formed a monophyletic group with similar average branch lengths separating the sets of sequences. In contrast to the star-shaped phylogeny characteristic of interinfant viral evolution, the shape of the phylogeny formed by sequences from the infants who developed AIDS tended to be linear. A computer program, DISTRATE, was written to analyze changes in DNA distance values over time. For the six infants, the rate of divergence from the initial variant was inversely correlated with CD4 cell counts averaged over the first 11 to 15 months of life (r = -0.87, P = 0.024). To uncover evolutionary relationships that might be dictated by protein structure and function, tree-building methods were applied to inferred amino acid sequences. Trees constructed from the full-length protein fragment (92 amino acids) showed that viruses from each infant formed a monophyletic group. Unexpectedly, V3 loop protein sequences (35 amino acids) that were found at later time points from the two infants who developed AIDS clustered together. Furthermore, these sequences uniquely shared amino acids that have been shown to confer a T-cell line tropic phenotype. The evolutionary pattern suggests that viruses from these infants with AIDS acquired similar and possibly more virulent phenotypes.

Acquired Immunodeficiency Syndrome↗

Predicting cesarean delivery with decision tree models.

OBJECTIVE: The purpose of this study was to determine whether decision tree-based methods can be used to predict cesarean delivery. STUDY DESIGN: This was a historical cohort study of women delivered of live-born singleton neonates in 1995 through 1997 (22,157). The frequency of cesarean delivery was 17%; 78 variables were used for analysis. Decision tree rule-based methods and logistic regression models were each applied to the same 50% of the sample to develop the predictive training models and these models were tested on the remaining 50%. RESULTS: Decision tree receiver operating characteristic curve areas were as follows: nulliparous, 0.82; parous, 0.93. Logistic receiver operating characteristic curve areas were as follows: nulliparous, 0.86; parous, 0.93. Decision tree methods and logistic regression methods used similar predictive variables; however, logistic methods required more variables and yielded less intelligible models. Among the 6 decision tree building methods tested, the strict minimum message length criterion yielded decision trees that were small yet accurate. Risk factor variables were identified in 676 nulliparous cesarean deliveries (69%) and 419 parous cesarean deliveries (47.6%). CONCLUSION: Decision tree models can be used to predict cesarean delivery. Models built with strict minimum message length decision trees have the following attributes: Their performance is comparable to that of logistic regression; they are small enough to be intelligible to physicians; they reveal causal dependencies among variables not detected by logistic regression; they can handle missing values more easily than can logistic methods; they predict cesarean deliveries that lack a categorized risk factor variable.

Adolescent↗