PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Tree building”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Evolutionary origins of the endocannabinoid system.

Endocannabinoid system evolution was estimated by searching for functional orthologs in the genomes of twelve phylogenetically diverse organisms: Homo sapiens, Mus musculus, Takifugu rubripes, Ciona intestinalis, Caenorhabditis elegans, Drosophila melanogaster, Saccharomyces cerevisiae, Arabidopsis thaliana, Plasmodium falciparum, Tetrahymena thermophila, Archaeoglobus fulgidus, and Mycobacterium tuberculosis. Sequences similar to human endocannabinoid exon sequences were derived from filtered BLAST searches, and subjected to phylogenetic testing with ClustalX and tree building programs. Monophyletic clades that agreed with broader phylogenetic evidence (i.e., gene trees displaying topographical congruence with species trees) were considered orthologs. The capacity of orthologs to function as endocannabinoid proteins was predicted with pattern profilers (Pfam, Prosite, TMHMM, and pSORT), and by examining queried sequences for amino acid motifs known to serve critical roles in endocannabinoid protein function (obtained from a database of site-directed mutagenesis studies). This novel transfer of functional information onto gene trees enabled us to better predict the functional origins of the endocannabinoid system. Within this limited number of twelve organisms, the endocannabinoid genes exhibited heterogeneous evolutionary trajectories, with functional orthologs limited to mammals (TRPV1 and GPR55), or vertebrates (CB2 and DAGLbeta), or chordates (MAGL and COX2), or animals (DAGLalpha and CB1-like receptors), or opisthokonta (animals and fungi, NAPE-PLD), or eukaryotes (FAAH). Our methods identified fewer orthologs than did automated annotation systems, such as HomoloGene. Phylogenetic profiles, nonorthologous gene displacement, functional convergence, and coevolution are discussed.

Animals↗

The unholy trinity: taxonomy, species delimitation and DNA barcoding.

Recent excitement over the development of an initiative to generate DNA sequences for all named species on the planet has in our opinion generated two major areas of contention as to how this 'DNA barcoding' initiative should proceed. It is critical that these two issues are clarified and resolved, before the use of DNA as a tool for taxonomy and species delimitation can be universalized. The first issue concerns how DNA data are to be used in the context of this initiative; this is the DNA barcode reader problem (or barcoder problem). Currently, many of the published studies under this initiative have used tree building methods and more precisely distance approaches to the construction of the trees that are used to place certain DNA sequences into a taxonomic context. The second problem involves the reaction of the taxonomic community to the directives of the 'DNA barcoding' initiative. This issue is extremely important in that the classical taxonomic approach and the DNA approach will need to be reconciled in order for the 'DNA barcoding' initiative to proceed with any kind of community acceptance. In fact, we feel that DNA barcoding is a misnomer. Our preference is for the title of the London meetings--Barcoding Life. In this paper we discuss these two concerns generated around the DNA barcoding initiative and attempt to present a phylogenetic systematic framework for an improved barcoder as well as a taxonomic framework for interweaving classical taxonomy with the goals of 'DNA barcoding'.

Animals↗

Accurate prediction of kidney allograft outcome based on creatinine course in the first 6 months posttransplant.

Most attempts to predict early kidney allograft loss are based on the patient and donor characteristics at baseline. We investigated how the early posttransplant creatinine course compares to baseline information in the prediction of kidney graft failure within the first 4 years after transplantation. Two approaches to create a prediction rule for early graft failure were evaluated. First, the whole data set was analysed using a decision-tree building software. The software, rpart, builds classification or regression models; the resulting models can be represented as binary trees. In the second approach, a Hill-Climbing algorithm was applied to define cut-off values for the median creatinine level and creatinine slope in the period between day 60 and 180 after transplantation. Of the 497 patients available for analysis, 52 (10.5%) experienced an early graft loss (graft loss within the first 4 years after transplantation). From the rpart algorithm, a single decision criterion emerged: Median creatinine value on days 60 to 180 higher than 3.1 mg/dL predicts early graft failure (accuracy 95.2% but sensitivity = 42.3%). In contrast, the Hill-Climbing algorithm delivered a cut-off of 1.8 mg/dL for the median creatinine level and a cut-off of 0.3 mg/dL per month for the creatinine slope (sensitivity = 69.5% and specificity 79.0%). Prediction rules based on median and slope of creatinine levels in the first half year after transplantation allow early identification of patients who are at risk of loosing their graft early after transplantation. These patients may benefit from therapeutic measures tailored for this high-risk setting.

Creatinine↗

Topology testing of phylogenies using least squares methods.

BACKGROUND: The least squares (LS) method for constructing confidence sets of trees is closely related to LS tree building methods, in which the goodness of fit of the distances measured on the tree (patristic distances) to the observed distances between taxa is the criterion used for selecting the best topology. The generalized LS (GLS) method for topology testing is often frustrated by the computational difficulties in calculating the covariance matrix and its inverse, which in practice requires approximations. The weighted LS (WLS) allows for a more efficient albeit approximate calculation of the test statistic by ignoring the covariances between the distances. RESULTS: The goal of this paper is to assess the applicability of the LS approach for constructing confidence sets of trees. We show that the approximations inherent to the WLS method did not affect negatively the accuracy and reliability of the test both in the analysis of biological sequences and DNA-DNA hybridization data (for which character-based testing methods cannot be used). On the other hand, we report several problems for the GLS method, at least for the available implementation. For many data sets of biological sequences, the GLS statistic could not be calculated. For some data sets for which it could, the GLS method included all the possible trees in the confidence set despite a strong phylogenetic signal in the data. Finally, contrary to WLS, for simulated sequences GLS showed undercoverage (frequent non-inclusion of the true tree in the confidence set). CONCLUSION: The WLS method provides a computationally efficient approximation to the GLS useful especially in exploratory analyses of confidence sets of trees, when assessing the phylogenetic signal in the data, and when other methods are not available.

Animals↗

Relaxed neighbor joining: a fast distance-based phylogenetic tree construction method.

Our ability to construct very large phylogenetic trees is becoming more important as vast amounts of sequence data are becoming readily available. Neighbor joining (NJ) is a widely used distance-based phylogenetic tree construction method that has historically been considered fast, but it is prohibitively slow for building trees from increasingly large datasets. We developed a fast variant of NJ called relaxed neighbor joining (RNJ) and performed experiments to measure the speed improvement over NJ. Since repeated runs of the RNJ algorithm generate a superset of the trees that repeated NJ runs generate, we also assessed tree quality. RNJ is dramatically faster than NJ, and the quality of resulting trees is very similar for the two algorithms. The results indicate that RNJ is a reasonable alternative to NJ and that it is especially well suited for uses that involve large numbers of taxa or highly repetitive procedures such as bootstrapping.

Algorithms↗

A parsimonay analysis of eukaryotic small subunit ribosomal RNA ("18 S") sequences.

Using the principles of Felsenstein's DNA parsimony analysis program (package PHYLIP 3.2), a program with the capacity to handle data sets of the magnitude of ss rRNA has been developed. With this program we applied an heuristic approach to 83 s rRNA sequences from the Antwerp data bank, compared along 4230 positions of their alignment. Branch and bound can also be applied to a smaller number of sequences. We suggest a heuristic tree for a large number of species might be more interesting than a proven minimal tree for a smaller set of species. Comparisons are also made with earlier studies using distance matrix methods; caution is urged for deciding on the evolutionary position of widely divergent groups, whatever the tree building technique used.

Base Sequence↗

Small subunit rDNA phylogeny of Bacillidium sp. (Microspora, Mrazekiidae) infecting oligochaets.

Small subunit (SSU) rDNA has been sequenced from a microsporidium, identified as a member of the genus Bacillidium obtained from an oligochaete. The length of the amplified PCR product was 1386 bp which is currently the longest microsporidium SSU sequence known. Phylogenetic analysis using 28 microsporidia SSU sequences, using 3 different tree-building methods indicated that Bacillidium sp. may be one of the earliest branches on the microsporidia tree. However, bootstrapping failed to give a high score (more than 50%) for the position of Bacillidium sp. The branch leading to Bacillidium sp. was long, indicating that this species is not closely related to any of the other microsporidia so far studied by means of rDNA.

Animals↗

Evolutionary relationships of avian Eimeria species among other Apicomplexan protozoa: monophyly of the apicomplexa is supported.

Direct, reverse transcriptase-mediated, partial sequencing of the small-subunit (16S-like) ribosomal RNA (srRNA) of Eimeria tenella and E. acervulina was performed. Sequences were aligned by eye with six previously published, partial or complete srRNA sequences of apicomplexan protists (Plasmodium berghei, Theileria annulata, Cryptosporidium sp., Toxoplasma gondii, Sarcocystis muris, and S. gigantea). Six eukaryotic protists (a slime mold, a yeast, two dinoflagellates, and two ciliates) acted as an outgroup for a parsimony-based phylogenetic analysis (PAUP Ver. 3.0). The 188 phylogenetically informative sites (i.e., those positions that neither were unvaried nor had only autapomorphic substitutions) supported a single tree topology 481 steps in length with a consistency index of 0.65 in which the monophyly of the Apicomplexa was supported. The two Eimeria species and S. muris, S. gigantea, and T. gondii formed a pair of monophyletic groups that were sister groups. The two Sarcocystis species were not hypothesized to be sister taxa. The genera Plasmodium and Cryptosporidium were hypothesized to form the sister group to these five coccidia and T. annulata. A priori data-editing techniques that deleted "variable" positions prior to analysis failed to recognize the monophyly of the Apicomplexa when the same parsimony-based tree-building algorithm was used. Inability of the outgroup taxa to root the well-supported ingroup tree (Apicomplexa) at a unique site when these taxa were used individually for this purpose reinforces the need for an appropriate, multiple-taxon outgroup in such analyses.

Animals↗

Total evidence, consensus, and bat phylogeny: A distance-based approach.

Resolution of the total evidence (i.e., character congruence) versus consensus (i.e., taxonomic congruence) debate has been impeded by (1) a failure to employ validation methods consistently across both tree-building and consensus analyses, (2) the incomparability of methods for constructing as opposed to those for combining trees, and (3) indifference to aspects of trees other than their topologies. We demonstrate a uniform, distance-based approach which allows for comparability among the results of character- and taxonomic-congruence studies, whether or not an identical suite of taxa has been included in all contributing data sets. Our results indicate that total-evidence and consensus trees differ little in topology if branch lengths are taken into account when combining two or more trees. In addition, when character-state data are converted to distances, our method permits their combination with information produced by techniques which generate distances directly. Moreover, treating all data sets or trees as distance matrices avoids the problem that different numbers of characters in contributing studies may confound the conclusions of a total-evidence or consensus analysis. Our protocol is illustrated with an example involving bats, in which the three component studies based on serology, DNA hybridization, and anatomy imply distinct phylogenies. However, the total-evidence and consensus trees support a fourth, somewhat different, topology resolved at all but one node and which conforms closely to the currently accepted higher category classification of Chiroptera.

Animals↗

Direct calculation of a tree length using a distance matrix.

Comparative studies of tree-building methods have shown minimum evolution to be in general an accurate criterion for selecting a true tree. To improve the use of this criterion, this paper proposes a method for rapidly and directly calculating a length of a dichotomous tree without having to resort to branch length calculations. This direct calculation (DC) method applies to the complete final topology, giving equal importance to each branch after a dichotomy. According to this method, the tree length S(DC) is S(DC) = sigma(i) sigma(j)(D(ij)/2(B(ij))) = (sigma(i<j) sigma D(ij)2(Bmax-B(ij)))/2(Bmax)(-1) where D(ij) is the observed distance between taxa i and j, B(ij) is the number of branches connecting i and j, Bmax is the greatest B(ij) in the tree, and the powers of two are due to the dichotomy of the tree. This tree length expression may be used as a rapid method for selecting the shortest tree from a set of hypothetical or subobtimal trees.

Algorithms↗

A Markov Chain Model of Coalescence with Recombination

Trees that describe the ancestry of DNA sequences sampled from a population may differ between loci because of genetic recombination. We seek to understand the relationship between such trees for loci that are linked with non-zero recombination rate. We consider a coalescent process model with recombination, as described by Hudson (1983; 1990). For two loci and a sample size of two sequences, a detailed analysis of this process yields the joint distribution of the two trees (one at each locus). A number of interesting results follow from this analysis, including the distribution of the number of recombination events in the history of the sample. For the general case of m loci and samples of size n, we describe an algorithm for simulating the tree building process. Because analytic results are difficult to obtain in this case, we use simulation to study properties of trees at multiple linked loci such as total tree time and number of recombination events. Copyright 1997 Academic Press

Journal Article↗

Phylogenetic relationships among Tetrahymena species determined using the polymerase chain reaction.

The species of the Tetrahymena pyriformis complex present a conundrum with regard to their highly conservative morphology and widely divergent molecular characteristics. We have investigated the phylogenetic relationships among these species using the nucleotide sequences from the histone H3II/H4II region of the genome. This region includes portions of the two histone coding sequences, as well as the intergenic region. The DNA sequences of these regions were amplified by the polymerase chain reaction (PCR) and the sequence of each was determined. Nucleotide substitutions and insertions/deletions within this set of sequences were compared to determine the phylogenetic relationships among the species of the complex. These data yield phylogenetic trees with identical topologies when different tree-building routines are used, indicating that the data are very robust. Glaucoma chattoni was used as an outgroup to root the trees for this analysis. The genome organization of G. chattoni and the divergence of its histone H3II/H4II region sequence relative to those of the complex clearly indicate that this species has diverged considerably from the complex. These results show that PCR amplification analysis is feasible over considerable evolutionary distances. However, DNA-DNA hybridization may be more useful than sequence analysis in resolving the relationships among the closely related species in the complex.

Amino Acid Sequence↗

Constructing a minimal diagnostic decision tree.

Classification trees and discriminant function analysis were employed in order to ascertain whether a small number of diagnostic decision rules could be extracted from a large inventory of items. Several models, involving up to 17 symptoms, that led to a broad psychiatric diagnosis were then tested on a small validation sample of 53 patients. All methods, with the exception of CART used without any pruning, generated identical trees involving four items. Almost 90% of the validation sample was able to be correctly classified by all methods although poor classification performance was noted in the case of one particular diagnosis, Schizoaffective Psychosis. In contrast, stepwise linear discriminant analysis originally selected 17 items, although three out of the first four items selected were identical to those chosen by the tree-building methods. Although more research is required, there are indications that the latter methods may be usefully employed in constructing parsimonious decision trees.

Algorithms↗

Three-dimensional segmentation and skeletonization to build an airway tree data structure for small animals.

Quantitative analysis of intrathoracic airway tree geometry is important for objective evaluation of bronchial tree structure and function. Currently, there is more human data than small animal data on airway morphometry. In this study, we implemented a semi-automatic approach to quantitatively describe airway tree geometry by using high-resolution computed tomography (CT) images to build a tree data structure for small animals such as rats and mice. Silicon lung casts of the excised lungs from a canine and a mouse were used for micro-CT imaging of the airway trees. The programming language IDL was used to implement a 3D region-growing threshold algorithm for segmenting out the airway lung volume from the CT data. Subsequently, a fully-parallel 3D thinning algorithm was implemented in order to complete the skeletonization of the segmented airways. A tree data structure was then created and saved by parsing through the skeletonized volume using the Python programming language. Pertinent information such as the length of all airway segments was stored in the data structure. This approach was shown to be accurate and efficient for up to six generations for the canine lung cast and ten generations for the mouse lung cast.

Algorithms↗

SEAVIEW and PHYLO_WIN: two graphic tools for sequence alignment and molecular phylogeny.

SEAVIEW and PHYLO_WIN are two graphic tools for X Windows-Unix computers dedicated to sequence alignment and molecular phylogenetics. SEAVIEW is a sequence alignment editor allowing manual or automatic alignment through an interface with CLUSTALW program. Alignment of large sequences with extensive length differences is made easier by a dot-plot-based routine. The PHYLO_WIN program allows phylogenetic tree building according to most usual methods (neighbor joining with numerous distance estimates, maximum parsimony, maximum likelihood), and a bootstrap analysis with any of them. Reconstructed trees can be drawn, edited, printed, stored, evaluated according to numerous criteria. Taxonomic species groups and sets of conserved regions can be defined by mouse and stored into sequence files, thus avoiding multiple data files. Both tools are entirely mouse driven. On-line help makes them easy to use. They are freely available by anonymous ftp at biom3.univ-lyon1.fr/pub/ mol_phylogeny or http:@acnuc.univ-lyon1.fr/, or by e-mail to galtier@biomserv.univ-lyon1.fr.

Animals↗

Phylogenetic supermatrix analysis of GenBank sequences from 2228 papilionoid legumes.

A comprehensive phylogeny of papilionoid legumes was inferred from sequences of 2228 taxa in GenBank release 147. A semiautomated analysis pipeline was constructed to download, parse, assemble, align, combine, and build trees from a pool of 11,881 sequences. Initial steps included all-against-all BLAST similarity searches coupled with assembly, using a novel strategy for building length-homogeneous primary sequence clusters. This was followed by a combination of global and local alignment protocols to build larger secondary clusters of locally aligned sequences, thus taking into account the dramatic differences in length of the heterogeneous coding and noncoding sequence data present in GenBank. Next, clusters were checked for the presence of duplicate genes and other potentially misleading sequences and examined for combinability with other clusters on the basis of taxon overlap. Finally, two supermatrices were constructed: a "sparse" matrix based on the primary clusters alone (1794 taxa x 53,977 characters), and a somewhat more "dense" matrix based on the secondary clusters (2228 taxa x 33,168 characters). Both matrices were very sparse, with 95% of their cells containing gaps or question marks. These were subjected to extensive heuristic parsimony analyses using deterministic and stochastic heuristics, including bootstrap analyses. A "reduced consensus" bootstrap analysis was also performed to detect cryptic signal in a subtree of the data set corresponding to a "backbone" phylogeny proposed in previous studies. Overall, the dense supermatrix appeared to provide much more satisfying results, indicated by better resolution of the bootstrap tree, excellent agreement with the backbone papilionoid tree in the reduced bootstrap consensus analysis, few problematic large polytomies in the strict consensus, and less fragmentation of conventionally recognized genera. Nevertheless, at lower taxonomic levels several problems were identified and diagnosed. A large number of methodological issues in supermatrix construction at this scale are discussed, including detection of annotation errors in GenBank sequences; the shortage of effective algorithms and software for local multiple sequence alignment; the difficulty of overcoming effects of fragmentation of data into nearly disjoint blocks in sparse supermatrices; and the lack of informative tools to assess confidence limits in very large trees.

Algorithms↗

Genome trees and the tree of life.

Genome comparisons indicate that horizontal gene transfer and differential gene loss are major evolutionary phenomena that, at least in prokaryotes, involve a large fraction, if not the majority, of genes. The extent of these events casts doubt on the feasibility of constructing a 'Tree of Life', because the trees for different genes often tell different stories. However, alternative approaches to tree construction that attempt to determine tree topology on the basis of comparisons of complete gene sets seem to reveal a phylogenetic signal that supports the three-domain evolutionary scenario and suggests the possibility of delineation of previously undetected major clades of prokaryotes. If the validity of these whole-genome approaches to tree building is confirmed by analyses of numerous new genomes, which are currently being sequenced at an increasing rate, it would seem that the concept of a universal 'species' tree is still appropriate. However, this tree should be reinterpreted as a prevailing trend in the evolution of genome-scale gene sets rather than as a complete picture of evolution.

Animals↗

Hunting for trees in binary character sets: efficient algorithms for extraction, enumeration, and optimization.

We describe a useful and efficient tree building technique that is closely related to the maximum compatible subset method. Given a set of binary characters C and a degree bound d, the algorithm we describe counts the number of trees T that satisfy (1) the edges of T correspond to characters contained in C, and (2) the vertices of T have degree bounded by d. If the characters are weighted then we can efficiently determine which of these trees have maximum summed edge weight, where the weight of an edge equals the weight of the corresponding binary character in C. The complexity of the algorithms equals O(nk + ndK(d-1), where n is the number of taxa, k is the number of characters, and K is the number of distinct characters. A number of new tree consensus methods based on the algorithm are introduced, including one that uses edge weighted trees. As well, we illustrate the applicability of the algorithm for large sequence data sets by analyzing sequences from the "Out of Africa" mtDNA data set.

Algorithms↗