PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Tree building”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

S-TREE: self-organizing trees for data clustering and online vector quantization.

This paper introduces S-TREE (Self-Organizing Tree), a family of models that use unsupervised learning to construct hierarchical representations of data and online tree-structured vector quantizers. The S-TREE1 model, which features a new tree-building algorithm, can be implemented with various cost functions. An alternative implementation, S-TREE2, which uses a new double-path search procedure, is also developed. The performance of the S-TREE algorithms is illustrated with data clustering and vector quantization examples, including a Gauss-Markov source benchmark and an image compression application. S-TREE performance on these tasks is compared with the standard tree-structured vector quantizer (TSVQ) and the generalized Lloyd algorithm (GLA). The image reconstruction quality with S-TREE2 approaches that of GLA while taking less than 10% of computer time. S-TREE1 and S-TREE2 also compare favorably with the standard TSVQ in both the time needed to create the codebook and the quality of image reconstruction.

Algorithms↗

BLASTing proteomes, yielding phylogenies.

We develop a procedure called RiPE (Retrieval-induced Phylogeny Environment) that automatically performs an evolutionary analysis of a protein (sub)family, (i) by retrieving the relevant sequences via a homology search, (ii) by using the search report to construct the alignment using only homologous subsequences (taking into account their neighborhood with a low chance of homology), (iii) by realigning, and (iv) by generating phylogenetic trees based on the alignment. In a first implementation of our scheme, we start with the available proteome data of model organisms, perform a PSI-BLAST search, use MView to convert hits into a multiple alignment, and perform realignment and tree building. As a test case, we have investigated the human ABC transporters of the subfamily G, starting with the five known human ABCG transporters. Our method retrieved homologous sequences not previously analyzed, generating a tree that is more plausible and better supported than previously published trees. The RiPE 0.1 prototype is available at the RiPE website, http://ifg-izkf.uni-muenster.de/fuellen/RiPE/ripe.html.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

Complete mitochondrial DNA genome sequences of extinct birds: ratite phylogenetics and the vicariance biogeography hypothesis.

The ratites have stimulated much debate as to how such large flightless birds came to be distributed across the southern continents, and whether they are a monophyletic group or are composed of unrelated lineages that independently lost the power of flight. Hypotheses regarding the relationships among taxa differ for morphological and molecular data sets, thus hindering attempts to test whether plate tectonic events can explain ratite biogeography. Here, we present the complete mitochondrial DNA genomes of two extinct moas from New Zealand, along with those of five extant ratites (the lesser rhea, the ostrich, the great spotted kiwi, the emu and the southern cassowary and two tinamous from different genera. The non-stationary base composition in these sequences violates the assumptions of most tree-building methods. When this bias is corrected using neighbour-joining with log-determinant distances and non-homogeneous maximum likelihood, the ratites are found to be monophlyletic, with moas basal, as in morphological trees. The avian sequences also violate a molecular clock, so we applied a non-parametric rate smoothing algorithm, which minimizes ancestor-descendant local rate changes, to date nodes in the tree. Using this method, most of the major ratite lineages fit the vicariance biogeography hypothesis, the exceptions being the ostrich and the kiwi, which require dispersal to explain their present distribution.

Animals↗

[The construction of an evolutionary tree for signal receptor proteins].

A method is suggested for evolutionary tree building for proteins with low level of similarity. The degree of similarity is computed by conservative parts for a given family. The method allows to reveal the relationship of amino acid sequences even in cases when conventional techniques yield negative results. Evolutionary tree for signal receptor proteins is built. All the investigated proteins (except STE2) display non-accidental similarity.

Amino Acid Sequence↗

A probabilistic treatment of phylogeny and sequence alignment.

Carrying out simultaneous tree-building and alignment of sequence data is a difficult computational task, and the methods currently available are either limited to a few sequences or restricted to highly simplified models of alignment and phylogeny. A method is given here for overcoming these limitations by Bayesian sampling of trees and alignments simultaneously. The method uses a standard substitution matrix model for residues together with a hidden Markov model structure that allows affine gap penalties. It escapes the heavy computational burdens of other models by using an approximation called the "*" rule, which replaces missing data by a sum over all possible values of variables. The behavior of the model is demonstrated on test sets of globins.

Bayes Theorem↗

Comparing sequences without using alignments: application to HIV/SIV subtyping.

BACKGROUND: In general, the construction of trees is based on sequence alignments. This procedure, however, leads to loss of informationwhen parts of sequence alignments (for instance ambiguous regions) are deleted before tree building. To overcome this difficulty, one of us previously introduced a new and rapid algorithm that calculates dissimilarity matrices between sequences without preliminary alignment. RESULTS: In this paper, HIV (Human Immunodeficiency Virus) and SIV (Simian Immunodeficiency Virus) sequence data are used to evaluate this method. The program produces tree topologies that are identical to those obtained by a combination of standard methods detailed in the HIV Sequence Compendium. Manual alignment editing is not necessary at any stage. Furthermore, only one user-specified parameter is needed for constructing trees. CONCLUSION: The extensive tests on HIV/SIV subtyping showed that the virus classifications produced by our method are in good agreement with our best taxonomic knowledge, even in non-coding LTR (Long Terminal Repeat) regions that are not tractable by regular alignment methods due to frequent duplications/insertions/deletions. Our method, however, is not limited to the HIV/SIV subtyping. It provides an alternative tree construction without a time-consuming aligning procedure.

Animals↗

RpoA: a useful gene for phylogenetic analysis in diatoms.

The aim of this study was to compare the usefulness of two chloroplast-encoded genes (rpoA and rbcL) and the nuclear-encoded small subunit (SSU) ribosomal RNA for reconstructing phylogenetic relationships among diatoms at lower taxonomic levels. To this end, the rpoA and rbcL genes for selected centric and pennate diatoms were sequenced. The new rpoA and rbcL sequences, and an existing nuclear-encoded SSU rRNA data set, were subjected to weighted/unweighted parsimony, maximum likelihood, minimum evolution, and Bayesian analyses. All of the tree-building methods employed showed, based on the support values, that the rpoA gene was the most useful, relative to the rbcL and SSU rRNA genes, in determining phylogenetic relationships among the sampled diatoms. The support values for the relationships among the pennate lineages were, in many instances, greater in the rpoA trees than in the SSU rRNA trees. These results suggest that rpoA might be of value in determining phylogenetic relationships among pennate lineages.

Base Sequence↗

On the quality of tree-based protein classification.

MOTIVATION: Phylogenetic analysis of protein sequences is widely used in protein function classification and delineation of subfamilies within larger families. In addition, the recent increase in the number of protein sequence entries with controlled vocabulary terms describing function (e.g. the Gene Ontology) suggests that it may be possible to overlay these terms onto phylogenetic trees to automatically locate functional divergence events in protein family evolution. Phylogenetic analysis of large datasets requires fast algorithms; and even 'fast', approximate distance matrix-based phylogenetic algorithms are slow on large datasets since they involve calculating maximum likelihood estimates of pairwise evolutionary distances. There have been many attempts to classify protein sequences on the family and subfamily level without reconstructing phylogenetic trees, but using hierarchical clustering with simpler distance measures, which also produce trees or dendrograms. How can these trees be compared in their ability to accurately classify protein sequences? RESULTS: Given a 'reference classification' or 'group membership labels' for a set of related protein sequences as well as a tree describing their relationships (e.g. a phylogenetic tree), we propose a method for dividing the tree into monophyletic or paraphyletic groups so as to optimize the correspondence between the reference groups and the tree-derived groups. We call the achieved optimal correspondence the 'accuracy of a tree-based classification (TBC)', which measures the ability of a tree to separate proteins of similar function into monophyletic or paraphyletic groups. We apply this measure to compare classical NJ and UPGMA phylogenetic trees with the trees obtained from hierarchical clustering using different protein similarity measures. Our preliminary analysis on a set of expert-curated protein families and alignments suggests that there is no uniformly superior algorithm, and that simple protein similarity measures combined with hierarchical clustering produce trees with reasonable and often the most accurate TBC. We used our measure to help us to design TIPS, a tree-building algorithm, based on agglomerative clustering with a similarity measure derived from profile scoring. TIPS is comparable with phylogenetic algorithms in terms of classification accuracy and is much faster on large protein families. Due to its time scalability and acceptable accuracy, TIPS is being used in the large-scale PANTHER protein classification project. The trees produced by different algorithms for different protein families can be viewed at http://panther.appliedbiosystems.com/pub/tree_quality/trees.jsp. For every tree and every level of classification granularity we provide the optimal TBC along with the reference classification. AVAILABILITY: The script that evaluates the accuracy of TBC is available at http://panther.appliedbiosystems.com/pub/tree_quality/index.jsp

Algorithms↗

Maximum likelihood methods reveal conservation of function among closely related kinesin families.

We have reconstructed the evolution of the anciently derived kinesin superfamily using various alignment and tree-building methods. In addition to classifying previously described kinesins from protists, fungi, and animals, we analyzed a variety of kinesin sequences from the plant kingdom including 12 from Zea mays and 29 from Arabidopsis thaliana. Also included in our data set were four sequences from the anciently diverged amitochondriate protist Giardia lamblia. The overall topology of the best tree we found is more likely than previously reported topologies and allows us to make the following new observations: (1) kinesins involved in chromosome movement including MCAK, chromokinesin, and CENP-E may be descended from a single ancestor; (2) kinesins that form complex oligomers are limited to a monophyletic group of families; (3) kinesins that crosslink antiparallel microtubules at the spindle midzone including BIMC, MKLP, and CENP-E are closely related; (4) Drosophila NOD and human KID group with other characterized chromokinesins; and (5) Saccharomyces SMY1 groups with kinesin-I sequences, forming a family of kinesins capable of class V myosin interactions. In addition, we found that one monophyletic clade composed exclusively of sequences with a C-terminal motor domain contains all known minus end-directed kinesins.

Animals↗

Ecdysozoan phylogeny and Bayesian inference: first use of nearly complete 28S and 18S rRNA gene sequences to classify the arthropods and their kin.

Relationships among the ecdysozoans, or molting animals, have been difficult to resolve. Here, we use nearly complete 28S+18S ribosomal RNA gene sequences to estimate the relations of 35 ecdysozoan taxa, including newly obtained 28S sequences from 25 of these. The tree-building algorithms were likelihood-based Bayesian inference and minimum-evolution analysis of LogDet-transformed distances, and hypotheses were tested wth parametric bootstrapping. Better taxonomic resolution and recovery of established taxa were obtained here, especially with Bayesian inference, than in previous parsimony-based studies that used 18S rRNA sequences (or 18S plus small parts of 28S). In our gene trees, priapulan worms represent the basal ecdysozoans, followed by nematomorphs, or nematomorphs plus nematodes, followed by Panarthropoda. Panarthropoda was monophyletic with high support, although the relationships among its three phyla (arthropods, onychophorans, tardigrades) remain uncertain. The four groups of arthropods-hexapods (insects and related forms), crustaceans, chelicerates (spiders, scorpions, horseshoe crabs), and myriapods (centipedes, millipedes, and relatives)-formed two well-supported clades: Hexapoda in a paraphyletic crustacea (Pancrustacea), and 'Chelicerata+Myriapoda' (a clade that we name 'Paradoxopoda'). Pycnogonids (sea spiders) were either chelicerates or part of the 'chelicerate+myriapod' clade, but not basal arthropods. Certain clades derived from morphological taxonomy, such as Mandibulata, Atelocerata, Schizoramia, Maxillopoda and Cycloneuralia, are inconsistent with these rRNA data. The 28S gene contained more signal than the 18S gene, and contributed to the improved phylogenetic resolution. Our findings are similar to those obtained from mitochondrial and nuclear (e.g., elongation factor, RNA polymerase, Hox) protein-encoding genes, and should revive interest in using rRNA genes to study arthropod and ecdysozoan relationships.

Animals↗

Rabbits, if anything, are likely Glires.

Rodentia (e.g., mice, rats, dormice, squirrels, and guinea pigs) and Lagomorpha (e.g., rabbits, hares, and pikas) are usually grouped into the Glires. Status of this controversial superorder has been evaluated using morphology, paleontology, and mitochondrial plus nuclear DNA sequences. This growing corpus of data has been favoring the monophyly of Glires. Recently, Misawa and Janke [Mol. Phylogenet. Evol. 28 (2003) 320] analyzed the 6441 amino acids of 20 nuclear proteins for six placental mammals (rat, mouse, rabbit, human, cattle, and dog) and two outgroups (chicken and xenopus), and observed a basal position of the two murine rodents among the former. They concluded that "the Glires hypothesis was rejected." We here reanalyzed [loc. cit.] data set under maximum likelihood and Bayesian tree-building approaches, using phylogenetic models that take into account among-site variation in evolutionary rates and branch-length variation among proteins. Our observations support both the association of rodents and lagomorphs and the monophyly of Euarchontoglires (=Supraprimates) as the most likely explanation of the protein alignments. We conducted simulation studies to evaluate the appropriateness of lissamphibian and avian outgroups to root the placental tree. When the outgroup-to-ingroup evolutionary distance increases, maximum parsimony roots the topology along the long Mus-Rattus branch. Maximum likelihood, in contrast, roots the topology along different branches as a function of their length. Maximum likelihood appears less sensitive to the "long-branch attraction artifact" than is parsimony. Our phylogenetic conclusions were confirmed by the analysis of a different protein data set using a similar sample of species but different outgroups. We also tested the effect of the addition of afrotherian and xenarthran taxa. Using the linearized tree method, [loc. cit.] estimated that mice and rats diverged about 35 million years ago. Molecular dating based on the Bayesian relaxed molecular clock method suggests that the 95% credibility interval for the split between mice and rats is 7-17 Mya. We here emphasize the need for appropriate models of sequence evolution (matrices of amino acid replacement, taking into account among-site rate variation, and independent parameters across independent protein partitions) and for a taxonomically broad sample, and conclude on the likelihood that rodents and lagomorphs together constitute a monophyletic group (Glires).

Animals↗

Mitochondrial phylogeny of the Cyprichromini, a lineage of open-water cichlid fishes endemic to Lake Tanganyika, East Africa.

We present a phylogeny of the Cyprichromini, a lineage of cichlid fishes from Lake Tanganyika, showing progressive adaptation towards pelagic life style. Our study is based upon three mitochondrial gene segments, 443 bp of the control region, 402 bp of the cytochrome b gene and the entire NADH dehydrogenase subunit 2 gene (1047 bp). The topologies obtained by different tree building methods subdivide the Cyprichromini into four distinct lineages: the Paracyprichromis-, the Cyprichromis zonatus-, the Cyprichromis microlepidotus-lineage, and a lineage comprising Cyprichromis pavo and Cyprichromis leptosoma. Our study thus corroborates the distinctness of C. zonatus which was recently described formally. Concerning ecology and mating behavior, a clear evolutionary trend towards progressive adaptation to the pelagic zone emerges during the evolution of the Cyprichromini. The linearized tree analysis further shows that the four lineages have split almost contemporaneously. The mean Kimura-2-parameter distance among the four lineages emerging from the primary radiation of the Cyprichromini amounts to 7.21% and is in close agreement to that previously found for the primary radiation of the tribe Tropheini (7.01%), a lineage of rock-dwelling cichlids endemic to Lake Tanganyika. To date, the influence of lake level fluctuations as promoters of diversification has been demonstrated only for rock-dwelling cichlids. Based on the agreement in temporary patterns of diversification, we suggest that Pleistocene lake level changes have left a similar genetic imprint in a group of cichlid fishes that progressively colonized the open water during their radiation.

Africa, Eastern↗

MULTICOMP: a program for preparing sequence data for phylogenetic analysis.

MULTICOMP is a program that assists in the phylogenetic analysis of DNA sequences. It streamlines sequence handling and analysis. Input is from either individual sequence files or a file of aligned sequences. It produces data on variation at DNA and amino acid sequence level and can also convert sequences to data formats suitable for PHYLIP, PAUP and MacClade phylogenetic inference programs. Further, two tree-building programs, NEIGHBOR and DNAPARS, of PHYLIP can be directly run from within it. MULTICOMP performs Sawyer's algorithm for detection of gene conversion. The program facilitates analysis using only part of the data of a data set or using two or more combined data sets.

Algorithms↗

Molecular phylogenetic relationships of Eastern Asian Cyprinidae (pisces: cypriniformes) inferred from cytochrome b sequences.

Complete mitochondrial cytochrome b sequences of 54 species, including 18 newly sequenced, were analyzed to infer the phylogenetic relationships within the family Cyprinidae in East Asia. Phylogenetic trees were generated using various tree-building methods, including Neighbor-joining (NJ), Maximum Parsimony (MP) and Maximum Likelihood (ML) methods, with Myxocyprinus asiaticus (family Catostomidae) as the designated outgroup. The results from NJ and ML methods were mostly similar, supporting some existing subfamilies within Cyprinidae as monophyletic, such as Cultrinae, Xenocyprinae and Gobioninae (including Gobiobotinae). However, genera within the subfamily "Danioninae" did not form a monophyletic group. The subfamily Leuciscinae was divided into two unrelated groups: the "Leuciscinae" in East Asia forming as a monophyletic group together with Cultrinae and Xenocyprinae, while the Leucisciriae in Europe, Siberia, and North America as another monophyletic group. The monophyly of subfamily Cyprininae sensu Howes was supported by NJ and ML trees and is basal in the tree. The position of Acheilognathinae, a widely accepted monophyletic group represented by Rhodeus sericeus, was not resolved.

Amino Acid Sequence↗

A constrained-syntax genetic programming system for discovering classification rules: application to medical data sets.

This paper proposes a new constrained-syntax genetic programming (GP) algorithm for discovering classification rules in medical data sets. The proposed GP contains several syntactic constraints to be enforced by the system using a disjunctive normal form representation, so that individuals represent valid rule sets that are easy to interpret. The GP is compared with C4.5, a well-known decision-tree-building algorithm, and with another GP that uses Boolean inputs (BGP), in five medical data sets: chest pain, Ljubljana breast cancer, dermatology, Wisconsin breast cancer, and pediatric adrenocortical tumor. For this last data set a new preprocessing step was devised for survival prediction. Computational experiments show that, overall, the GP algorithm obtained good results with respect to predictive accuracy and rule comprehensibility, by comparison with C4.5 and BGP.

Algorithms↗

Euclidian space and grouping of biological objects.

MOTIVATION: Biological objects tend to cluster into discrete groups. Objects within a group typically possess similar properties. It is important to have fast and efficient tools for grouping objects that result in biologically meaningful clusters. Protein sequences reflect biological diversity and offer an extraordinary variety of objects for polishing clustering strategies. Grouping of sequences should reflect their evolutionary history and their functional properties. Visualization of relationships between sequences is of no less importance. Tree-building methods are typically used for such visualization. An alternative concept to visualization is a multidimensional sequence space. In this space, proteins are defined as points and distances between the points reflect the relationships between the proteins. Such a space can also be a basis for model-based clustering strategies that typically produce results correlating better with biological properties of proteins. RESULTS: We developed an approach to classification of biological objects that combines evolutionary measures of their similarity with a model-based clustering procedure. We apply the methodology to amino acid sequences. On the first step, given a multiple sequence alignment, we estimate evolutionary distances between proteins measured in expected numbers of amino acid substitutions per site. These distances are additive and are suitable for evolutionary tree reconstruction. On the second step, we find the best fit approximation of the evolutionary distances by Euclidian distances and thus represent each protein by a point in a multidimensional space. The Euclidian space may be projected in two or three dimensions and the projections can be used to visualize relationships between proteins. On the third step, we find a non-parametric estimate of the probability density of the points and cluster the points that belong to the same local maximum of this density in a group. The number of groups is controlled by a sigma-parameter that determines the shape of the density estimate and the number of maxima in it. The grouping procedure outperforms commonly used methods such as UPGMA and single linkage clustering.

Algorithms↗

Interior-branch and bootstrap tests of phylogenetic trees.

We have compared statistical properties of the interior-branch and bootstrap tests of phylogenetic trees when the neighbor-joining tree-building method is used. For each interior branch of a predetermined topology, the interior-branch and bootstrap tests provide the confidence values, PC and PB, respectively, that indicate the extent of statistical support of the sequence cluster generated by the branch. In phylogenetic analysis these two values are often interpreted in the same way, and if PC and PB are high (say, > or = 0.95), the sequence cluster is regarded as reliable. We have shown that PC is in fact the complement of the P-value used in the standard statistical test, but PB is not. Actually, the bootstrap test usually underestimates the extent of statistical support of species clusters. The relationship between the confidence values obtained by the two tests varies with both the topology and expected branch lengths of the true (model) tree. The most conspicuous difference between PC and PB is observed when the true tree is starlike, and there is a tendency for the difference to increase as the number of sequences in the tree increases. The reason for this is that the bootstrap test tends to become progressively more conservative as the number of sequences in the tree increases. Unlike the bootstrap, the interior-branch test has the same statistical properties irrespective of the number of sequences used when a predetermined tree is considered. Therefore, the interior-branch test appears to be preferable to the bootstrap test as long as unbiased estimators of evolutionary distances are used. However, when the interior-branch is applied to a tree estimated from a given data set, PC may give an overestimate of statistical confidence. For this case, we developed a method for computing a modified version (P'C) of the PC value and showed that this P'C tends to give a conservative estimate of statistical confidence, though it is not as conservative as PB. In this paper we have introduced a model in which evolutionary distances between sequences follow a multivariate normal distribution. This model allowed us to study the relationships between the two tests analytically.

Computer Simulation↗

Nonisotopic single-strand conformation polymorphism analysis of sequence variability in ribosomal DNA expansion segments within the genus Trichinella (Nematoda: Adenophorea).

A nonisotopic single-strand conformation polymorphism (SSCP) approach was employed to 'fingerprint' sequence variability in the expansion segment 5 (ES5) of domain IV and the D3 domain of nuclear ribosomal DNA within and/or among isolates and individual muscle (first-stage) larvae representing all currently recognized species/genotypes of Trichinella. In addition, phylogenetic analyses of the D3 sequence data set, employing three different tree-building algorithms, examined the relationships among all of them. These analyses showed strong support that the encapsulated species T. spiralis and T. nelsoni formed a group to the exclusion of the other encapsulated species T. britovi and its related genotypes Trichinella T8 and T9 and T. murrelli, and T. nativa and Trichinella T6, and strong support that T. nativa and Trichinella T6 grouped together. Also, these eight encapsulated members grouped to the exclusion of the nonencapsulated species T. papuae and T. zimbabwensis and the three representatives of T. pseudospiralis investigated. The findings showed that nonencapsulated species constitute a complex group which is distinct from the encapsulated species and supported the current hypothesis that encapsulated Trichinella group external to the nonencapsulated forms, in accordance with independent biological and biochemical data sets.

Animals↗