PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Tree building”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Mitochondrial DNA sequences of five squamates: phylogenetic affiliation of snakes.

Complete or nearly complete mitochondrial DNA sequences were determined from four lizards (Western fence lizard, Warren's spinytail lizard, Terrestrial arboreal alligator lizard, and Chinese crocodile lizard) and a snake (Texas blind snake). These genomes had a typical gene organization found in those of most mammals and fishes, except for a translocation of the glutamine tRNA gene in the blind snake and a tandem duplication of the threonine and proline tRNA genes in the spinytail lizard. Although previous work showed the existence of duplicate control regions in mitochondrial DNAs of several snakes, the blind snake did not have this characteristic. Phylogenetic analyses based on different tree-building methods consistently supported that the blind snake and a colubrid snake (akamata) make a sister clade relative to all the lizard taxa from six different families. An alternative hypothesis that snakes evolved from a lineage of varanoids was not favored and nearly statistically rejected by the Kishino-Hasegawa test. It is therefore likely that the apparent similarity of the tongue structure between snakes and varanoids independently evolved and that the duplication of the control region occurred on a snake lineage after divergence of the blind snake.

Animals↗

Phylogenetically informative length polymorphism and sequence variability in mitochondrial DNA of Australian songbirds (Pomatostomus).

A combination of restriction analysis and direct sequencing via the polymerase chain reaction (PCR) was used to build trees relating mitochondrial DNAs (mtDNAs) from 50 individuals belonging to five species of Australian babblers (Pomatostomus). The trees served as a quantitative framework for analyzing the direction and tempo of evolution of an intraspecific length polymorphism from a third mitochondrial ancestor. The length polymorphism lies between the cytochrome b and 12S rRNA (srRNA) genes. Screening of mtDNAs within and between the five species with restriction enzymes showed that Pomatosomus temporalis was polymorphic for two smaller size classes (M and S) that are completely segregated geographically, whereas mtDNAs from the other four species were exclusively of a third, larger size (L). Inter- and intraspecific phylogenetic trees relating mtDNAs based on restriction maps, cytochrome b sequences obtained via PCR, and the two data sets combined were compared to one another statistically and were broadly similar except for the phylogenetic position of Pomatosomus halli. Both sets of phylogenies imply that only two deletion events can account for the observed intraspecific distribution of the three length types. High levels of base-substitutional divergence were detected within and between northern and southern lineages of P. temporalis, which implies a low level of gene flow between northern and southern regions as well as a low rate of length mutation. These conclusions were confirmed by applying coalescent theory to the statistical framework provided by the phylogenetic analyses.

Animals↗

Network analysis provides insights into evolution of 5S rDNA arrays in Triticum and Aegilops.

We have used network analysis to study gene sequences of the Triticum and Aegilops 5S rDNA arrays, as well as the spacers of the 5S-DNA-A1 and 5S-DNA-2 loci. Network analysis describes relationships between 5S rDNA sequences in a more realistic fashion than conventional tree building because it makes fewer assumptions about the direction of evolution, the extent of sexual isolation, and the pattern of ancestry and descent. The networks show that the 5S rDNA sequences of Triticum and Aegilops species are related in a reticulate manner around principal nodal sequences. The spacer networks have multiple principal nodes of considerable antiquity but the gene network has just one principal node, corresponding to the correct gene sequence. The networks enable orthologous groups of spacer sequences to be identified. When orthologs are compared it is seen that the patterns of intra- and interspecific diversity are similar for both genes and spacers. We propose that 5S rDNA arrays combine sequence conservation with a large store of mutant variations, the number of correct gene copies within an array being the result of neutral processes that act on gene and spacer regions together.

Algorithms↗

QNet: an agglomerative method for the construction of phylogenetic networks from weighted quartets.

We present QNet, a method for constructing split networks from weighted quartet trees. QNet can be viewed as a quartet analogue of the distance-based Neighbor-Net (NNet) method for network construction. Just as NNet, QNet works by agglomeratively computing a collection of circular weighted splits of the taxa set which is subsequently represented by a planar split network. To illustrate the applicability of QNet, we apply it to a previously published Salmonella data set. We conclude that QNet can provide a useful alternative to NNet if distance data are not available or a character-based approach is preferred. Moreover, it can be used as an aid for determining when a quartet-based tree-building method may or may not be appropriate for a given data set. QNet is freely available for download.

Algorithms↗

The Ribosomal Database Project (RDP-II): previewing a new autoaligner that allows regular updates and the new prokaryotic taxonomy.

The Ribosomal Database Project-II (RDP-II) pro-vides data, tools and services related to ribosomal RNA sequences to the research community. Through its website (http://rdp.cme.msu.edu), RDP-II offers aligned and annotated rRNA sequence data, analysis services, and phylogenetic inferences (trees) derived from these data. RDP-II release 8.1 contains 16 277 prokaryotic, 5201 eukaryotic, and 1503 mitochondrial small subunit rRNA sequences in aligned and annotated format. The current public beta release of 9.0 debuts a new regularly updated alignment of over 50 000 annotated (eu)bacterial sequences. New analysis services include a sequence search and selection tool (Hierarchy Browser) and a phylogenetic tree building and visualization tool (Phylip Interface). A new interactive tutorial guides users through the basics of rRNA sequence analysis. Other services include probe checking, phylogenetic placement of user sequences, screening of users' sequences for chimeric rRNA sequences, automated alignment, production of similarity matrices, and services to plan and analyze terminal restriction fragment polymorphism (T-RFLP) experiments. The RDP-II email address for questions or comments is rdpstaff@msu.edu.

Animals↗

Detecting non-orthology in the COGs database and other approaches grouping orthologs using genome-specific best hits.

Correct orthology assignment is a critical prerequisite of numerous comparative genomics procedures, such as function prediction, construction of phylogenetic species trees and genome rearrangement analysis. We present an algorithm for the detection of non-orthologs that arise by mistake in current orthology classification methods based on genome-specific best hits, such as the COGs database. The algorithm works with pairwise distance estimates, rather than computationally expensive and error-prone tree-building methods. The accuracy of the algorithm is evaluated through verification of the distribution of predicted cases, case-by-case phylogenetic analysis and comparisons with predictions from other projects using independent methods. Our results show that a very significant fraction of the COG groups include non-orthologs: using conservative parameters, the algorithm detects non-orthology in a third of all COG groups. Consequently, sequence analysis sensitive to correct orthology assignments will greatly benefit from these findings.

Algorithms↗

Greene SCPrimer: a rapid comprehensive tool for designing degenerate primers from multiple sequence alignments.

Polymerase chain reaction (PCR) is widely applied in clinical and environmental microbiology. Primer design is key to the development of successful assays and is often performed manually by using multiple nucleic acid alignments. Few public software tools exist that allow comprehensive design of degenerate primers for large groups of related targets based on complex multiple sequence alignments. Here we present a method for designing such primers based on tree building followed by application of a set covering algorithm, and demonstrate its utility in compiling Multiplex PCR primer panels for detection and differentiation of viral pathogens.

Algorithms↗

Structure-based phylogenetic analysis of short-chain alcohol dehydrogenases and reclassification of the 17beta-hydroxysteroid dehydrogenase family.

Short-chain alcohol dehydrogenases (SCAD) constitute a large and diverse family of ancient origin. Several of its members play an important role in human physiology and disease, especially in the metabolism of steroid substrates (e.g., prostaglandins, estrogens, androgens, and corticosteroids). Their involvement in common human disorders such as endocrine-related cancer, osteoporosis, and Alzheimer disease makes them an important candidate for drug targets. Recent phylogenetic analysis of SCAD is incomplete and does not allow any conclusions on very ancient divergences or on a functional characterization of novel proteins within this complex family. We have developed a 3D structure-based approach to establish the deep-branching pattern within the SCAD family. In this approach, pairwise superpositions of X-ray structures were used to calculate similarity scores as an input for a tree-building algorithm. The resulting phylogeny was validated by comparison with the results of sequence-based algorithms and biochemical data. It was possible to use the 3D data as a template for the reliable determination of the phylogenetic position of novel proteins as a first step toward functional predictions. We were able to discern new patterns in the phylogenetic relationships of the SCAD family, including a basal dichotomy of the 17beta-hydroxysteroid dehydrogenases (17beta-HSDs). These data provide an important contribution toward the development of type-specific inhibitors for 17beta-HSDs for the treatment and prevention of disease. Our structure-based phylogenetic approach can also be applied to increase the reliability of evolutionary reconstructions in other large protein families.

17-Hydroxysteroid Dehydrogenases↗

Testing the new animal phylogeny: first use of combined large-subunit and small-subunit rRNA gene sequences to classify the protostomes.

Although the small-subunit ribosomal RNA (SSU rRNA) gene is widely used in the molecular systematics, few large-subunit (LSU) rRNA gene sequences are known from protostome animals, and the value of the LSU gene for invertebrate systematics has not been explored. The goal of this study is to test whether combined LSU and SSU rRNA gene sequences support the division of protostomes into Ecdysozoa (molting forms) and Lophotrochozoa, as was proposed by Aguinaldo et al. (1997) (Nature 387:489) based on SSU rRNA sequences alone. Nearly complete LSU gene sequences were obtained, and combined LSU + SSU sequences were assembled, for 15 distantly related protostome taxa plus five deuterostome outgroups. When the aligned LSU + SSU sequences were analyzed by tree-building methods (minimum evolution analysis of LogDet-transformed distances, maximum likelihood, and maximum parsimony) and by spectral analysis of LogDet distances, both Ecdysozoa and Lophotrochozoa were indeed strongly supported (e.g., bootstrap values >90%), with higher support than from the SSU sequences alone. Furthermore, with the LogDet-based methods, the LSU + SSU sequences resolved some accepted subgroups within Ecdysozoa and Lophotrochozoa (e.g., the polychaete sequence grouped with the echiuran, and the annelid sequences grouped with the mollusc and lophophorates)-subgroups that SSU-based studies do not reveal. Also, the mollusc sequence grouped with the sequences from lophophorates (brachiopod and phoronid). Like SSU sequences, our LSU + SSU sequences contradict older hypotheses that grouped annelids with arthropods as Articulata, that said flatworms and nematodes were basal bilateralians, and considered lophophorates, nemerteans, and chaetognaths to be deuterostomes. The position of chaetognaths within protostomes remains uncertain: our chaetognath sequence associated with that of an onychophoran, but this was unstable and probably artifactual. Finally, the benefits of combining LSU with SSU sequences for phylogenetic analyses are discussed: LSU adds signal, it can be used at lower taxonomic levels, and its core region is easy to align across distant taxa-but its base frequencies tend to be nonstationary across such taxa. We conclude that molecular systematists should use combined LSU + SSU rRNA genes rather than SSU alone.

Animals↗

Evaluating hypotheses of deuterostome phylogeny and chordate evolution with new LSU and SSU ribosomal DNA data.

We investigated evolutionary relationships among deuterostome subgroups by obtaining nearly complete large-subunit ribosomal RNA (LSU rRNA)-gene sequences for 14 deuterostomes and 3 protostomes and complete small-subunit (SSU) rRNA-gene sequences for five of these animals. With the addition of previously published sequences, we compared 28 taxa using three different data sets (LSU only, SSU only, and combined LSU + SSU) under minimum evolution (with LogDet distances), maximum likelihood, and maximum parsimony optimality criteria. Additionally, we analyzed the combined LSU + SSU sequences with spectral analysis of LogDet distances, a technique that measures the amount of support and conflict within the data for every possible grouping of taxa. Overall, we found that (1) the LSU genes produced a tree very similar to the SSU gene tree, (2) adding LSU to SSU sequences strengthened the bootstrap support for many groups above the SSU-only values (e.g., hemichordates plus echinoderms as Ambulacraria; lancelets as the sister group to vertebrates), (3) LSU sequences did not support SSU-based hypotheses of pterobranchs evolving from enteropneusts and thaliaceans evolving from ascidians, and (4) the combined LSU + SSU data are ambiguous about the monophyly of chordates. No tree-building algorithm united urochordates conclusively with other chordates, although spectral analysis did so, providing our only evidence for chordate monophyly. With spectral analysis, we also evaluated several major hypotheses of deuterostome phylogeny that were constructed from morphological, embryological, and paleontological evidence. Our rRNA-gene analysis refutes most of these hypotheses and thus advocates a rethinking of chordate and vertebrate origins.

Animals↗

Extreme differences in rates of molecular evolution of foraminifera revealed by comparison of ribosomal DNA sequences and the fossil record.

Foraminifera have one of the best known fossil records among the unicellular eukaryotes. However, the origin and phylogenetic relationships of the extant foraminiferal lineages are poorly understood. To test the current paleontological hypotheses on evolution of foraminifera, we sequenced about 1,000 base pairs from the 3' end of the small subunit rRNA gene (SSU rDNA) in 22 species representing all major taxonomic groups. Phylogenies were derived using neighbor-joining, maximum-parsimony, and maximum-likelihood methods. All analyses confirm the monophyletic origin of foraminifera. Evolutionary relationships within foraminifera inferred from rDNA sequences, however, depend on the method of tree building and on the choice of analyzed sites. In particular, the position of planktonic foraminifera shows important variations. We have shown that these changes result from the extremely high rate of rDNA evolution in this group. By comparing the number of substitutions with the divergence times inferred from the fossil record, we have estimated that the rate of rDNA evolution in planktonic foraminifera is 50 to 100 times faster than in some benthic foraminifera. The use of the maximum-likelihood method and limitation of analyzed sites to the most conserved parts of the SSU rRNA molecule render molecular and paleontological data generally congruent.

Animals↗

Weighted neighbor joining: a likelihood-based approach to distance-based phylogeny reconstruction.

We introduce a distance-based phylogeny reconstruction method called "weighted neighbor joining," or "Weighbor" for short. As in neighbor joining, two taxa are joined in each iteration; however, the Weighbor criterion for choosing a pair of taxa to join takes into account that errors in distance estimates are exponentially larger for longer distances. The criterion embodies a likelihood function on the distances, which are modeled as correlated Gaussian random variables with different means and variances, computed under a probabilistic model for sequence evolution. The Weighbor criterion consists of two terms, an additivity term and a positivity term, that quantify the implications of joining the pair. The first term evaluates deviations from additivity of the implied external branches, while the second term evaluates confidence that the implied internal branch has a positive branch length. Compared with maximum-likelihood phylogeny reconstruction, Weighbor is much faster, while building trees that are qualitatively and quantitatively similar. Weighbor appears to be relatively immune to the "long branches attract" and "long branch distracts" drawbacks observed with neighbor joining, BIONJ, and parsimony.

Animals↗

Mammalian phylogeny: comparison of morphological and molecular results.

In an attempt to resolve the "bushy" part at the root of the eutherian tree, 182 nondental morphological characters from 100 species (79 extant and 21 extinct; 98 mammalian and 2 nonmammalian) were analyzed using two maximum-parsimony tree-building algorithms. Parallel analyses of 2,258 pairwise immunodiffusion comparisons with chicken antisera on 101 mammalian species and of amino acid sequence data of alpha and beta hemoglobins and other published protein sequences were also carried out. The morphological and molecular phylogenies agree in depicting the infraclass Eutheria as consisting of five major clades (thus resolving part of the "bush"). Rates of evolution were also found to be similar in the two types of phylogenies.

Amino Acid Sequence↗

Pseudomonas argentinensis sp. nov., a novel yellow pigment-producing bacterial species, isolated from rhizospheric soil in Cordoba, Argentina.

During a study in the Argentinian region of Chaco (Cordoba), some strains were isolated from the rhizosphere of grasses growing in semi-desertic arid soils. Two of these strains, one isolated from the rhizospheric soil of Chloris ciliata (strain CH01(T)) and the other from Pappophorum caespitosum (strain PA01), were Gram-negative, strictly aerobic rods, which formed yellow round colonies on nutrient agar. They produced a water-insoluble yellow pigment, and a fluorescent pigment was also detected. A polyphasic taxonomic approach was used to characterize the strains. Comparison of the 16S rRNA gene sequences showed a similarity of 99.3 % between them, and phylogenetic analysis revealed that the strains belong to the genus Pseudomonas, within the gamma-subclass of the Proteobacteria. The closest related species is Pseudomonas straminea IAM 1598(T) (similarity of 99.0 % to strain CH01(T) and 98.8 % to strain PA01), clustering in a separate branch with the various methods of tree building used. Strains CH01(T) and PA01 both had a single polar flagellum, like other yellow pigment-producing pseudomonads related to them. Both strains produced catalase and oxidase. Similar to P. straminea, they did not hydrolyse gelatin or casein. The G+C DNA contents determined were 57.5 mol% for CH01(T) and 58.0 mol% for PA01. DNA-DNA hybridization results showed 81 % relatedness between them, and only 40-44 % relatedness with respect to the type strain of P. straminea. These results, together with other phenotypic characteristics, support the conclusion that both isolates belong to the same species, and should be described as representing a novel species within the genus Pseudomonas, for which the name Pseudomonas argentinensis sp. nov. is proposed. The type strain is CH01(T) (=LMG 22563(T) = CECT 7010(T)).

Aerobiosis↗

Tenofovir resistance and resensitization.

Human immunodeficiency viruses in 321 samples from tenofovir-naïve patients were retrospectively evaluated for resistance to this nucleotide analogue. All virus strains with insertions between amino acids 67 and 70 of the reverse transcriptase (n = 6) were highly resistant. Virus strains with the Q151M mutation were divided into susceptible (n = 12) and highly resistant (n = 8) viruses. This difference was due to the absence or presence of the K65R mutation, which was confirmed by site-directed mutagenesis. Viral clones with various combinations of the mutations M41L, K70R, L210W, and T215F or T215Y were analyzed for cross-resistance induced by thymidine analogue mutations (TAMs). The levels of increased resistance induced by single, double, and triple mutations at the indicated positions could be ranked as follows: for mutants with single mutations, mutations at positions 41 > 215 > 70; for mutants with double mutations, mutations at positions 41 and 215 > 70 and 215 = 210 and 215 > 41 and 70; for mutants with triple mutations, mutations at positions 41, 210, and 215 > 41, 70, and 215. Viral clones with M184V or M184I exhibited slightly increased susceptibilities to tenofovir (0.7-fold). Almost all clones with TAM-induced resistance were resensitized when M184V was present (P < 0.001). Among the viruses in the clinical samples, the rate of tenofovir resistance significantly increased with the number of TAMs both in the samples with 184M and in those with 184V (P = 0.005 and P = 0.003, respectively). A resensitizing effect of M184V was confirmed for all samples exhibiting at least one TAM (P = 0.03). However, accumulation of at least two TAMs resulted in more than 2.0-fold reduced susceptibility to tenofovir, irrespective of the presence of M184V. Decision tree building, a classical machine learning technique, was used to generate models for the interpretation of mutations with respect to tenofovir resistance. The application of previously proposed cutoffs for a reduced response to therapy and treatment failure demonstrated the central roles of positions 215 and 65 for 1.5- and 4.0-fold reduced susceptibilities, respectively. Thus, clinically relevant resistance may be conferred by the accumulation of TAMs, and the resensitizing effect of M184V should be considered only minor.

Adenine↗

TT virus infection in French hemodialysis patients: study of prevalence and risk factors.

The TT virus (TTV) is a recently discovered DNA virus which was first identified in patients with non-A to -G hepatitis following blood transfusion. In this study, we tested 150 attendees of two hemodialysis (HD) units of the public hospitals of Marseilles, France, for the presence of TTV genome by using a PCR-based methodology. The overall prevalence of TTV viremia was 28% (compared to 5.3% in blood donors from the same region). We demonstrated the existence of chronic infections and superinfections by strains belonging to different genotypes. The prevalence of infection was higher in patients originating from Africa, in patients with previous blood transfusion or organ transplantation, in patients with antibody to hepatitis B core antigen, and in those with diabetes mellitus. A high prevalence of TTV infection (50%) was also observed in a population of patients with diabetes mellitus but without renal disease. No significant relationship was found between TTV viremia and hepatitis C virus or GB virus C, transaminases, age, sex, and duration of HD treatment. The PCR amplification products (located in open reading frame 1 of the TTV genome) were sequenced. These genomic sequences were submitted to phylogenetic analysis by using the Jukes-Cantor algorithm for distance determination and the neighbor-joining method for tree building. In several instances, sequences from viruses isolated in a HD unit were grouped in the same phylogenetic cluster. These results together with the different distribution of cases in the two HD units suggest there is viral transmission within each.

Adult↗

Randomly amplified polymorphic DNA as a valuable tool for epidemiological studies of Paracoccidioides brasiliensis.

Randomly amplified polymorphic DNA (RAPD) has been successfully used to detect genetic variations among isolates of Paracoccidioides brasiliensis. However, the usefulness of this technique for assessing important parasitic properties is still unconfirmed. In the present work we further investigated the applicability of RAPD in revealing important intrinsic and extrinsic features of this fungus associated with geographical origin, time of isolation, source of clinical specimen, clinical forms of human disease and also in vitro and in vivo susceptibility to antimicrobial and antifungal drugs. The RAPD patterns allowed us to distinguish all of the analyzed strains, which included 26 clinical isolates, 2 animal isolates, and 1 environmental isolate of P. brasiliensis obtained from different geographic regions, confirming the strong discriminating power of this technique. A phenetic tree, build from the RAPD data, showed that although the two nonclinical Brazilian strains were set together the majority of the clinical Brazilian strains were randomly distributed through different sub-branches of a major cluster without any correlation to any of the parameters analyzed. A second major cluster, however, has grouped isolates from Mato Grosso and Roraima (Brazil) that not only were susceptible in vitro to trimethoprim-sulfamethoxazole but also produced a good in vivo response. These results open new vistas for epidemiological and clinical studies of P. brasiliensis.

Amphotericin B↗

SIMPROT: using an empirically determined indel distribution in simulations of protein evolution.

BACKGROUND: General protein evolution models help determine the baseline expectations for the evolution of sequences, and they have been extensively useful in sequence analysis and for the computer simulation of artificial sequence data sets. RESULTS: We have developed a new method of simulating protein sequence evolution, including insertion and deletion (indel) events in addition to amino-acid substitutions. The simulation generates both the simulated sequence family and a true sequence alignment that captures the evolutionary relationships between amino acids from different sequences. Our statistical model for indel evolution is based on the empirical indel distribution determined by Qian and Goldstein. We have parameterized this distribution so that it applies to sequences diverged by varying evolutionary times and generalized it to provide flexibility in simulation conditions. Our method uses a Monte-Carlo simulation strategy, and has been implemented in a C++ program named Simprot. CONCLUSION: Simprot will be useful for testing methods of analysis of protein sequence families particularly alignment methods, phylogenetic tree building, detection of recombination and horizontal gene transfer, and homology detection, where knowing the true course of sequence evolution is essential.

Amino Acid Substitution↗