PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Tree building”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Classification tree method for bacterial source tracking with antibiotic resistance analysis data.

Various statistical classification methods, including discriminant analysis, logistic regression, and cluster analysis, have been used with antibiotic resistance analysis (ARA) data to construct models for bacterial source tracking (BST). We applied the statistical method known as classification trees to build a model for BST for the Anacostia Watershed in Maryland. Classification trees have more flexibility than other statistical classification approaches based on standard statistical methods to accommodate complex interactions among ARA variables. This article describes the use of classification trees for BST and includes discussion of its principal parameters and features. Anacostia Watershed ARA data are used to illustrate the application of classification trees, and we report the BST results for the watershed.

Animals↗

Local similarity in RNA secondary structures.

We present a systematic treatment of alignment distance and local similarity algorithms on trees and forests. We build upon the tree alignment algorithm for ordered trees given by Jiang et. al (1995) and extend it to calculate local forest alignments, which is essential for finding local similar regions in RNA secondary structures. The time complexity of our algorithm is O(|F(1)| |F(2) deg(F(1)) deg(F(2)) (deg(F(1)) + deg(F(2))) where |F(i)| is the number of nodes in forest F(i) and deg (F(i)) is the degree of F(i). We provide carefully engineered dynamic programming implementations using dense, two-dimensional tables which considerably reduces the space requirement. We suggest a new representation of RNA secondary structures as forests that allow reasonable scoring of edit operations on RNA secondary structures. The comparison of RNA secondary structures is facilitated by a new visualization technique for RNA secondary structure alignments. Finally, we show how potential regulatory motifs can be discovered solely by their structural preservation, and independent of their sequence conservation and position.

Algorithms↗

A general approach to proving the minimality of phylogenetic trees illustrated by an example with a set of 23 vertebrates.

We have recently described a method of building phylogenetic trees and have outlined an approach for proving whether a particular tree is optimal for the data used. In this paper we describe in detail the method of establishing lower bounds on the length of a minimal tree by partitioning the data set into subsets. All characters that could be involved in duplications in the data are paired with all other such characters. A matching algorithm is then used to obtain the pairing of characters that reveals the most duplications in the data. This matching may still not account for all nucleotide substitutions on the tree. The structure of the tree is then used to help select subsets of three or more characters until the lower bound found by partitioning is equal to the length of the tree. The tree must then be a minimal tree since no tree can exist with a length less than that of the lower bound. The method is demonstrated using a set of 23 vertebrate cytochrome c sequences with the criterion of minimizing the total number of nucleotide substitutions. There are 131130 7045768798 96033440625 topologically distinct trees that can be constructed from this data set. The method described in this paper does identify 144 minimal tree variants. The method is general in the sense that it can be used for other data and other criteria of length. It need not however always be possible to prove a treee minimal but the method will give an upper and lower bound on the length of minimal trees.

Amino Acid Sequence↗

Trees for bees.

Controversy over the origins and evolution of social behaviour in the major groups of social bees (the corbiculate bees) has fuelled arguments over different approaches for building evolutionary trees. However, the application of different analytical methodologies does not explain why molecular and morphological data suggest strikingly different hypotheses for the evolution of eusociality in bees. Determining the phylogenetic root is expected to help resolve the question of the social evolution of corbiculate bees. However, this requires that the long branch attraction problem is overcome. This phenomenon affects both molecular and morphological data for corbiculate bees.

Journal Article↗

Applications of redundancy analysis for the detection of chemical response patterns to air pollution in lichen.

The lichens Ramalina celastri (Spreng.) Krog & Swinsc., Punctelia microsticta (Müll. Arg.) Krog and Canomaculina pilosa (Stizenb.) Elix & Hale were transplanted simultaneously to 17 urban-industrial sites in a northwestern area of Córdoba city, Argentina. The transplantation sites were set according to different environmental conditions: traffic, industries, tree cover, building height, topographic level, position in the block and distances from the river and from the power plant. Three months later, chlorophyll a, chlorophyll b, phaeophytin a, soluble proteins, hydroperoxy conjugated dienes, malondialdehyde concentration and sulfur accumulation were determined, and a pollution index was calculated for each sampling site. Redundancy analysis was applied to detect the variation pattern of the lichen variables that can be 'best' explained by the environmental variables considered. The present study provides information about both the specific pattern response of each species to atmospheric pollution, and environmental conditions that determine it. As regards pollutants emission sources R. celastri showed a chemical response associated mainly with pollutant released by the power plant and traffic. P. microsticta and C. pilosa responded mainly to industrial sources. Regarding environmental conditions that affect the spreading of air pollutants and their incidence on the bioindicator, the topographic level and tree cover surrounding the sampling site were found to be important for R. celastri, tree cover surrounding the sampling site and the building height affected P. microsticta, while building height did so for C. pilosa.

Air Pollutants↗

TreeDyn: towards dynamic graphics and annotations for analyses of trees.

BACKGROUND: Analyses of biomolecules for biodiversity, phylogeny or structure/function studies often use graphical tree representations. Many powerful tree editors are now available, but existing tree visualization tools make little use of meta-information related to the entities under study such as taxonomic descriptions or gene functions that can hardly be encoded within the tree itself (if using popular tree formats). Consequently, a tedious manual analysis and post-processing of the tree graphics are required if one needs to use external information for displaying or investigating trees. RESULTS: We have developed TreeDyn, a tool using annotations and dynamic graphical methods for editing and analyzing multiple trees. The main features of TreeDyn are 1) the management of multiple windows and multiple trees per window, 2) the export of graphics to several standard file formats with or without HTML encapsulation and a new format called TGF, which enables saving and restoring graphical analysis, 3) the projection of texts or symbols facing leaf labels or linked to nodes, through manual pasting or by using annotation files, 4) the highlight of graphical elements after querying leaf labels (or annotations) or by selection of graphical elements and information extraction, 5) the highlight of targeted trees according to a source tree browsed by the user, 6) powerful scripts for automating repetitive graphical tasks, 7) a command line interpreter enabling the use of TreeDyn through CGI scripts for online building of trees, 8) the inclusion of a library of packages dedicated to specific research fields involving trees. CONCLUSION: TreeDyn is a tree visualization and annotation tool which includes tools for tree manipulation and annotation and uses meta-information through dynamic graphical operators or scripting to help analyses and annotations of single trees or tree collections.

Computer Graphics↗

Locating protein coding regions in human DNA using a decision tree algorithm.

Genes in eukaryotic DNA cover hundreds or thousands of base pairs, while the regions of those genes that code for proteins may occupy only a small percentage of the sequence. Identifying the coding regions is of vital importance in understanding these genes. Many recent research efforts have studied computational methods for distinguishing between coding and noncoding regions, and several promising results have been reported. We describe here a new approach, using a machine learning system that builds decision trees from the data. This approach combines several coding measures to produce classifiers with consistently higher accuracies than previous methods, on DNA sequences ranging from 54 to 162 base pairs in length. The algorithm is very efficient, and it can easily be adapted to different sequence lengths. Our conclusion is that decision trees are a highly effective tool for identifying protein coding regions.

Algorithms↗

Decision tree-based formation of consensus protein secondary structure prediction.

MOTIVATION: Prediction of protein secondary structure provides information that is useful for other prediction methods like fold recognition and ab initio 3D prediction. A consensus prediction constructed from the output of several methods should yield more reliable results than each of the individual methods. METHOD: We present an approach that reveals subtle but systematic differences in the output of different secondary structure prediction methods allowing the derivation of coherent consensus predictions. The method uses a machine learning technique that builds decision trees from existing data. RESULTS: The first results of our analysis show that consensus prediction of protein secondary structure may be improved both quantitatively and qualitatively.

Algorithms↗

A bi-dimensional regression tree approach to the modeling of gene expression regulation.

MOTIVATION: The transcriptional regulation of a gene depends on the binding of cis-regulatory elements on its promoter to some transcription factors and the expression levels of the transcription factors. Most existing approaches to studying transcriptional regulation model these dependencies separately, i.e. either from promoters to gene expression or from the expression levels of transcription factors to the expression levels of genes. Little effort has been devoted to a single model for integrating both dependencies. RESULTS: We propose a novel method to model gene expression using both promoter sequences and the expression levels of putative regulators. The proposed method, called bi-dimensional regression tree (BDTree), extends a multivariate regression tree approach by applying it simultaneously to both genes and conditions of an expression matrix. The method produces hypotheses about the condition-specific binding motifs and regulators for each gene. As a side-product, the method also partitions the expression matrix into small submatrices in a way similar to bi-clustering. We propose and compare several splitting functions for building the tree. When applied to two microarray datasets of the yeast Saccharomyces cerevisiae, BDTree successfully identifies most motifs and regulators that are known to regulate the biological processes underlying the datasets. Comparing with an existing algorithm, BDTree provides a higher prediction accuracy in cross-validations.

Algorithms↗

Comprehensible evaluation of prognostic factors and prediction of wound healing.

We analyzed the data of a controlled clinical study of the chronic wound healing acceleration as a result of electrical stimulation. The study involved a conventional conservative treatment, sham treatment, biphasic pulsed current, and direct current electrical stimulation. Data was collected over 10 years and suffices for an analysis with machine learning methods. So far, only a limited number of studies have investigated the wound and patient attributes which affect the chronic wound healing. There is none to our knowledge to include treatment attributes. The aims of our study are to determine effects of the wound, patient and treatment attributes on the wound healing process and to propose a system for prediction of the wound healing rate. First we analyzed which wound and patient attributes play a predominant role in the wound healing process and investigated a possibility to predict the wound healing rate at the beginning of the treatment based on the initial wound, patient and treatment attributes. Later we tried to enhance the wound healing rate prediction accuracy by predicting it after a few weeks of the wound healing follow-up. Using the attribute estimation algorithms ReliefF and RReliefF we obtained a ranking of the prognostic factors which was comprehensible to experts. We used regression and classification trees to build models for prediction of the wound healing rate. The obtained results are encouraging and may form a basis for an expert system for the chronic wound healing rate prediction. If the wound healing rate is known, then the provided information can help to formulate the appropriate treatment decisions and orient resources towards individuals with poor prognosis.

Algorithms↗

Clinical and electrophysiological predictors of respiratory failure in Guillain-Barré syndrome: a prospective study.

BACKGROUND: Respiratory failure is the most serious short-term complication of Guillain-Barré syndrome and can require invasive mechanical ventilation in 20-30% of patients. We sought to identify clinical and electrophysiological predictors of respiratory failure in the disease. METHODS: We prospectively assessed electrophysiological data and clinical factors, including identified predictors of delay between disease onset and admission, inability to lift head, and vital capacity, in patients admitted with Guillain-Barré syndrome. We related these factors to subsequent need for ventilatory support. Neurophysiological findings were classified as demyelinating, axonal, equivocal, unexcitable, or normal. Predictive values of clinical and electrophysiological data were tested using classification trees to build up a predictive model. This model was initially built up in a two-third (fitting set) then validated in a one-third (validation set) of the total sample. The fitting and validation sets were randomly selected. We also assessed the predictive value of this model for disability at 6 months. FINDINGS: From 1998, to 2006, 154 patients with Guillain-Barré syndrome were included in the study and 34 (22%) were subsequently ventilated. Demyelinating Guillain-Barré syndrome was more common in patients who went on to be ventilated than in those who were not (85%vs 51%, p=0.0003). Vital capacity and the proximal/distal compound muscular amplitude potential (p/dCMAP) ratio of the common peroneal nerve were retained in the tree model, with a probability of needing ventilation of less than 2.5% in patients with a ratio of greater than 55.6% and a vital capacity more than 81% of predicted. A p/dCMAP ratio of the peroneal nerve less than 55.6% and age older than 40 years were retained as independent predictors of disability at 6 months. INTERPRETATION: Neurophysiological testing is helpful for assessing risk of respiratory failure, which is highest in patients with evidence of demyelination and very low in those without both 55.6% conduction block of the common peroneal nerve and a 20% reduction in vital capacity.

Action Potentials↗

Obtaining maximal concatenated phylogenetic data sets from large sequence databases.

To improve the accuracy of tree reconstruction, phylogeneticists are extracting increasingly large multigene data sets from sequence databases. Determining whether a database contains at least k genes sampled from at least m species is an NP-complete problem. However, the skewed distribution of sequences in these databases permits all such data sets to be obtained in reasonable computing times even for large numbers of sequences. We developed an exact algorithm for obtaining the largest multigene data sets from a collection of sequences. The algorithm was then tested on a set of 100,000 protein sequences of green plants and used to identify the largest multigene ortholog data sets having at least 3 genes and 6 species. The distribution of sizes of these data sets forms a hollow curve, and the largest are surprisingly small, ranging from 62 genes by 6 species, to 3 genes by 65 species, with more symmetrical data sets of around 15 taxa by 15 genes. These upper bounds to sequence concatenation have important implications for building the tree of life from large sequence databases.

Algorithms↗

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis↗

Evolutionary relationships of a "primitive" shark (Heterodontus) assessed by micro-complement fixation of serum transferrin.

The evolutionary relationships of six sharks were investigated by comparing their transferrins using the micro-complement fixation method. The immunological distances observed were used to build a tree that confirms that the squaloid and galeoid species examined belong to two separate groups and that Heterodontus, a genus of hitherto uncertain position, belongs with the galeoids. The divergence time estimated from the transferrin comparisons is roughly 240 +/- 65 million years between Heterodontus and galeoids.

Animals↗

Structure of the 16 S ribosomal RNA of the thermophilic cyanobacterium Chlorogloeopsis HTF ('Mastigocladus laminosus HTF') strain PCC7518, and phylogenetic analysis.

The thermophilic cyanobacterial strain, PCC7518, originally identified as 'Mastigocladus laminosus HTF' does not show branchings or heterocysts. The absence of branchings supports the later assignment to the genus Chlorogloeopsis. The absence of heterocysts may be the result of a mutation because heterocysts were observed in the original isolate. Alternatively, contamination may have happened. To solve this problem, the 16 S rRNA sequence was determined and used to infer a secondary structure model and build distance trees. The trees showed that strain PCC7518 belongs to the cluster of heterocystous species and has most probably lost the ability to produce heterocysts by mutation. It is only distantly related to Chlorogloeopsis fritschii PCC6718.

Base Sequence↗

Three-dimensional reconstruction of intracranial vessels from biplane projection views.

The three-dimensional (3D) reconstruction of intracerebral vessels from two-dimensional (2D) projection views is an important clinical problem that, so far, has eluded solution. This report describes a new approach that uses projection images to build arterial trees progressively from an underlying 3D network, using a new method to pair shadow images on widely separated projection views. As a test of our general methodology, we have reconstructed a middle cerebral arterial tree from two projection views of a magnetic resonance dataset and have tested the accuracy of reconstruction against the original 3D dataset. This report describes the general approach to 3D vascular reconstruction and the computer program used to perform the final reconstruction step. The results suggest that accurate, 3D reconstruction of intracranial vessels is indeed possible from as few as two projection views.

Cerebral Arteries↗

Fast and high precision algorithms for optimization in large-scale genomic problems.

There are several very difficult problems related to genetic or genomic analysis that belong to the field of discrete optimization in a set of all possible orders. With n elements (points, markers, clones, sequences, etc.), the number of all possible orders is n!/2 and only one of these is considered to be the true order. A classical formulation of a similar mathematical problem is the well-known traveling salesperson problem model (TSP). Genetic analogues of this problem include: ordering in multilocus genetic mapping, evolutionary tree reconstruction, building physical maps (contig assembling for overlapping clones and radiation hybrid mapping), and others. A novel, fast and reliable hybrid algorithm based on evolution strategy and guided local search discrete optimization was developed for TSP formulation of the multilocus mapping problems. High performance and high precision of the employed algorithm named guided evolution strategy (GES) allows verification of the obtained multilocus orders based on different computing-intensive approaches (e.g., bootstrap or jackknife) for detection and removing unreliable marker loci, hence, stabilizing the resulting paths. The efficiency of the proposed algorithm is demonstrated on standard TSP problems and on simulated data of multilocus genetic maps up to 1000 points per linkage group.

Algorithms↗

Domain rearrangements in protein evolution.

Most eukaryotic proteins are multi-domain proteins that are created from fusions of genes, deletions and internal repetitions. An investigation of such evolutionary events requires a method to find the domain architecture from which each protein originates. Therefore, we defined a novel measure, domain distance, which is calculated as the number of domains that differ between two domain architectures. Using this measure the evolutionary events that distinguish a protein from its closest ancestor have been studied and it was found that indels are more common than internal repetition and that the exchange of a domain is rare. Indels and repetitions are common at both the N and C-terminals while they are rare between domains. The evolution of the majority of multi-domain proteins can be explained by the stepwise insertions of single domains, with the exception of repeats that sometimes are duplicated several domains in tandem. We show that domain distances agree with sequence similarity and semantic similarity based on gene ontology annotations. In addition, we demonstrate the use of the domain distance measure to build evolutionary trees. Finally, the evolution of multi-domain proteins is exemplified by a closer study of the evolution of two protein families, non-receptor tyrosine kinases and RhoGEFs.

Databases, Protein↗