PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Phylogenetic analysis of the Monocotylidae (Monogenea) inferred from 28S rDNA sequences.

The current classification of the Monocotylidae (Monogenea) is based on a phylogeny generated from morphological characters. The present study tests the morphological phylogenetic hypothesis using molecular methods. Sequences from domains C2 and D1 and the partial domains C1 and D2 from the 28S rDNA gene for 26 species of monocotylids from six of the seven subfamilies were used. Trees were generated using maximum parsimony, neighbour joining and maximum likelihood algorithms. The maximum parsimony tree, with branches showing less than 70% bootstrap support collapsed, had a topology identical to that obtained using the maximum likelihood analysis. The neighbour joining tree, with branches showing less than 70% support collapsed, differed only in its placement of Heterocotyle capricornensis as the sister group to the Decacotylinae clade. The molecular tree largely supports the subfamilies established using morphological characters. Differences are primarily how the subfamilies are related to each other. The monophyly of the Calicotylinae and Merizocotylinae and their sister group relationship is supported by high bootstrap values in all three methods, but relationships within the Merizocotylinae are unclear. Merizocotyle is paraphyletic and our data suggest that Mycteronastes and Thaumatocotyle, which were synonymized with Merizocotyle after the morphological cladistic analysis, should perhaps be resurrected as valid genera. The monophyly of the Monocotylinae and Decacotylinae is also supported by high bootstrap values. The Decacotylinae, which was considered previously to be the sister group to the Calicotylinae plus Merizocotylinae, is grouped in an unresolved polychotomy with the Monocotylinae and members of the Heterocotylinae. According to our molecular data, the Heterocotylinae is paraphyletic. Molecular data support a sister group relationship between Troglocephalus rhinobatidis and Neoheterocotyle rhinobatidis to the exclusion of the other species of Neoheterocotyle and recognition of Troglocephalus renders Neoheterocotyle paraphyletic. We propose Troglocephalus incertae sedis. An updated classification and full species list of the Monocotylidae is provided.

Animals↗

Phylogenetic analysis of the Monocotylidae (Monogenea) inferred from 28S rDNA sequences.

The current classification of the Monocotylidae (Monogenea) is based on a phylogeny generated from morphological characters. The present study tests the morphological phylogenetic hypothesis using molecular methods. Sequences from domains C2 and D1 and the partial domains C1 and D2 from the 28S rDNA gene for 26 species of monocotylids from six of the seven subfamilies were used. Trees were generated using maximum parsimony, neighbour joining and maximum likelihood algorithms. The maximum parsimony tree, with branches showing less than 70% bootstrap support collapsed, had a topology identical to that obtained using the maximum likelihood analysis. The neighbour joining tree, with branches showing less than 70% support collapsed, differed only in its placement of Heterocotyle capricornensis as the sister group to the Decacotylinae clade. The molecular tree largely supports the subfamilies established using morphological characters. Differences are primarily how the subfamilies are related to each other. The monophyly of the Calicotylinae and Merizocotylinae and their sister group relationship is supported by high bootstrap values in all three methods, but relationships within the Merizocotylinae are unclear. Merizocotyle is paraphyletic and our data suggest that Mycteronastes and Thaumatocotyle, which were synonymized with Merizocotyle after the morphological cladistic analysis, should perhaps be resurrected as valid genera. The monophyly of the Monocotylinae and Decacotylinae is also supported by high bootstrap values. The Decacotylinae, which was considered previously to be the sister group to the Calicotylinae plus Merizocotylinae, is grouped in an unresolved polychotomy with the Monocotylinae and members of the Heterocotylinae. According to our molecular data, the Heterocotylinae is paraphyletic. Molecular data support a sister group relationship between Troglocephalus rhinobatidis and Neoheterocotyle rhinobatidis to the exclusion of the other species of Neoheterocotyle and recognition of Troglocephalus renders Neoheterocotyle paraphyletic. We propose Troglocephalus incertae sedis. An updated classification and full species list of the Monocotylidae is provided.

Animals↗

Dealing with skewed data: an example using asthma-related costs of medicaid clients.

BACKGROUND: Cost data often are nonnormally distributed due to a few very high cost values that may not necessarily be dismissed as outliers. Researchers have not reached agreement on how to appropriately deal with skewed cost data. OBJECTIVES: This study presents an example of skewed cost data that were collected retrospectively from the Texas Medicaid database. Common methods of dealing with skewed cost distributions are discussed. Data were analyzed using various methods, and the statistical results of each test were compared. METHODS: Prescription and medical claims data extracted from the Texas Medicaid database were analyzed using the Mann-Whitney U test and t tests of untransformed, log-transformed, and bootstrapped data. RESULTS: All distributions of the untransformed cost data were nonnormally distributed, and comparison groups had unequal variances. The Mann-Whitney U test negated the effect of the high-cost patients and gave a significant result for overall cost differences between groups, but in the opposite direction of the mean. The t tests on raw data and log-transformed data may not have been optimal because distributions of both raw costs and log-costs were nonnormal. CONCLUSIONS: The bootstrap method does not need to meet the assumptions of normality and equal variances. In analyses of small sample sizes with skewed cost data, the bootstrap method may offer an alternative to the more traditional nonparametric or log-transformation techniques.

Asthma↗

Prognostic factor studies in oncology: osteosarcoma as a clinical example.

OBJECTIVE: Prognostic factor studies are abundant in oncology. Nevertheless, most of them have very limited impact on clinical practice, in part because many of them have a low statistical power. The importance of statistical power is illustrated using bootstrap resampling of data from a series of osteosarcoma patients. METHODS AND MATERIALS: Osteosarcoma is a rare disease, the incidence being just a few cases per million person-years in the Western World. Very few Phase III studies have been conducted in the disease, and much knowledge on therapeutic progress has come from Phase II studies. This has caused controversy concerning the validity of historical controls, which again has stimulated interest in the identification of prognostic factors in this disease. A literature search in the National Library of Medicine MEDLINE database was performed to identify prognostic factor studies in osteosarcoma published between 1975 and 1998. Monte Carlo methods, so-called bootstrap resampling, are used to investigate the importance of sample size based on survival data for a previously published series of 158 osteosarcomas treated with surgery alone. RESULTS: Most published studies are too small to provide useful information on prognostic factors in osteosarcoma. Three-quarters of the papers reviewed included less than 100 patients with osteosarcoma. More than 20 different potential prognostic factors were included in these papers. Inherent differences between studies and poor reporting hamper a synthesis of information from various studies. The results from the bootstrap resampling illustrate how the majority of published studies would miss even quite significant prognostic factors. CONCLUSIONS: An effort is needed to improve the design, conduct, and reporting of prognostic studies in oncology.

Bone Neoplasms↗

Clustering ensembles of neural network models.

We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models' outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. A parallel application is the prediction of newspaper sales, where we learn a series of parallel tasks. The results indicate that it is not necessary to store all samples in the ensembles: a small number of representative models generally matches, or even surpasses, the performance of the full ensemble. The clustered representation of the ensemble obtained thus is much better suitable for qualitative analysis, and will be shown to yield new insights into the data.

Algorithms↗

Internal validation of predictive models: efficiency of some procedures for logistic regression analysis.

The performance of a predictive model is overestimated when simply determined on the sample of subjects that was used to construct the model. Several internal validation methods are available that aim to provide a more accurate estimate of model performance in new subjects. We evaluated several variants of split-sample, cross-validation and bootstrapping methods with a logistic regression model that included eight predictors for 30-day mortality after an acute myocardial infarction. Random samples with a size between n = 572 and n = 9165 were drawn from a large data set (GUSTO-I; n = 40,830; 2851 deaths) to reflect modeling in data sets with between 5 and 80 events per variable. Independent performance was determined on the remaining subjects. Performance measures included discriminative ability, calibration and overall accuracy. We found that split-sample analyses gave overly pessimistic estimates of performance, with large variability. Cross-validation on 10% of the sample had low bias and low variability, but was not suitable for all performance measures. Internal validity could best be estimated with bootstrapping, which provided stable estimates with low bias. We conclude that split-sample validation is inefficient, and recommend bootstrapping for estimation of internal validity of a predictive logistic regression model.

Aged↗

Impact of number of isoenzyme loci on the robustness of intraspecific phylogenies using multilocus enzyme electrophoresis: consequences for typing of Trypanosoma cruzi.

Thirty-one stocks of Trypanosoma cruzi, the agent of Chagas disease, representative of the genetic variability of the 2 principal lineages, that subdivide T. cruzi, were selected on the basis of previous multilocus enzyme electrophoresis analysis using 21 loci. Analyses were performed with lower numbers of loci to explore the impact of the number of loci on the robustness of the phylogenies obtained, and to identify the loci that have more impact on the phylogeny. Analyses were performed with numerical (UPGMA) and cladistical (Wagner parsimony analysis) methods for all sets of loci. Robustness of the phylogenies obtained was estimated by bootstrap analysis. Low numbers of randomly selected loci (6) were sufficient to demonstrate genetic heterogeneity among the stocks studied. However, they were unable to give reliable phylogenetic information. A higher number of randomly selected loci (15 and more) were required to reach this goal. All loci did not convey equivalent information. The more variable loci detected a greater genetic heterogeneity among the stocks, whereas the least variable loci were better for robust clustering. Finally, analysis was performed with only 5 and 9 loci bearing synapomorphic allozyme characters previously identified among larger samples of stocks. A set of 9 such loci was able to uncover both genetic heterogeneity among the stocks and to build robust phylogenies. It can therefore be recommended as a minimum set of isoenzyme loci that bring maximal information for all studies aiming to explore the phylogenetic diversity of a new set of T. cruzi stocks and for any preliminary genetic typing. Moreover, our results show that bootstrap analysis, like any statistics, is highly dependent upon the information available and that absolute bootstrap figures should be cautiously interpreted.

Animals↗

Cost-effectiveness of a preventive counseling and support package for postnatal depression.

OBJECTIVES: This study reports the cost-effectiveness of a preventive intervention, consisting of counseling and specific support for the mother-infant relationship, targeted at women at high risk of developing postnatal depression. METHODS: A prospective economic evaluation was conducted alongside a pragmatic randomized controlled trial in which women considered at high risk of developing postnatal depression were allocated randomly to the preventive intervention (n = 74) or to routine primary care (n = 77). The primary outcome measure was the duration of postnatal depression experienced during the first 18 months postpartum. Data on health and social care use by women and their infants up to 18 months postpartum were collected, using a combination of prospective diaries and face-to-face interviews, and then were combined with unit costs ( pound, year 2000 prices) to obtain a net cost per mother-infant dyad. The nonparametric bootstrap method was used to present cost-effectiveness acceptability curves and net benefit statistics at alternative willingness to pay thresholds held by decision makers for preventing 1 month of postnatal depression. RESULTS: Women in the preventive intervention group were depressed for an average of 2.21 months (9.57 weeks) during the study period, whereas women in the routine primary care group were depressed for an average of 2.70 months (11.71 weeks). The mean health and social care costs were estimated at pounds sterling 2,396.9 per mother-infant dyad in the preventive intervention group and pounds sterling 2,277.5 per mother-infant dyad in the routine primary care group, providing a mean cost difference of pounds sterling 119.5 (bootstrap 95 percent confidence interval [CI], -535.4, 784.9). At a willingness to pay threshold of pounds sterling 1,000 per month of postnatal depression avoided, the probability that the preventive intervention is cost-effective is .71 and the mean net benefit is pounds sterling 383.4 (bootstrap 95 percent CI, - pounds sterling 863.3- pounds sterling 1,581.5). CONCLUSIONS: The preventive intervention is likely to be cost-effective even at relatively low willingness to pay thresholds for preventing 1 month of postnatal depression during the first 18 months postpartum. Given the negative impact of postnatal depression on later child development, further research is required that investigates the longer-term cost-effectiveness of the preventive intervention in high risk women.

Cost-Benefit Analysis↗

Quantification of variability and uncertainty for air toxic emission inventories with censored emission factor data.

Probabilistic emission inventories were developed for urban air toxic emissions of benzene, formaldehyde, chromium, and arsenic for the example of Houston. Variability and uncertainty in emission factors were quantified for 71-97% of total emissions, depending upon the pollutant and data availability. Parametric distributions for interunit variability were fit using maximum likelihood estimation (MLE), and uncertainty in mean emission factors was estimated using parametric bootstrap simulation. For data sets containing one or more nondetected values, empirical bootstrap simulation was used to randomly sample detection limits for nondetected values and observations for sample values, and parametric distributions for variability were fit using MLE estimators for censored data. The goodness-of-fit for censored data was evaluated by comparison of cumulative distributions of bootstrap confidence intervals and empirical data. The emission inventory 95% uncertainty ranges are as small as -25% to +42% for chromium to as large as -75% to +224% for arsenic with correlated surrogates. Uncertainty was dominated by only a few source categories. Recommendations are made for future improvements to the analysis.

Air Pollutants↗

Estimating risk assessment exposure point concentrations when the data are not normal or lognormal.

The U.S. Environmental Protection Agency (EPA) recommends the use of the one-sided 95% upper confidence limit of the arithmetic mean based on either a normal or lognormal distribution for the contaminant (or exposure point) concentration term in the Superfund risk assessment process. When the data are not normal or lognormal this recommended approach may overestimate the exposure point concentration (EPC) and may lead to unecessary cleanup at a hazardous waste site. The EPA concentration term only seems to perform like alternative EPC methods when the data are well fit by a lognormal distribution. Several alternative methods for calculating the EPC are investigated and compared using soil data collected from three hazardous waste sites in Montana, Utah, and Colorado. For data sets that are well fit by a lognormal distribution, values for the Chebychev inequality or the EPA concentration term may be appropriate EPCs. For data sets where the soil concentration data are well fit by gamma distributions, Wong's method may be used for calculating EPCs. The studentized bootstrap-t and Hall's bootstrap-t transformation are recommended for EPC calculation when all distribution fits are poor. If a data set is well fit by a distribution, parametric bootstrap may provide a suitable EPC.

Evaluation Studies as Topic↗

Determining the number of clusters by sampling with replacement.

A split-sample replication criterion originally proposed by J. E. Overall and K. N. Magee (1992) as a stopping rule for hierarchical cluster analysis is applied to multiple data sets generated by sampling with replacement from an original simulated primary data set. An investigation of the validity of this bootstrap procedure was undertaken using different combinations of the true number of latent populations, degrees of overlap, and sample sizes. The bootstrap procedure enhanced the accuracy of identifying the true number of latent populations under virtually all conditions. Increasing the size of the resampled data sets relative to the size of the primary data set further increased accuracy. A computer program to implement the bootstrap stopping rule is made available via a referenced Web site.

Cluster Analysis↗

The root of the angiosperms revisited.

Most recent phylogenetic analyses of basal angiosperms have converged on the placement of Amborella as sister to all other extant angiosperms. However, certain recent studies suggest that Amborella and Nymphaeales (water lilies) form a clade sister to all remaining angiosperms or that Nymphaeales alone are the sister to the remaining angiosperms. We report here (i) maximum parsimony, maximum likelihood, and Bayesian phylogenetic analyses of 11 genes (>15,000 bp per taxon) for 16 taxa, (ii) maximum parsimony analysis for a subset of these genes for 104 taxa, and (iii) tests of alternative rootings with the nonparametric bootstrap and the likelihood ratio test with the parametric bootstrap. In addition, we use simulation analyses to examine the amount of bias that may be present in our methods of phylogeny estimation. Amborella continues to receive strong bootstrap support as the sister to all other extant angiosperms, and three of four tests reject alternative hypotheses of the angiosperm root. Although we cannot conclusively choose between Amborella vs. Amborella + Nymphaeales as sister to all other angiosperms, most analyses favor the former rooting.

Evolution, Molecular↗

Selection bias in gene extraction on the basis of microarray gene-expression data.

In the context of cancer diagnosis and treatment, we consider the problem of constructing an accurate prediction rule on the basis of a relatively small number of tumor tissue samples of known type containing the expression data on very many (possibly thousands) genes. Recently, results have been presented in the literature suggesting that it is possible to construct a prediction rule from only a few genes such that it has a negligible prediction error rate. However, in these results the test error or the leave-one-out cross-validated error is calculated without allowance for the selection bias. There is no allowance because the rule is either tested on tissue samples that were used in the first instance to select the genes being used in the rule or because the cross-validation of the rule is not external to the selection process; that is, gene selection is not performed in training the rule at each stage of the cross-validation process. We describe how in practice the selection bias can be assessed and corrected for by either performing a cross-validation or applying the bootstrap external to the selection process. We recommend using 10-fold rather than leave-one-out cross-validation, and concerning the bootstrap, we suggest using the so-called .632+ bootstrap error estimate designed to handle overfitted prediction rules. Using two published data sets, we demonstrate that when correction is made for the selection bias, the cross-validated error is no longer zero for a subset of only a few genes.

Discriminant Analysis↗

Evolution of translational elongation factor (EF) sequences: reliability of global phylogenies inferred from EF-1 alpha(Tu) and EF-2(G) proteins.

The EF-2 coding genes of the Archaea Pyrococcus woesei and Desulfurococcus mobilis were cloned and sequenced. Global phylogenies were inferred by alternative tree-making methods from available EF-2(G) sequence data and contrasted with phylogenies constructed from the more conserved but shorter EF-1 alpha(Tu) sequences. Both the monophyly (sensu Hennig) of Archaea and their subdivision into the kingdoms Crenarchaeota and Euryarchaeota are consistently inferred by analysis of EF-2(G) sequences, usually at a high bootstrap confidence level. In contrast, EF-1 alpha(Tu) phylogenies tend to be inconsistent with one another and show low bootstrap confidence levels. While evolutionary distance and DNA maximum parsimony analyses of EF-1 alpha(Tu) sequences do show archaeal monophyly, protein parsimony and DNA maximum-likelihood analyses of these data do not. In no case, however, do any of the tree topologies inferred from EF-1 alpha(Tu) sequence analyses receive significant bootstrap support.

Amino Acid Sequence↗

An approximately unbiased test of phylogenetic tree selection.

An approximately unbiased (AU) test that uses a newly devised multiscale bootstrap technique was developed for general hypothesis testing of regions in an attempt to reduce test bias. It was applied to maximum-likelihood tree selection for obtaining the confidence set of trees. The AU test is based on the theory of Efron et al. (Proc. Natl. Acad. Sci. USA 93:13429-13434; 1996), but the new method provides higher-order accuracy yet simpler implementation. The AU test, like the Shimodaira-Hasegawa (SH) test, adjusts the selection bias overlooked in the standard use of the bootstrap probability and Kishino-Hasegawa tests. The selection bias comes from comparing many trees at the same time and often leads to overconfidence in the wrong trees. The SH test, though safe to use, may exhibit another type of bias such that it appears conservative. Here I show that the AU test is less biased than other methods in typical cases of tree selection. These points are illustrated in a simulation study as well as in the analysis of mammalian mitochondrial protein sequences. The theoretical argument provides a simple formula that covers the bootstrap probability test, the Kishino-Hasegawa test, the AU test, and the Zharkikh-Li test. A practical suggestion is provided as to which test should be used under particular circumstances.

Animals↗

Molecular systematics and adaptive radiation of Hawaii's endemic Damselfly genus Megalagrion (Odonata: Coenagrionidae).

Damselflies of the endemic Hawaiian genus Megalagrion have radiated into a wide variety of habitats and are an excellent model group for the study of adaptive radiation. Past phylogenetic analysis based on morphological characters has been problematic. Here, we examine relationships among 56 individuals from 20 of the 23 described species using maximum likelihood (ML) and Bayesian phylogenetic analysis of mitochondrial (1287 bp) and nuclear (1039 bp) DNA sequence data. Models of evolution were chosen using the Akaike information criterion. Problems with distant outgroups were accommodated by constraining the best ML ingroup topology but allowing the outgroups to attach to any ingroup branch in a bootstrap analysis. No strong contradictions were obtained between either data partition and the combined data set. Areas of disagreement are mainly confined to clades that are strongly supported by the mitochondrial DNA and weakly supported by the elongation factor 1alpha data because of lack of changes. However, the combined analysis resulted in a unique tree. Correlation between Bayesian posterior probabilities and bootstrap percentages decreased in concert with decreasing information in the data partitions. In cases where nodes were supported by single characters bootstrap proportions were dramatically reduced compared with posterior probabilities. Two speciation patterns were evident from the phylogenetic analysis. First, most speciation is interisland and occurred as members of established ecological guilds colonized new volcanoes after they emerged from the sea. Second, there are several instances of rapid radiation into a variety of specialized habitats, in one case entirely within the island of Kauai. Application of a local clock procedure to the mitochondrial DNA topology suggests that two of these radiations correspond to the development of habitat on the islands of Kauai and Oahu. About 4.0 million years ago, species simultaneously moved into fast streams and plant leaf axils on Kauai, and about 1.5 million years later another group moved simultaneously to seeps and terrestrial habitats on Oahu. Results from the local clock analysis also strongly suggest that Megalagrion arrived in Hawaii about 10 million years ago, well before the emergence of Kauai. Date estimates were more sensitive to the particular node that was fixed in time than to the model of local branch evolution used. We propose a general model for the development of endemic damselfly species on Hawaiian Islands and document five potential cases of hybridization (M. xanthomelas x M. pacificum, M. eudytum x M. vagabundum, M. orobates x M. oresitrophum, M. nesiotes x M. oahuense, and M. mauka x M. paludicola).

Adaptation, Biological↗

Branch-length prior influences Bayesian posterior probability of phylogeny.

The Bayesian method for estimating species phylogenies from molecular sequence data provides an attractive alternative to maximum likelihood with nonparametric bootstrap due to the easy interpretation of posterior probabilities for trees and to availability of efficient computational algorithms. However, for many data sets it produces extremely high posterior probabilities, sometimes for apparently incorrect clades. Here we use both computer simulation and empirical data analysis to examine the effect of the prior model for internal branch lengths. We found that posterior probabilities for trees and clades are sensitive to the prior for internal branch lengths, and priors assuming long internal branches cause high posterior probabilities for trees. In particular, uniform priors with high upper bounds bias Bayesian clade probabilities in favor of extreme values. We discuss possible remedies to the problem, including empirical and full Bayesian methods and subjective procedures suggested in Bayesian hypothesis testing. Our results also suggest that the bootstrap proportion and Bayesian posterior probability are different measures of accuracy, and that the bootstrap proportion, if interpreted as the probability that the clade is true, can be either too liberal or too conservative.

Bayes Theorem↗

Relevant and redundant lung function parameters in discriminating asthma from COPD.

A relevant set of lung function parameters, derived from spirometry, flow-volume curves, diffusion capacity and bodyplethysmography, to discriminate asthma from COPD was established via logistic regression analysis. All new patients, referred to the outpatient clinic and later defined as asthma or COPD, underwent extensive lung function testing with reversibility testing. Logistic regression was used to calculate the probability to be a COPD or asthma patient. Selection of relevant parameters was done via 1] forward-, 2] backward-, 3] stepwise selection and 4] the best score method. All four methods were supplemented by bootstrapping to obtain a validated selection and estimation of the logistic regression parameters. The area under the ROC curve (mean+/-sd) for respectively the forward, backward, stepwise and best score selection method is 0.9348+/-0.0115, 0.9346+/-0.0115, 0.9348+/-0.0115 and 0.9296+/-0.0121. The TLCO, VA and the postdilator MEF50, VC and PEF were selected as the most relevant parameters in discriminating asthma from COPD: they appeared most often as relevant discriminators in 500 bootstrapped samples: TLco was present in all bootstrapped samples and VA, postdilator MEF50, VC and PEF in resp. 70.8%, 46.2%, 42.8% and 36.8%. Bodyplethysmography derived parameters turned out to be of limited value. Diffusion capacity testing and spirometry/flow-volume curve after administration of bronchodilators are the methods of choicewhen having to chose between asthma or COPD.

Adult↗