PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Bootstrap confidence levels for phylogenetic trees.

Evolutionary trees are often estimated from DNA or RNA sequence data. How much confidence should we have in the estimated trees? In 1985, Felsenstein [Felsenstein, J. (1985) Evolution 39, 783-791] suggested the use of the bootstrap to answer this question. Felsenstein's method, which in concept is a straightforward application of the bootstrap, is widely used, but has been criticized as biased in the genetics literature. This paper concerns the use of the bootstrap in the tree problem. We show that Felsenstein's method is not biased, but that it can be corrected to better agree with standard ideas of confidence levels and hypothesis testing. These corrections can be made by using the more elaborate bootstrap method presented here, at the expense of considerably more computation.

Algorithms↗

Bootstrap confidence levels for phylogenetic trees.

Evolutionary trees are often estimated from DNA or RNA sequence data. How much confidence should we have in the estimated trees? In 1985, Felsenstein [Felsenstein, J. (1985) Evolution 39, 783-791] suggested the use of the bootstrap to answer this question. Felsenstein's method, which in concept is a straightforward application of the bootstrap, is widely used, but has been criticized as biased in the genetics literature. This paper concerns the use of the bootstrap in the tree problem. We show that Felsenstein's method is not biased, but that it can be corrected to better agree with standard ideas of confidence levels and hypothesis testing. These corrections can be made by using the more elaborate bootstrap method presented here, at the expense of considerably more computation.

Animals↗

Stability of questionnaire items in sport and exercise psychology: bootstrap limits of agreement.

We describe an alternative method for assessing the stability of individual questionnaire items in sport and exercise psychology. To date, the Pearson product-moment correlation coefficient has been widely used in psychometrics. We propose an alternative non-parametric method based on proportion of agreement. Ninety-two male university students completed the revised 9-item Social Physique Anxiety Scale on two occasions, separated by a 2-week interval. Point estimates of the proportion of direct within-individual agreement between the two occasions were calculated separately for each item of the Social Physique Anxiety Scale. Estimates of uncertainty of the agreement were calculated using a bootstrapping resampling technique. For each item, 2000 bootstrap samples (each n = 92 pairs) were redrawn from the original sample. The sample statistic was calculated for each bootstrap sample to provide a bootstrap sampling distribution. The 95% confidence intervals (CI) were then calculated using the percentile method. The three most problematic items were items 7, 8 and 10 (as labelled in the original 12-item scale). These items demonstrated an agreement of 0.46 (95% CI= 0.36 0.56), 0.42 (95% CI = 0.33-0.52) and 0.41 (95% CI = 0.32-0.51) respectively. Our proposed method measures absolute agreement between test-retest responses, is free of normal assumptions, does not depend on high between-individuals variance, and can be applied successfully to individual items in the development of psychological tests.

Adult↗

Reliability of Bayesian posterior probabilities and bootstrap frequencies in phylogenetics.

Many empirical studies have revealed considerable differences between nonparametric bootstrapping and Bayesian posterior probabilities in terms of the support values for branches, despite claimed predictions about their approximate equivalence. We investigated this problem by simulating data, which were then analyzed by maximum likelihood bootstrapping and Bayesian phylogenetic analysis using identical models and reoptimization of parameter values. We show that Bayesian posterior probabilities are significantly higher than corresponding nonparametric bootstrap frequencies for true clades, but also that erroneous conclusions will be made more often. These errors are strongly accentuated when the models used for analyses are underparameterized. When data are analyzed under the correct model, nonparametric bootstrapping is conservative. Bayesian posterior probabilities are also conservative in this respect, but less so.

Bayes Theorem↗

Sampling properties of the bootstrap support in molecular phylogeny: influence of nonindependence among sites.

The influence of nonindependence among sites on phylogenetic reconstructions and bootstrap scores was investigated both analytically and empirically. First, the sampling properties of the bootstrap support in the four-species case was derived for the maximum-parsimony method, assuming either independently or nonindependently evolving sites. The influence of various models of departure from the independence assumption was quantified. Second, trees and bootstrap scores estimated from subsets of consecutive (potentially coevolving) versus dispersed (presumably independent) sites of a ribosomal RNA data set were contrasted. The two approaches consistently suggest that a departure from the assumption of independent sites tends to reduce the amount of phylogenetic information contained in the data, but to increase the apparent statistical support for reconstructed trees, as measured by the bootstrap. In particular, nonindependence can lead to strongly supported wrong internal branches.

Classification↗

Increasing data transparency and estimating phylogenetic uncertainty in supertrees: Approaches using nonparametric bootstrapping.

The estimation of ever larger phylogenies requires consideration of alternative inference strategies, including divide-and-conquer approaches that decompose the global inference problem to a set of smaller, more manageable component problems. A prominent locus of research in this area is the development of supertree methods, which estimate a composite tree by combining a set of partially overlapping component topologies. Although promising, the use of component tree topologies as the primary data dissociates supertrees from complexities within the underling character data and complicates the evaluation of phylogenetic uncertainty. We address these issues by exploring three approaches that variously incorporate nonparametric bootstrapping into a common supertree estimation algorithm (matrix representation with parsimony, although any algorithm might be used), including bootstrap-weighting, source-tree bootstrapping, and hierarchical bootstrapping. We illustrate these procedures by means of hypothetical and empirical examples. Our preliminary experiments suggest that these methods have the potential to improve the correspondence of supertree estimates to those derived from simultaneous analysis of the combined data and to allow uncertainty in supertree topologies to be quantified. The ability to increase the transparency of supertrees to the underlying character data has several practical implications and sheds new light on an old debate. These methods have been implemented in the freely available program, tREeBOOT.

Classification↗

Culling avalanches in bootstrap percolation.

We study the culling avalanches which occur after the "death" of a single randomly chosen site in a network where sites are unstable, and are culled, if they have coordination less than an integer parameter m. Avalanche distributions are presented for triangular and cubic lattices for values of m where the associated bootstrap transitions are either first or second order. In second order cases, the culling avalanche distribution is found to be exponential, while in first order cases it follows a power law. We present an exact relation between culling avalanches and conventional bootstrap percolation and show that a relation proposed by Manna [Physica A 261, 351 (1998)] can be a good approximation for strongly first order bootstrap transitions but not for continuous bootstrap transitions.

Journal Article↗

Minimizing model fitting objectives that contain spurious local minima by bootstrap restarting.

Objective functions that arise when fitting nonlinear models often contain local minima that are of little significance except for their propensity to trap minimization algorithms. The standard methods for attempting to deal with this problem treat the objective function as fixed and employ stochastic minimization approaches in the hope of randomly jumping out of local minima. This article suggests a simple trick for performing such minimizations that can be employed in conjunction with most conventional nonstochastic fitting methods. The trick is to stochastically perturb the objective function by bootstrapping the data to be fit. Each bootstrap objective shares the large-scale structure of the original objective but has different small-scale structure. Minimizations of bootstrap objective functions are alternated with minimizations of the original objective function starting from the parameter values with which minimization of the previous bootstrap objective terminated. An example is presented, fitting a nonlinear population dynamic model to population dynamic data and including a comparison of the suggested method with simulated annealing. Convergence diagnostics are discussed.

Algorithms↗

Locus-specific heritability estimation via the bootstrap in linkage scans for quantitative trait loci.

BACKGROUND/AIMS: In genome-wide linkage analysis of quantitative trait loci (QTL), locus-specific heritability estimates are biased when the original data are used to both localize linkage and estimate effects, due to maximization of the LOD score over the genome. Positive bias is increased by adoption of stringent significance levels to control genome-wide type I error. We propose multi-locus bootstrap resampling estimators for bias reduction in the situation in which linkage peaks at more than one QTL are of interest. METHODS: Bootstrap estimates were based on repeated sample splitting in the original dataset. We conducted simulation studies in nuclear families with 0 to 5 QTLs and applied the methods in a genome-wide analysis of a blood pressure phenotype in extended pedigrees from the Framingham Heart Study (FHS). RESULTS: Compared to naïve estimates in the original simulation samples, bootstrap estimates had reduced bias and smaller mean squared error. In the FHS pedigrees, the bootstrap yielded heritability estimates as much as 70% smaller than in the original sample. CONCLUSIONS: Because effect estimates obtained in an initial study are typically inflated relative to those expected in an independent replication study, successful replication will be more likely when sample size requirements are based on bias-reduced estimates.

Cardiovascular Diseases↗

Uncertainty in decision models analyzing cost-effectiveness: the joint distribution of incremental costs and effectiveness evaluated with a nonparametric bootstrap method.

PURPOSE: To illustrate the use of a nonparametric bootstrap method in the evaluation of uncertainty in decision models analyzing cost-effectiveness. METHODS: The authors reevaluated a previously published cost-effectiveness analysis that used a Markov model comparing initial percutaneous transluminal angioplasty with bypass surgery for femoropopliteal lesions. Each probability in the model was simulated with a first-order Monte Carlo simulation to represent sampling uncertainty. Superimposed on this, a second-order Monte Carlo simulation was performed to represent parameter uncertainty, drawing the probability values from nonparametric distributions based on published data or from primary collected data as available. After simulation of a mixed (i.e., non-identical) cohort of 30,000 patients, 3,000 bootstrap samples of 1,000 patients each were drawn and the joint distribution of mean incremental costs and mean effectiveness gained was evaluated. RESULTS: Using a bootstrap sample size of 1,000 patients, 92.7% of the joint distribution of mean incremental costs and mean effectiveness gained fell in the quadrant where angioplasty dominated bypass surgery. Another 6.9% of samples demonstrated either greater effectiveness with an incremental cost-effectiveness ratio of at most $20,000/QALY gained, or cost savings with a ratio of at least $20,000 saved/QALY lost. CONCLUSION: A nonparametric bootstrap method can be used to estimate the joint distribution of mean incremental costs and mean effectiveness gained, and the results can provide an understanding of the uncertainty in a cost-effectiveness analysis based on a decision model.

Angioplasty, Balloon↗

The use of bootstrap methods for analysing Health-Related Quality of Life outcomes (particularly the SF-36).

Health-Related Quality of Life (HRQoL) measures are becoming increasingly used in clinical trials as primary outcome measures. Investigators are now asking statisticians for advice on how to analyse studies that have used HRQoL outcomes.HRQoL outcomes, like the SF-36, are usually measured on an ordinal scale. However, most investigators assume that there exists an underlying continuous latent variable that measures HRQoL, and that the actual measured outcomes (the ordered categories), reflect contiguous intervals along this continuum. The ordinal scaling of HRQoL measures means they tend to generate data that have discrete, bounded and skewed distributions. Thus, standard methods of analysis such as the t-test and linear regression that assume Normality and constant variance may not be appropriate. For this reason, conventional statistical advice would suggest that non-parametric methods be used to analyse HRQoL data. The bootstrap is one such computer intensive non-parametric method for analysing data. We used the bootstrap for hypothesis testing and the estimation of standard errors and confidence intervals for parameters, in four datasets (which illustrate the different aspects of study design). We then compared and contrasted the bootstrap with standard methods of analysing HRQoL outcomes. The standard methods included t-tests, linear regression, summary measures and General Linear Models.Overall, in the datasets we studied, using the SF-36 outcome, bootstrap methods produce results similar to conventional statistical methods. This is likely because the t-test and linear regression are robust to the violations of assumptions that HRQoL data are likely to cause (i.e. non-Normality). While particular to our datasets, these findings are likely to generalise to other HRQoL outcomes, which have discrete, bounded and skewed distributions. Future research with other HRQoL outcome measures, interventions and populations, is required to confirm this conclusion.

Arthritis, Rheumatoid↗

Bootstrap scree tests: a Monte Carlo simulation and applications to published data.

A non-parametric procedure for Cattell's scree test is proposed, using the bootstrap method. Bentler and Yuan developed parametric tests for the linear trend of scree eigenvalues in principal component analysis. The proposed method is for cases where parametric assumptions are not realistic. We define the break in the scree trend in several ways, based on linear slopes defined with two or three consecutive eigenvalues, or all eigenvalues after the k largest. The resulting scree test statistics are evaluated under various data conditions, among which Gorsuch and Nelson's bootstrap CNG performs best and is reasonably consistent and efficient under leptokurtic and skewed conditions. We also examine the bias-corrected and accelerated bootstrap method for these statistics, and the bias correction is found to be too unstable to be useful. Using seven published data sets which Bentler and Yuan analysed, we compare the bootstrap approach to the scree test with the parametric linear trend test.

Humans↗

Standard errors in covariance structure models: asymptotics versus bootstrap.

Commonly used formulae for standard error (SE) estimates in covariance structure analysis are derived under the assumption of a correctly specified model. In practice, a model is at best only an approximation to the real world. It is important to know whether the estimates of SEs as provided by standard software are consistent when a model is misspecified, and to understand why if not. Bootstrap procedures provide nonparametric estimates of SEs that automatically account for distribution violation. It is also necessary to know whether bootstrap estimates of SEs are consistent. This paper studies the relationship between the bootstrap estimates of SEs and those based on asymptotics. Examples are used to illustrate various versions of asymptotic variance-covariance matrices and their validity. Conditions for the consistency of the bootstrap estimates of SEs are identified and discussed. Numerical examples are provided to illustrate the relationship of different estimates of SEs and covariance matrices.

Adolescent↗

Bootstrap analyses of cost effectiveness in antidepressant pharmacotherapy.

In this study, we describe 'bootstrap' methodology for placing statistical confidence limits around an incremental cost effectiveness ratio (ICER). This approach was applied to a retrospective study of annual charges for patients undergoing pharmacotherapy for depression. We used MarketScanSM (service mark) data from 1990 to 1992, which includes medical and pharmacy claims for a privately insured group of employed individuals and their families in the US. Our primary effectiveness measure was the proportion of patients who remained stable on their initial antidepressant medication for at least 6 consecutive months. Our primary cost measure was the total annual charge incurred by patients taking the selective serotonin reuptake inhibitor fluoxetine, a tricyclic antidepressant or a heterocyclic antidepressant. On average, fluoxetine pharmacotherapy tended to decrease annual charges by $US16.48 per patient for each percentage increase in depressed patients remaining stable on initial pharmacotherapy for 6 months, resulting in a negative ICER point-estimate. However, the upper ICER confidence limit is positive, which means that fluoxetine treatment may possibly increase annual per patient charges. With 95% confidence, any such increase was no more than $US130 per patient for each percentage increase in patients remaining stable on initial pharmacotherapy for at least 6 months. One advantage of using a bootstrap approach to ICER analysis is that it does not require restrictive distributional assumptions about cost and outcome measures. Bootstrapping also yields a dramatic graphical display of the variability in cost and effectiveness outcomes that result when a study is literally 'redone' hundreds of times. This graphic also displays the ICER confidence interval as a 'wedge-shaped' region on the cost-effectiveness plane. In fact, bootstrapping is easier to explain and appreciate than the elaborate calculations and approximations otherwise involved in ICER estimation. Our discussion addresses key technical questions, such as the role of logarithmic transformation in symmetrising highly skewed cost distributions. We hope that our discussion contributes to a dialogue, leading ultimately to a consensus on analysis of ICERs.

Adult↗

Better bootstrap estimation of hazardous concentration thresholds for aquatic assemblages.

The introduction of species sensitivity distribution (SSD) approaches to ecological risk assessment offers the potential for a more transparent scientific basis for the derivation of predicted no-effect concentrations. However, conventional SSD methodologies have relied on standard distributions (e.g., log logistic, log normal) that are not necessarily based on sound ecological or statistical grounds. More recently, bootstrap resampling techniques that do not rely on distributional assumptions have been applied to the problem. Here we describe how a more advanced bootstrap methodology may be applied to derive better point estimates and confidence intervals for SSD estimates of safe environmental concentrations. Motivated by the fact that the true SSD may not fit any standard model category, we go on to consider a hybrid bootstrap regression approach. This can yield a substantially different estimate for the SSD when compared with both the basic bootstrap and the more frequently used parametric curve approaches. With increasing use of SSDs in ecological risk assessment, it is now imperative that the scientific community develops agreement over appropriate methods for their derivation.

Data Interpretation, Statistical↗

Bootstrapped ultrasound calibration.

This paper introduces an enhanced (bootstrapped) method for tracked ultrasound probe calibration. Prior to calibration, a position sensor is used to track an ultrasound probe in 3D space, while the US image is used to determine calibration target locations within the image. From this information, an estimate of the transformation matrix of the scan plane with respect to the position sensor is computed. While all prior calibration methods terminate at this phase, we use this initial calibration estimate to bootstrap an additional optimization of the transformation matrix on independent data to yield the minimum reconstruction error on calibration targets. The bootstrapped workflow makes use of a closed-form calibration solver and associated sensitivity analysis, allowing for rapid and robust convergence to an optimal calibration matrix. Bootstrapping demonstrates superior reconstruction accuracy.

Calibration↗

Regression with bounded outcome score: evaluation of power by bootstrap and simulation in a chronic myelogenous leukaemia clinical trial.

Evaluation of the treatment effect on cytogenetic ordered categorical response is considered in patients treated for chronic myelogenous leukaemia (CML) in a clinical trial initiated by the East German Group for Hematology and Oncology. A simulation model for the cytogenetic response (per cent of Philadelphia chromosome positive metaphases) serially measured in CML patients was constructed to describe roughly the sparse information available in medical literature. The model was used to construct a summary measure of response and to formulate the treatment effect as a regression with U-shape distributed ordered categorical data. Two simple models (vertical shift model and pooled conditional response model) were specifically designed to model the treatment effect 'observed' in a simulated 'pilot' data set. The powers were contrasted with the traditional proportional odds and binary models. The comparison was based both on repeated sampling from the simulated model and on bootstrap of 'given' pilot data set. We show that the specific models that address the treatment effect directly (as anticipated from pilot data) can gain in power as compared to the traditional proportional odds model when evaluated by bootstrap. However, the proportional odds model appears to be better with repeated sampling from the simulation model. To explain this discrepancy we generated 'pilot data sets' repeatedly from the simulation model and showed that the ordering of the bootstrap power estimates is unstable with reasonably complex models dependent on the random fall of the pilot data sets. This phenomenon clearly limits the usefulness of subtle modelling the form of the treatment difference observed in a small pilot data set.

Antineoplastic Combined Chemotherapy Protocols↗

Bootstrap confidence intervals for the sensitivity of a quantitative diagnostic test.

We examine bootstrap approaches to the analysis of the sensitivity of quantitative diagnostic test data. Methods exist for inference concerning the sensitivity of one or more tests for fixed levels of specificity, taking into account the variability in the sensitivity due to variability in the test values for normal subjects. However, parametric methods do not adequately account for error, particularly when the data are non-normally distributed, and non-parametric methods have low power. We implement bootstrap methods for confidence limits for the sensitivity of a test for a fixed specificity and demonstrate that under certain circumstances the bootstrap method gives more accurate confidence intervals than do other methods, while it performs at least as well as other methods in many standard situations.

Algorithms↗