PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A bootstrap test for the analysis of microarray experiments with a very small number of replications.

In microarray studies it is common that the number of replications (i.e. the sample size) is small and that the distribution of expression values differs from normality. In this situation, permutation and bootstrap tests may be appropriate for the identification of differentially expressed genes. However, unlike bootstrap tests, permutation tests are not suitable for very small sample sizes, such as three per group. A variety of different bootstrap tests exists. For example, it is possible to adjust the data to have a common mean before the bootstrap samples are drawn. For small significance levels, which can occur when a large number of genes is investigated, the original bootstrap test, as well as a bootstrap test suggested for the Behrens-Fisher problem, have no power in cases of very small sample sizes. In contrast, the modified test based on adjusted data is powerful. Using a Monte Carlo simulation study, we demonstrate that the difference in power can be huge. In addition, the different tests are illustrated using microarray data.

Algorithms↗

The psychometric function: II. Bootstrap-based confidence intervals and sampling.

The psychometric function relates an observer's performance to an independent variable, usually a physical quantity of an experimental stimulus. Even if a model is successfully fit to the data and its goodness of fit is acceptable, experimenters require an estimate of the variability of the parameters to assess whether differences across conditions are significant. Accurate estimates of variability are difficult to obtain, however, given the typically small size of psychophysical data sets: Traditional statistical techniques are only asymptotically correct and can be shown to be unreliable in some common situations. Here and in our companion paper (Wichmann & Hill, 2001), we suggest alternative statistical techniques based on Monte Carlo resampling methods. The present paper's principal topic is the estimation of the variability of fitted parameters and derived quantities, such as thresholds and slopes. First, we outline the basic bootstrap procedure and argue in favor of the parametric, as opposed to the nonparametric, bootstrap. Second, we describe how the bootstrap bridging assumption, on which the validity of the procedure depends, can be tested. Third, we show how one's choice of sampling scheme (the placement of sample points on the stimulus axis) strongly affects the reliability of bootstrap confidence intervals, and we make recommendations on how to sample the psychometric function efficiently. Fourth, we show that, under certain circumstances, the (arbitrary) choice of the distribution function can exert an unwanted influence on the size of the bootstrap confidence intervals obtained, and we make recommendations on how to avoid this influence. Finally, we introduce improved confidence intervals (bias corrected and accelerated) that improve on the parametric and percentile-based bootstrap confidence intervals previously used. Software implementing our methods is available.

Confidence Intervals↗

Interpretation of bootstrap values in phylogenetic analysis.

Bootstrap Analysis is a common tool in cladistics, and consequently many authors tend to believe that it could be close to a test of monophyly. In fact, it is only a procedure to calculate the redundancy of a certain character pattern among taxa. To demonstrate this, we set up a study with questionable data: Four skulls of great apes and humans were digitally photographed, and the pixels' brightness values were simply transformed to a one-zero-matrix, which was then used to calculate a Wagner tree with PHYLIP. As a rule, the higher the resolution of the photos is, the higher are the bootstrap values of supported taxa (and the lower are the bootstrap values of non-supported data). Redundancy of intertaxic information might indeed be an indicator of phylogenetic relationship, but can also be due to other reasons, like functional-adaptive needs in morphology, or semantic needs in a DNA-code. As a result, we tend to believe that high bootstrap values are actually less important than low ones. It is safer, based on a low bootstrap value, to claim that a certain taxon is not well supported by certain data. Therefore, we recommend discussions of low bootstrap values in future publications.

Animals↗

Bootstrap confidence intervals for cost-effectiveness ratios: some simulation results.

Recently, a number of papers have brought up the issue of how to make cost-effectiveness (CE) studies stochastic, i.e. how to obtain confidence intervals for CE ratios. In this note we present a bootstrap procedure for estimating bias-corrected confidence intervals for CE ratios. The bootstrap procedure is tested in a simulation study based on the assumptions made in a recent paper by Wakker and Klaassen in this journal. We test two variants of CE ratio bootstrap confidence intervals. The first is a bootstrap analogue of the parametric method proposed by Wakker and Klaassen which gives results similar to those obtained with the parametric method. However, computing bootstrap confidence intervals directly for the CE ratio produce results closer to the predetermined significance level.

Algorithms↗

On studentising and blocklength selection for the bootstrap on time series.

For independent data, non-parametric bootstrap is realised by resampling the data with replacement. This approach fails for dependent data such as time series. If the data generating process is at least stationary and mixing, the blockwise bootstrap by drawing subsamples or blocks of the data saves the concept. For the blockwise bootstrap a blocklength has to be selected. We propose a method for selecting the optimal blocklength. To improve the finite size properties of the blockwise bootstrap, studentised statistics is considered. If the statistic can be represented as a smooth function model this studentisation can be approximated efficiently. The studentised blockwise bootstrap method is applied for testing hypotheses on medical time series.

Algorithms↗

The bootstrap and identification of prognostic factors via Cox's proportional hazards regression model.

This paper describes the use of the bootstrap, a new computer-based statistical methodology, to help validate a regression model resulting from the fitting of Cox's proportional hazards model to a set of censored survival data. As an example, we define a prognostic model for outcome in childhood acute lymphocytic leukemia with the Cox model and use of a training set of 224 patients. To validate the accuracy of the model, we use a bootstrap resampling technique to mimic the population under study in two stages. First, we select the important prognostic factors via a stepwise regression procedure with 100 bootstrap samples. Secondly we estimate the corresponding regression parameters for these important factors with 400 bootstrap samples. The bootstrap result suggests that the model constructed from the training set is reasonable.

Child↗

Exploring the order of odds ratios using the bootstrap.

We show the use of the bootstrap resampling method to examine further the order of a series of odds ratios. Specifically, the bootstrap provides a method for estimating the probabilities that one would find in subsequent independent samples from the same population the observed odds ratio rankings. To illustrate this use of the bootstrap, we modelled the responses of 77 white male physicians to an ethical dilemma involving hypothetical patients. Would the physician report positive HIV status to the health department or would he maintain patient confidentiality? To see if a patient's sex, race, or sexual preference would influence the physicians' decisions, each physician received one of eight randomly selected descriptions of a hypothetical patient. To evaluate the initial order of the patient categories, we constructed 1000 bootstrap samples. Black heterosexual males ranked first or second in 92.2 per cent of the bootstrap samples; black homosexual males ranked first, second or third in 88.6 per cent; and white homosexual females ranked sixth or seventh in 82.9 per cent. Thus we would expect to observe these rankings of the categories in a high percentage of subsequent independent samples.

Black or African American↗

Estimating inestimable standard errors in population pharmacokinetic studies: the bootstrap with Winsorization.

A simulation study was performed to determine how inestimable standard errors could be obtained when population pharmacokinetic analysis is performed with the NONMEM software on data from small sample size phase I studies. Plausible sets of concentration-time data for nineteen subjects were simulated using an incomplete longitudinal population pharmacokinetic study design, and parameters of a drug in development that exhibits two compartment linear pharmacokinetics with single dose first order input. They were analyzed with the NONMEM program. Standard errors for model parameters were computed from the simulated parameter values to serve as true standard errors of estimates. The nonparametric bootstrap approach was used to generate replicate data sets from the simulated data and analyzed with NONMEM. Because of the sensitivity of the bootstrap to extreme values, winsorization was applied to parameter estimates. Winsorized mean parameters and their standard errors were computed and compared with their true values as well as the non-winsorized estimates. Percent bias was used to judge the performance of the bootstrap approach (with or without winsorization) in estimating inestimable standard errors of population pharmacokinetic parameters. Winsorized standard error estimates were generally more accurate than non-winsorized estimates because the distribution of most parameter estimates were skewed, sometimes with heavy tails. Using the bootstrap approach combined with winsorization, inestimable robust standard errors can be obtained for NONMEM estimated population pharmacokinetic parameters with > or = 150 bootstrap replicates. This approach was also applied to a real data set and a similar outcome was obtained. This investigation provides a structural framework for estimating inestimable standard errors when NONMEM is used for population pharmacokinetic modeling involving small sample sizes.

Clinical Trials, Phase I as Topic↗

Bootstrapping model of the origin of life.

The origin of life is analyzed in terms of the self-facilitating aspect of evolution. According to the self-facilitation (or bootstrap) principle the structure of biological systems becomes increasingly suited to effective evolutionary search through the process of evolution. The principle may be extended to primitive collections of polymers with catalytic properties. The origin of the code may be based on a form of bootstrapping evolution. Once a primitive code appeared it could become more sophisticated through a multi-coding mechanism together with classical Darwinian mechanisms. The bootstrapping principle is also formulated in terms of the fluctuation-instability framework of Prigogine, Nicolis, and Babloyantz. Primitive collections of polymers accumulate evolution-enhancing redundancies which tend to reduce the extent to which they change when destabilized by fluctuation, but which increase the chances that these changes will lead to new predominant regimes. We show that the bootstrapping of evolutionary amenability is accompanied by the accumulation of a thermodynamic load, that is, a free energy cost which reduces the mechanistic efficiency. The bootstrapping effect suggests that the origin of life is most fruitfully approached as a long process during which the capacity to evolve facilitates itself in a step by step fashion rather than as a series of low probability events.

Genetic Code↗

Giving the boot to the bootstrap: how not to learn the natural numbers.

According to one theory about how children learn the concept of natural numbers, they first determine that "one", "two", and "three" denote the size of sets containing the relevant number of items. They then make the following inductive inference (the Bootstrap): The next number word in the counting series denotes the size of the sets you get by adding one more object to the sets denoted by the previous number word. For example, if "three" refers to the size of sets containing three items, then "four" (the next word after "three") must refer to the size of sets containing three plus one items. We argue, however, that the Bootstrap cannot pick out the natural number sequence from other nonequivalent sequences and thus cannot convey to children the concept of the natural numbers. This is not just a result of the usual difficulties with induction but is specific to the Bootstrap. In order to work properly, the Bootstrap must somehow restrict the concept of "next number" in a way that conforms to the structure of the natural numbers. But with these restrictions, the Bootstrap is unnecessary.

Humans↗

Inferential estimation of polymer quality using bootstrap aggregated neural networks.

Inferential estimation of polymer quality in a batch polymerisation reactor using bootstrap aggregated neural networks is studied in this paper. Number average molecular weight and weight average molecular weight are estimated from the on-line measurements of reactor temperature, jacket inlet and outlet temperatures, coolant flow rate through the jacket, monomer conversion, and the initial batch conditions. Bootstrap aggregated neural networks are used to enhance the accuracy and robustness of neural network models built from a limited amount of training data. The training data set is re-sampled using bootstrap re-sampling with replacement to form several sets of training data. For each set of training data, a neural network model is developed. The individual neural networks are then combined together to form a bootstrap aggregated neural network. Determination of appropriate weights for combining individual networks using principal component regression is proposed in this paper. Confidence bounds for neural network predictions can also be obtained using the bootstrapping technique. The techniques have been successfully applied to the simulation of a batch methyl methacrylate polymerisation reactor.

Journal Article↗

Bootstrap hypothesis tests for evolutionary trees and other dendrograms.

The bootstrap computer-intensive statistical technique is frequently applied to statistical analyses of phylogenetic trees. The widely used rule that a group is supported significantly if it appears in at least 95% of bootstrap trees is conservative in most situations. This paper describes three ways of using the bootstrap to carry out statistical inference on phylogenies. The first method tests whether there is nonrandom support for a single group or tree. The second method compares the support for two groups or trees. The third method tests whether a single group or tree has better support than the set of all possible alternatives; this may be a replacement for the "95% rule." These tests generally require fewer bootstrap trees to be estimated than do other methods of bootstrapping phylogenies. A simple, sequential statistical method can be used to increase the efficiency further. These methods can be applied to tests of multiple hypotheses about a single phylogeny. Parsimony analyses of 5S rRNA sequences of plants and cluster analyses of randomly amplified polymorphic DNA bands in three pathotypes of the cereal eyespot fungus are used as illustrative examples. The tests can be used to analyze dendrograms in subjects other than taxonomy.

Base Sequence↗

Improved bootstrap confidence limits in large-scale phylogenies, with an example from Neo-Astragalus (Leguminosae).

Phylogenetic analyses of large data sets pose special challenges, including the apparent tendency for the bootstrap support for a clade to decline with increased taxon sampling of that clade. We document this decline in data sets with increasing numbers of taxa in Astragalus, the most species-rich angiosperm genus. Support for one subclade, Neo-Astragalus, declined monotonically with increased sampling of taxa inside Neo-Astragalus, irrespective of whether parsimony or neighbor-joining methods were used or of which particular heuristic search algorithm was used (although more stringent algorithms tended to yield higher support). Three possible explanations for this decline were examined, including (1) mistaken assignment of the most recent common ancestor of the taxon sample (and its bootstrap support) with the most recent common ancestor of the clade from which it was sampled; (2) computational limitations of heuristic search strategies; and (3) statistical bias in bootstrap proportions, especially that from random homoplasy distributed among taxa. The best explanation appears to be (3), although computational shortcomings (2) may explain some of the problem. The bootstrap proportion, as currently used in phylogenetic analysis, does not accurately capture the classical notion of confidence assessments on the null hypothesis of nonmonophyly, especially in large data sets. More accurate assessments of confidence as type I error levels (relying on iterated bootstrap methods) remove most of the monotonic decline in confidence with increasing numbers of taxa.

Algorithms↗

Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap.

MOTIVATION: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap. RESULTS: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices. AVAILABILITY: The sequence sets and code for performing these analyses are available from http://compbio.berkeley.edu/. CONTACT: brenner@compbio.berkeley.edu.

Algorithms↗

Improved confidence intervals in quantitative trait loci mapping by permutation bootstrapping.

The nonparametric bootstrap approach is known to be suitable for calculating central confidence intervals for the locations of quantitative trait loci (QTL). However, the distribution of the bootstrap QTL position estimates along the chromosome is peaked at the positions of the markers and is not tailed equally. This results in conservativeness and large width of the confidence intervals. In this study three modified methods are proposed to calculate nonparametric bootstrap confidence intervals for QTL locations, which compute noncentral confidence intervals (uncorrected method I), correct for the impact of the markers (weighted method I), or both (weighted method II). Noncentral confidence intervals were computed with an analog of the highest posterior density method. The correction for the markers is based on the distribution of QTL estimates along the chromosome when the QTL is not linked with any marker, and it can be obtained with a permutation approach. In a simulation study the three methods were compared with the original bootstrap method. The results showed that it is useful, first, to compute noncentral confidence intervals and, second, to correct the bootstrap distribution of the QTL estimates for the impact of the markers. The weighted method II, combining these two properties, produced the shortest and less biased confidence intervals in a large number of simulated configurations.

Chromosome Mapping↗

Construction and bootstrap analysis of DNA fingerprinting-based phylogenetic trees with the freeware program FreeTree: application to trichomonad parasites.

The Win95/98/NT program FreeTree for computation of distance matrices and construction of phylogenetic or phenetic trees on the basis of random amplified polymorphic DNA (RAPD), RFLP and allozyme data is presented. In contrast to other similar software, the program FreeTree (available at http://www.natur.cuni.cz/~flegr/programs/freetree or http://ijs.sgmjournals.org/content/vol51/issue3/) can also assess the robustness of the tree topology by bootstrap, jackknife or operational taxonomic unit-jackknife analysis. Moreover, the program can be also used for the analysis of data obtained in several independent experiments performed with non-identical subsets of taxa. The function of the program was demonstrated by an analysis of RAPD data from 42 strains of 10 species of trichomonads. On the phylogenetic tree constructed using FreeTree, the high bootstrap values and short terminal branches for the Tritrichomonas foetus/suis 14-strain branch suggested relatively recent and probably clonal radiation of this species. At the same time, the relatively lower bootstrap values and long terminal branches for the Trichomonas vaginalis 20-strain branch suggested more ancient radiation of this species and the possible existence of genetic recombination (sexual reproduction) in this human pathogen. The low bootstrap values and the star-like topology of the whole Trichomonadidae tree confirm that the RAPD method is not suitable for phylogenetic analysis of protozoa at the level of higher taxa. It is proposed that the repeated bootstrap analysis should be an obligatory part of any RAPD study. It makes it possible to assess the reliability of the tree obtained and to adjust the amount of collected data (the number of random primers) to the amount of phylogenetic signals in the RAPD data of the taxon analysed. The FreeTree program makes such analysis possible.

Animals↗

Bootstrapping: applications to psychophysiology.

This paper presents the statistical technique known as the bootstrap to the general audience of psychophysiologists. The bootstrap, introduced by Efron (1979), allows data analysts to study the distribution of sample statistics that might otherwise be too complicated to consider. The technique, which requires simple calculations, involves drawing repeated samples (with replacement) from the empirical--or the actual--data distribution and then building a distribution for a statistic by calculating a value of the statistic for each sample. The bootstrap can be used to obtain confidence intervals, standard errors, and even higher moments for the statistic. It is similar to the well-known jackknife of Quenouille and Tukey. After discussing the history and theory of both the bootstrap and the jackknife, we illustrate the use of the bootstrap in the statistical analysis of correlation coefficients and the general linear model.

Algorithms↗

A bootstrap method to avoid the effect of concurvity in generalised additive models in time series studies of air pollution.

BACKGROUND: In recent years a great number of studies have applied generalised additive models (GAMs) to time series data to estimate the short term health effects of air pollution. Lately, however, it has been found that concurvity--the non-parametric analogue of multicollinearity--might lead to underestimation of standard errors of the effects of independent variables. Underestimation of standard errors means that for concurvity levels commonly present in the data, the risk of committing type I error rises by over threefold. METHODS: This study developed a conditional bootstrap methology that consists of assuming that the outcome in any observation is conditional upon the values of the set of independent variables used. It then tested this procedure by means of a simulation study using a Poisson additive model. The response variable of this model is a function of an unobserved confounding variable (that introduces trend and seasonality), real black smoke data, and temperature. Scenarios were created with different coefficients and degrees of concurvity. RESULTS: Conditional bootstrap provides confidence intervals with coverages close to nominal (95%), irrespective of the degree of concurvity, number of variables in the model or magnitude of the coefficient to be estimated (for example, for a concurvity of 0.85, bootstrap confidence interval coverage is 95% compared with 71% in the case of the asymptotic interval obtained directly with S-plus gam function). CONCLUSIONS: The bootstrap method avoids the problem of concurvity in time series studies of air pollution, and is easily generalised to non-linear dose-risk effects. All bootstrap calculations described in this paper can be performed using S-Plus gam.boot software.

Air Pollution↗