PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

When being "most likely" is not enough: examining the performance of three uses of the parametric bootstrap in phylogenetics.

I show that three parametric-bootstrap (PB) applications that have been proposed for phylogenetic analysis, can be misleading as currently implemented. First, I show that simulating a topology estimated from preliminary data in order to determine the sequence length that should allow the best tree obtained from more extensive data to be correct with a desired probability, delivers an accurate estimate of this length only in topological situations in which most preliminary trees are expected to be both correct and statistically significant, i.e. when no further analysis would be needed. Otherwise, one obtains strong underestimates of the length or similarly biased values for incorrect trees. Second, I show that PB-based topology tests that use as null hypothesis the most likely tree congruent with a pre-specified topological relationship alternative to the unconstrained most likely tree, and simulate this tree for P value estimation, produce excessive type I error (from 50% to 600% and higher) when they are applied to null data generated by star-shaped or dichotomous four-taxon topologies. Simulating the most likely star topology for P value estimation results instead in correct type-I-error production even when the null data are generated by a dichotomous topology. This is a strong indication that the star topology is the correct default null hypothesis for phylogenies. Third, I show that PB-estimated confidence intervals (CIs) for the length of a tree branch are generally accurate, although in some situations they can be strongly over- or under-estimated relative to the "true" CI. Attempts to identify a biased CI through a further round of simulations were unsuccessful. Tracing the origin and propagation of parameter estimate error through the CI estimation exercise, showed that the sparseness of site-patterns which are crucial to the estimation of pivotal parameters, can allow homoplasy to bias these estimates and ultimately the PB-based CI estimation. Concluding, I stress that statistical techniques that simulate models estimated from limited data need to be carefully calibrated, and I defend the point that pattern-sparseness assessment will be the next frontier in the statistical analysis of phylogenies, an effort that will require taking advantage of the merits of black-box maximum-likelihood approaches and of insights from intuitive, site-pattern-oriented approaches like parsimony.

Algorithms↗

Testing population genetic structure using parametric bootstrapping and MIGRATE-N.

We present a method for investigating genetic population structure using sequence data. Our hypothesis states that the parameters most responsible for the formation of genetic structure among different populations are the relative rates of mutation (micro) and migration (M). The evolution of genetic structure among different populations requires rates of M << p because this allows population-specific mutation to accumulate. Rates of micro << M will result in populations that are effectively panmictic because genetic differentiation will not develop among demes. Our test is implemented by using a parametric bootstrap to create the null distribution of the likelihood of the data having been produced under an appropriate model of sequence evolution and a migration rate sufficient to approximate panmixia. We describe this test, then apply it to mtDNA data from 243 plethodontid salamanders. We are able to reject the null hypothesis of no population structure on all but smallest geographic scales, a result consistent with the apparent lack of migration in Plethodon idahoensis. This approach represents a new method of investigating population structure with haploid DNA, and as such may be particularly useful for preliminary investigation of non-model organisms in which multi-locus nuclear data are not available.

Animals↗

What sort of innate structure is needed to "bootstrap" into syntax?

The paper starts from Pinker's theory of the acquisition of phrase structure; it shows that it is possible to drop all the assumptions about innate syntactic structure from this theory. These assumptions can be replaced by assumptions about the basic structure of semantic representation available at the outset of language acquisition, without penalizing the acquisition of basic phrase structure rules. Essentially, the role played by X-bar theory in Pinker's model would be played by the (presumably innate) structure of the language of thought in the revised parallel model. Bootstrapping and semantic assimilation theories are shown to be formally very similar, though making different primitive assumptions. In their primitives, semantic assimilation theories have the advantage that they can offer an account of the origin of syntactic categories instead of postulating them as primitive. Ways of improving on the semantic assimilation version of Pinker's theory are considered, including a way of deriving the NP-VP constituent division that appears to have a better fit than Pinker's to evidence on language variation.

Child, Preschool↗

The use of multiple frames in verb learning via syntactic bootstrapping.

Following the original Syntactic Bootstrapping proposal of Landau and Gleitman (1985), this study investigated whether young 2-year-old children (mean age = 28 months) can use multiple syntactic frames, in addition to the extralinguistic scene, to help focus on the meaning of a novel verb. The multiple frames tested were combinations of transitive and intransitive frames in two alternation patterns, Causative and Omitted Object. By hypothesis, the Causative alternation would be more predictive of actions involving physical causation and the Omitted Object alternation more predictive of actions involving repeated physical contact without causation. Subjects were presented with videos depicting both actions, together with a novel verb. The actions were subsequently separated, and the children were asked to select which action was the referent of the novel verb. The novel verb was presented either in transitive and intransitive frames in the Causative alternation (CS: The duck is sebbing the frog, the frog is sebbing) or the Omitted Object alternation (OO: The duck is sebbing the frog, the duck is sebbing), or in intransitive frames only (IO: The duck is sebbing), or without a frame (FF: Sebbing!). In the CS, IO, and FF conditions, children preferred the causative action as the referent of the verb. However, the girls in the OO condition showed a significantly different preference, and looked more toward the contact actions than their peers in the other conditions did. This study thus provides the first experimental evidence that young 2-year-old children can use multiple syntactic frames to help determine the meaning of a novel verb.

Child Development↗

Comparison of receiver operating curves derived from the same population: a bootstrapping approach.

The receiver operating curve (ROC) gives a representation of sensitivity and specificity of a prediction model when varying the cutpoint of a decision rule on a whole spectrum. Evaluation of two models established (or tested) in the same population of patients warrants a valid statistical comparison of their ROC curves. Hanley et al. recently provided a method for overall comparison of ROC curves (J. A. Hanley and B. J. McNeil, Radiology 148, 839-843, 1983). Often ROC curves cross, or differ in only a part of their courses. Bootstrapping of ROC curves is proposed as a graphical check for the statistical significance of differences confined to a part of the curve. An example comparing two models of prediction of coronary artery disease progression is given to illustrate this new approach.

Coronary Angiography↗

A conditional bootstrap procedure for reconstruction of the incubation period of AIDS.

Data on the incubation period of AIDS patients are often fragmented and censored. parametric models have been proposed in the literature to impute the missing segment of the incubation period. The numerical results vary widely with the parametric models used. We propose a nonparametric conditional bootstrap (CB) procedure for imputation. The quality of the CB data is studied by checking the asymptotic accuracy of the CB estimators. We establish the asymptotic accuracy of the CB procedure for two basic nonparametric estimators: the empirical distribution function and a kernel-type conditional empirical distribution function. The rates of convergence of the CB approximation are obtained. The results for the kernel-type estimators hold also for the nearest-neighbor-type estimators.

Acquired Immunodeficiency Syndrome↗

Application of a statistical bootstrapping technique to calculate growth rate variance for modelling psychrotrophic pathogen growth.

The inherent variability or 'variance' of growth rate measurements is critical to the development of accurate predictive models in food microbiology. A large number of measurements are typically needed to estimate variance. To make these measurements requires a significant investment of time and effort. If a single growth rate determination is based on a series of independent measurements, then a statistical bootstrapping technique can be used to simulate multiple growth rate measurements from a single set of experiments. Growth rate variances were calculated for three large datasets (Listeria monocytogenes, Listeria innocua, and Yersinia enterocolitica) from our laboratory using this technique. This analysis revealed that the population of growth rate measurements at any given condition are not normally distributed, but instead follow a distribution that is between normal and Poisson. The relationship between growth rate and temperature was modeled by response surface models using generalized linear regression. It was found that the assumed distribution (i.e. normal, Poisson, gamma or inverse normal) of the growth rates influenced the prediction of each of the models used. This research demonstrates the importance of variance and assumptions about the statistical distribution of growth rates on the results of predictive microbiological models.

Bacteria↗

A bootstrap method to compare the shapes of two scalp fields.

A method is described to compare two evoked potential scalp fields in order to decide if the two fields are the same or different. The method uses Efron's bootstrap technique which avoids potential errors due to assumptions about the underlying stochastic process. It is configured to focus only on the shape of the evoked potential scalp field. The method is applied to a simple visually evoked potential paradigm and results are compared to the chi-square test using data from 7 normal subjects.

Evoked Potentials, Visual↗

Evaluation of screening methods for Down's syndrome using bootstrap comparison of ROC curves.

This paper concerns the prediction of fetal Down's syndrome in pregnant women. Down's syndrome is the most common congenital cause of severe mental retardation. We elaborate two predictive functions of trisomy 21, combining maternal age and a maternal serum marker. We evaluated them by means of receiver operating characteristic (ROC) curves which give a representation of sensitivity and specificity of a prediction model when varying the cutoff of the predictor on the whole spectrum. Since normal statistical methods for comparison of ROC curves rely on distributional assumptions which were not verified, we used bootstrapping of ROC curves as a check for the statistical significance of differences between the areas under the curves.

Algorithms↗

Using permutation tests and bootstrap confidence limits to analyze repeated events data from clinical trials.

In clinical trials comparing treatments for superficial bladder cancer, patients are at risk of repeated recurrences of their disease. Statistical methods of analyzing such data are required. This article presents a nonparametric approach. A statistical test to compare the recurrence or tumor rates in two treatment groups, using the randomization distribution, is described. Confidence intervals for the rate ratio are determined from the bootstrap distribution. The implementation of both requires Monte Carlo methods. Computer simulations support the use of these nonparametric methods when there are more than 60 recurrences in each treatment group. An example illustrating their use is given. The strategy adopted for analysis of these data could be applied to other clinical trials where standard methodology is inappropriate.

Biometry↗

An attenuation of the 'normal' category effect in patients with Alzheimer's disease: a review and bootstrap analysis.

There is a consensus that Alzheimer's disease (AD) impairs semantic information, with one of the first markers being anomia i.e. an impaired ability to name items. Doubts remain, however, about whether this naming impairment differentially affects items from the living and nonliving knowledge domains. Most studies have reported an impairment for naming living things (e.g. animals or plants), a minority have found an impairment for nonliving things (e.g. tools or vehicles), and some have found no category-specific effect. A survey of the literature reveals that this lack of agreement may reflect a failure to control for intrinsic variables (such as familiarity) and the problems associated with ceiling effects in the control data. Investigating picture naming in 32 AD patients and 34 elderly controls, we used bootstrap techniques to deal with the abnormal distributions in both groups. Our analyses revealed the previously reported impairment for naming living things in AD patients and that this persisted even when intrinsic variables were covaried; however, covarying control performance eliminated the significant category effect. Indeed, the within-group comparison of living and nonliving naming revealed a larger effect size for controls than patients. We conclude that the category effect in Alzheimer's disease is no larger than is expected in the healthy brain and may even represent a small diminution of the normal profile.

Aged↗

A weighted bootstrap method for the determination of probability density functions of freshwater distribution coefficients (Kds) of Co, Cs, Sr and I radioisotopes.

The objective of the study was to provide global probability density functions (PDFs) representing the uncertainty of distribution coefficients (Kds) in freshwater for radioisotopes of Co, Cs, Sr and I. A comprehensive database containing Kd values referenced in 61 articles was first built and quality scores were affected to each data point according to various criteria (e.g. presentation of data, contact times, pH, solid-to-liquid ratio, expert judgement). A weighted bootstrapping procedure was then set up in order to build PDFs, in such a way that more importance is given to the most relevant data points (i.e. those corresponding to typical natural environments). However, it was also assessed that the relevance and the robustness of the PDFs determined by our procedure depended on the number of Kd values in the database. Owing to the large database, conditional PDFs were also proposed, for site studies where some parametric information is known (e.g. pH, contact time between radionuclides and particles, solid-to-liquid ratio). Such conditional PDFs reduce the uncertainty on the Kd values. These global and conditional PDFs are useful for end-users of dose models because the uncertainty and sensitivity of Kd values are taking into account.

Cesium↗

Bootstrap confidence intervals for the mode of the hazard function.

In many applications of lifetime data analysis, it is important to perform inferences about the mode of the hazard function in situations of lifetime data modeling with unimodal hazard functions. For lifetime distributions where the mode of the hazard function can be analytically calculated, its maximum likelihood estimator is easily obtained from the invariance properties of the maximum likelihood estimators. From the asymptotical normality of the maximum likelihood estimators, confidence intervals can be obtained. However, these results might not be very accurate for small sample sizes and/or large proportion of censored observations. Considering the log-logistic distribution for the lifetime data with shape parameter beta>1, we present and compare the accuracy of asymptotical confidence intervals with two confidence intervals based on bootstrap simulation. The alternative methodology of confidence intervals for the mode of the log-logistic hazard function are illustrated in three numerical examples.

Confidence Intervals↗

Identifying barriers to the effective use of clinical reminders: bootstrapping multiple methods.

Advances in electronic medical record capabilities enable clinical reminders to inform providers when recommended actions are "due" for a patient. Despite evidence that they improve adherence to guidelines, the Veteran's Health Administration (VHA) has experienced challenges in having providers consistently use clinical reminders as intended. In this paper, we describe how multiple methods were used to opportunistically triangulate, or "bootstrap," an understanding of barriers to the effective use of clinical reminders in the VHA. In an initial study using ethnographic observations and semi-structured interviews of HIV clinical reminders, we identified six barriers to effective use: workload, time to remove inapplicable reminders, false alarms, training, reduced eye contact, and the use of paper forms rather than software. In a second study, we collected open-ended and closed-ended data regarding barriers and facilitators to the use of clinical reminders in general in the VHA through a survey of 261 participants at a national informatics meeting, where 104 of 142 VHA health care facilities were represented. The findings from the second study extended our understanding of the previously identified barriers. In addition, four new barriers were identified: ease of use issues, accessibility of workstations, resident physicians and trainees, and administration benefiting more than providers from clinical reminder use. We discuss potential implications regarding the similarities and differences in study findings for factors to consider in planning interventions to improve clinical reminder use.

Appointments and Schedules↗

A global goodness-of-fit test for receiver operating characteristic curve analysis via the bootstrap method.

OBJECTIVE: Medical classification accuracy studies often yield continuous data based on predictive models for treatment outcomes. A popular method for evaluating the performance of diagnostic tests is the receiver operating characteristic (ROC) curve analysis. The main objective was to develop a global statistical hypothesis test for assessing the goodness-of-fit (GOF) for parametric ROC curves via the bootstrap. DESIGN: A simple log (or logit) and a more flexible Box-Cox normality transformations were applied to untransformed or transformed data from two clinical studies to predict complications following percutaneous coronary interventions (PCIs) and for image-guided neurosurgical resection results predicted by tumor volume, respectively. We compared a non-parametric with a parametric binormal estimate of the underlying ROC curve. To construct such a GOF test, we used the non-parametric and parametric areas under the curve (AUCs) as the metrics, with a resulting p value reported. RESULTS: In the interventional cardiology example, logit and Box-Cox transformations of the predictive probabilities led to satisfactory AUCs (AUC=0.888; p=0.78, and AUC=0.888; p=0.73, respectively), while in the brain tumor resection example, log and Box-Cox transformations of the tumor size also led to satisfactory AUCs (AUC=0.898; p=0.61, and AUC=0.899; p=0.42, respectively). In contrast, significant departures from GOF were observed without applying any transformation prior to assuming a binormal model (AUC=0.766; p=0.004, and AUC=0.831; p=0.03), respectively. CONCLUSIONS: In both studies the p values suggested that transformations were important to consider before applying any binormal model to estimate the AUC. Our analyses also demonstrated and confirmed the predictive values of different classifiers for determining the interventional complications following PCIs and resection outcomes in image-guided neurosurgery.

Adolescent↗

A method of focused classification, based on the bootstrap 3D variance analysis, and its application to EF-G-dependent translocation.

The bootstrap-based method for calculation of the 3D variance in cryo-EM maps reconstructed from sets of their projections was applied to a dataset of functional ribosomal complexes containing the Escherichia coli 70S ribosome, tRNAs, and elongation factor G (EF-G). The variance map revealed regions of high variability in the intersubunit space of the ribosome: in the locations of tRNAs, in the putative location of EF-G, and in the vicinity of the L1 protein. This result indicated heterogeneity of the dataset. A method of focused classification was put forward in order to sort out the projection data into approximately homogenous subsets. The method is based on the identification and localization of a region of high variance that a subsequent classification step can be focused on by the use of a 3D spherical mask. After initial classification, template volumes are created and are subsequently refined using a multireference 3D projection alignment procedure. In the application to the ribosome dataset, the two resulting structures were interpreted as resulting from ribosomal complexes with bound EF-G and an empty A site, or, alternatively, from complexes that had no EF-G bound but had both A and P sites occupied by tRNA. The proposed method of focused classification proved to be a successful tool in the analysis of the heterogeneous cryo-EM dataset. The associated calculation of the correlations within the density map confirmed the conformational variability of the complex, which could be interpreted in terms of the ribosomal elongation cycle.

Cryoelectron Microscopy↗

Estimation of variance in single-particle reconstruction using the bootstrap technique.

Density maps of a molecule obtained by single-particle reconstruction from thousands of molecule projections exhibit strong changes in local definition and reproducibility, as a consequence of conformational variability of the molecule and non-stoichiometry of ligand binding. These changes complicate the interpretation of density maps in terms of molecular structure. A three-dimensional (3-D) variance map provides an effective tool to assess the structural definition in each volume element. In this work, the different contributions to the 3-D variance in a single-particle reconstruction are discussed, and an effective method for the estimation of the 3-D variance map is proposed, using a bootstrap technique of sampling. Computations with test data confirm the viability, computational efficiency, and accuracy of the method under conditions encountered in practical circumstances.

Algorithms↗

Classification of technical pitfalls in objective universal hearing screening by otoacoustic emissions, using an ARMA model of the stimulus waveform and bootstrap cross-validation.

Transient-evoked otoacoustic emissions (TEOAE) are widely used for objective hearing screening in neonates. Their main shortcoming is their sensitivity to adverse conditions for sound transmission through the middle-ear, to and from the cochlea. We study here whether a close examination of the stimulus waveform (SW) recorded in the ear canal in the course of a screening test can pinpoint the most frequent middle-ear dysfunctions, thus allowing screeners to avoid misclassifying the corresponding babies as deaf for lack of TEOAE. Three groups of SWs were defined in infants (6-36 months of age) according to middle-ear impairment as assessed by independent testing procedures, and analyzed in the frequency domain where their properties are more readily interpreted than in the time domain. Synthetic SW parameters were extracted with the help of an autoregressive and moving average (ARMA) model, then classified using a maximum likelihood criterion and a bootstrap cross-validation. The best classification performance was 79% with a lower limit (with 90% confidence) of 60%, showing the results' consistency. We therefore suggest that new parameters and methodology based upon a more thorough analysis of SWs can improve the efficiency of TEOAE-based tests by helping the most frequent technical pitfalls to be identified.

Acoustic Impedance Tests↗