PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Evaluation of uncertainty parameters estimated by different population PK software and methods.

The uncertainty associated with parameter estimations is essential for population model building, evaluation, and simulation. Summarized by the standard error (SE), its estimation is sometimes questionable. Herein, we evaluate SEs provided by different non linear mixed-effect estimation methods associated with their estimation performances. Methods based on maximum likelihood (FO and FOCE in NONMEM, nlme in Splus, and SAEM in MONOLIX) and Bayesian theory (WinBUGS) were evaluated on datasets obtained by simulations of a one-compartment PK model using 9 different designs. Bootstrap techniques were applied to FO, FOCE, and nlme. We compared SE estimations, parameter estimations, convergence, and computation time. Regarding SE estimations, methods provided concordant results for fixed effects. On random effects, SAEM and WinBUGS, tended respectively to under or over-estimate them. With sparse data, FO provided biased estimations of SE and discordant results between bootstrapped and original datasets. Regarding parameter estimations, FO showed a systematic bias on fixed and random effects. WinBUGS provided biased estimations, but only with sparse data. SAEM and WinBUGS converged systematically while FOCE failed in half of the cases. Applying bootstrap with FOCE yielded CPU times too large for routine application and bootstrap with nlme resulted in frequent crashes. In conclusion, FO provided bias on parameter estimations and on SE estimations of random effects. Methods like FOCE provided unbiased results but convergence was the biggest issue. Bootstrap did not improve SEs for FOCE methods, except when confidence interval of random effects is needed. WinBUGS gave consistent results but required long computation times. SAEM was in-between, showing few under-estimated SE but unbiased parameter estimations.

Bayes Theorem↗

Advances in the computational study of language acquisition.

This paper provides a tutorial introduction to computational studies of how children learn their native languages. Its aim is to make recent advances accessible to the broader research community, and to place them in the context of current theoretical issues. The first section locates computational studies and behavioral studies within a common theoretical framework. The next two sections review two papers that appear in this volume: one on learning the meanings of words and one or learning the sounds of words. The following section highlights an idea which emerges independently in these two papers and which I have dubbed autonomous bootstrapping. Classical bootstrapping hypotheses propose that children begin to get a toc-hold in a particular linguistic domain, such as syntax, by exploiting information from another domain, such as semantics. Autonomous bootstrapping complements the cross-domain acquisition strategies of classical bootstrapping with strategies that apply within a single domain. Autonomous bootstrapping strategies work by representing partial and/or uncertain linguistic knowledge and using it to analyze the input. The next two sections review two more more contributions to this special issue: one on learning word meanings via selectional preferences and one on algorithms for setting grammatical parameters. The final section suggests directions for future research.

Algorithms↗

A new variant of GB virus C/hepatitis G virus (GBV-C/HGV) from South Africa.

Phylogenetic analysis of the 5' non-coding region (5'NCR) sequences has demonstrated that GB virus C/hepatitis G virus (GBV-C/HGV) can be separated into three major groups that correlate with the geographic origin of the isolate. Sequence analysis of the 5'NCR of 54 GBV-C/HGV isolates from 31 blood donors, 11 haemodialysis patients and 12 patients with chronic liver disease suggests the presence of a new variant of GBV-C/HGV in the province of KwaZulu Natal, South Africa. Eleven isolates grouped as group 1 variants (bootstrap support, 90%) found predominantly in West and Central Africa, a further six isolates grouped as group 2 variants (bootstrap support, 58%) found in Europe and North America; five of which grouped as 2a (bootstrap support, 91%) and one as 2b (bootstrap support, 87%), the latter also includes isolates from Japan, East Africa and Pakistan. Although the remaining 37 GBV-C/HGV isolates were more closely related to group 1 variants (bootstrap support, 90%), they formed a cluster, which was distinct from all other known GBV-C/HGV sequences. None of the South African isolates grouped with group 3 variants described from Southeast Asia. Three variants of GBV-C/HGV exist in KwaZulu Natal: groups 1, 2 and a new variant, which is distinct from other African isolates.

Base Sequence↗

[Comparison of 2 methods for calculating uncertainty in laboratory analysis].

OBJECTIVE: To compare two methods in the estimation of the uncertainty in laboratory quality control. METHODS: A computerized simulation is performed to compare the delta method (suggested by the International Organization for Standardization and the Entidad Nacional de Acreditación) and a bootstrap-based method. The simulation includes several situations with different environmental conditions and different relationships between the analyzed variables. RESULTS: The mean in the coverage obtained by the estimated confidence intervals is higher and closer to the nominal using the bootstrap than using the delta method. The most important differences are observed in the coverage percent distribution: while using the bootstrap, a great number of simulations obtain coverage near the nominal of 95%; using the delta method the coverage are more dispersed, including coverage of 100% in some occasions and lesser than 80% in others. The bootstrap offers very similar results under different conditions, including in the presence of unknown and unmeasured variables or when the analyzed variables are correlated. The delta method shows poorer results in both situations. CONCLUSION: The uncertainty in the laboratory quality control can be estimated more accurately with bootstrapping than with the delta method.

Clinical Laboratory Techniques↗

Internal and external validation of predictive models: a simulation study of bias and precision in small samples.

We performed a simulation study to investigate the accuracy of bootstrap estimates of optimism (internal validation) and the precision of performance estimates in independent validation samples (external validation). We combined two data sets containing children presenting with fever without source (n=376+179=555; 120 bacterial infections). Random samples were drawn from this combined data set for the development (n=376) and validation (n=179) of logistic regression models. The models included statistically significant predictors for infection selected from a set of 57 candidate predictors. Model development, including the selection of predictors, and validation were repeated in a bootstrapping procedure. The resulting expected optimism estimate in the receiver operating characteristic (ROC) area was compared with the observed optimism according to independent validation samples. The average apparent ROC area was 0.74, which was expected (based on bootstrapping) to decrease by 0.07 to 0.67, whereas the observed decrease in the validation samples was 0.09 to 0.65. Omitting the selection of predictors from the bootstrap procedure led to a severe underestimation of the optimism (decrease 0.006). The standard error of the observed ROC area in the independent validation samples was large (0.05). We recommend bootstrapping for internal validation because it gives reasonably valid estimates of the expected optimism in predictive performance provided that any selection of predictors is taken into account. For external validation, substantial sample sizes should be used for sufficient power to detect clinically important changes in performance as compared with the internally validated estimate.

Bias↗

Parametric and non-parametric statistical analysis of DT-MRI data.

In this work parametric and non-parametric statistical methods are proposed to analyze Diffusion Tensor Magnetic Resonance Imaging (DT-MRI) data. A Multivariate Normal Distribution is proposed as a parametric statistical model of diffusion tensor data when magnitude MR images contain no artifacts other than Johnson noise. We test this model using Monte Carlo (MC) simulations of DT-MRI experiments. The non-parametric approach proposed here is an implementation of bootstrap methodology that we call the DT-MRI bootstrap. It is used to estimate an empirical probability distribution of experimental DT-MRI data, and to perform hypothesis tests on them. The DT-MRI bootstrap is also used to obtain various statistics of DT-MRI parameters within a single voxel, and within a region of interest (ROI); we also use the bootstrap to study the intrinsic variability of these parameters in the ROI, independent of background noise. We evaluate the DT-MRI bootstrap using MC simulations and apply it to DT-MRI data acquired on human brain in vivo, and on a phantom with uniform diffusion properties.

Artifacts↗

Pvclust: an R package for assessing the uncertainty in hierarchical clustering.

SUMMARY: Pvclust is an add-on package for a statistical software R to assess the uncertainty in hierarchical cluster analysis. Pvclust can be used easily for general statistical problems, such as DNA microarray analysis, to perform the bootstrap analysis of clustering, which has been popular in phylogenetic analysis. Pvclust calculates probability values (p-values) for each cluster using bootstrap resampling techniques. Two types of p-values are available: approximately unbiased (AU) p-value and bootstrap probability (BP) value. Multiscale bootstrap resampling is used for the calculation of AU p-value, which has superiority in bias over BP value calculated by the ordinary bootstrap resampling. In addition the computation time can be enormously decreased with parallel computing option.

Algorithms↗

A comparison of three methods for calculating confidence intervals for the benchmark dose.

Various methods exist to calculate confidence intervals for the benchmark dose in risk analysis. This study compares the performance of three such methods in fitting nonlinear dose-response models: the delta method, the likelihood-ratio method, and the bootstrap method. A data set from a developmental toxicity test with continuous, ordinal, and quantal dose-response data is used for the comparison of these methods. Nonlinear dose-response models, with various shapes, were fitted to these data. The results indicate that a few thousand runs are generally needed to get stable confidence limits when using the bootstrap method. Further, the bootstrap and the likelihood-ratio method were found to give fairly similar results. The delta method, however, resulted in some cases in different (usually narrower) intervals, and appears unreliable for nonlinear dose-response models. Since the bootstrap method is more time consuming than the likelihood-ratio method, the latter is more attractive for routine dose-response analysis. In the context of a probabilistic risk assessment the bootstrap method has the advantage that it directly links to Monte Carlo analysis.

Animals↗

Statistical analysis of reduced pair correlation functions of capillaries in the prostate gland.

Blood capillaries are thread-like structures that may be considered as an example of a spatial fibre process in three dimensions. At light microscopy, the capillary profiles appear as a planar point process on sections. It has recently been shown that the observed pair correlation function g(r) of the centres of the fibre profiles on two-dimensional sections may be used to estimate the reduced pair correlation function of stationary and isotropic fibre processes in three dimensions. In the present study, we explored how this approach may be extended to statistical analysis of reduced g-functions of capillaries from multiple specimens of different groups and with replicated observations. The methods were applied to normal prostatic tissue compared with prostate cancer. Confidence intervals for the mean reduced g-functions of groups were estimated for fixed r-values parametrically using the t-distribution, and by bootstrap methods. Each estimated reduced g-function was furthermore characterized in terms of its first maximum and minimum. The mean length of capillaries per unit tissue volume was significantly higher in prostate cancer tissue than in normal prostate tissue. Significant differences between the mean reduced g-functions of malignant and benign lesions could be demonstrated for two domains of r-values. In general, bootstrap-based confidence intervals were slightly wider than parametrically estimated confidence intervals. Falsely negative lower bounds of the intervals, which sometimes arose using the parametric approach, could be avoided by the bootstrap method. Testing of group mean values for significant differences by the bootstrap method yielded more conservative results than multiple t-tests. The functional value of the first maximum of the reduced g-function and a global statistical parameter of short-range ordering was significantly reduced in the carcinoma group. Prostate cancer tissue is more densely supplied with capillaries than normal prostate tissue and the three-dimensional arrangement of the vessels differs with respect to interaction at various distance ranges. In the local approach used here, bootstrap methods can be used as a robust statistical tool for the computation of confidence intervals and group comparisons of mean reduced g-functions at specific ranges of interaction.

Capillaries↗

Estimating upper confidence limits for extra risk in quantal multistage models.

Multistage models are frequently applied in carcinogenic risk assessment. In their simplest form, these models relate the probability of tumor presence to some measure of dose. These models are then used to project the excess risk of tumor occurrence at doses frequently well below the lowest experimental dose. Upper confidence limits on the excess risk associated with exposures at these doses are then determined. A likelihood-based method is commonly used to determine these limits. We compare this method to two computationally intensive "bootstrap" methods for determining the 95% upper confidence limit on extra risk. The coverage probabilities and bias of likelihood-based and bootstrap estimates are examined in a simulation study of carcinogenicity experiments. The coverage probabilities of the nonparametric bootstrap method fell below 95% more frequently and by wider margins than the better-performing parametric bootstrap and likelihood-based methods. The relative bias of all estimators are seen to be affected by the amount of curvature in the true underlying dose-response function. In general, the likelihood-based method has the best coverage probability properties while the parametric bootstrap is less biased and less variable than the likelihood-based method. Ultimately, neither method is entirely satisfactory for highly curved dose-response patterns.

Animals↗

A comparison of different strategies for computing confidence intervals of the linkage disequilibrium measure D'.

Many linkage disequilibrium (LD) measures have been used to study LD patterns and for haplotype block partitioning. We examine the properties of one of these measures, Lewontin's D', in order to understand the dependency of its confidence interval (CI) to allele frequency and sample size as well as its applications in defining haplotype blocks. This measure and its CIs were used to partition haplotypes into blocks by Gabriel et al. as well as in many other applications. Gabriel et al. utilized a bootstrap approach to calculate the CI for D'. Under this method, over 1,000 bootstrap samples may be needed to obtain an accurate estimate of the CI for each pair of single nucleotide polymorphism (SNP) markers which can be very computationally intensive, particularly when many SNP markers are involved. We develop two alternative methods for calculating the CI for D' without bootstrap: one based on the approximate variance of D' given by Zapata et al. and the other based on a maximum likelihood estimate (MLE) of D' together with Fisher Information theory. Both methods depend on normal approximation for the estimates of D' for large sample sizes. We assess and compare the coverage of the CIs using the three methods through extensive simulations. We define the coverage as the fraction of times the estimated CI contains the true value of D'. In general, the average coverage of the bootstrap method is less than the pre-specified coverage. When the sample size is small (< or = 100), the remaining two methods slightly under estimate the coverage with MLE approach having smaller standard error compared to Zapata's method. When the sample size is large (> or = 200) , the estimated coverage from both Zapata's and MLE methods are very close to the pre-specified coverage with the MLE method having the smallest standard error among all three methods. In most typical scenarios, we recommend the use of MLE method for all sample sizes. Only under rare specific cases, would the bootstrap method be better suited for determining the CI, i.e. small sample size, at extreme allele frequencies and -3 < D' < 0.

Computational Biology↗

Validation of models for predicting the use of health technologies.

Validation must be carried out before a model can be used confidently as a tool of managerial decision making in health care. The authors describe a bootstrap approach to validating models for predicting the utilization of four technologies used in neonatal care: measurement of blood gases (gasometry), the oxygen hood, continuous positive airway pressure (CPAP), and mechanical ventilation. These models were fitted by stepwise multiple linear regression from 20 prognostic covariates of 193 neonates. One hundred bootstrap samples were generated to validate the choices of covariates in the models based on their frequencies of selection. This approach validated the models for the oxygen hood and CPAP. The regression coefficients and standard deviations for the CPAP and oxygen hood models were estimated using 200 additional bootstrap samples. A close agreement between stepwise and bootstrap estimates was observed for both models. These results suggest that bootstrap can be useful for validating models for predicting the utilization of health technologies.

Blood Gas Analysis↗

Very fast algorithms for evaluating the stability of ML and Bayesian phylogenetic trees from sequence data.

Evolutionary trees sit at the core of all realistic models describing a set of related sequences, including alignment, homology search, ancestral protein reconstruction and 2D/3D structural change. It is important to assess the stochastic error when estimating a tree, including models using the most realistic likelihood-based optimizations, yet computation times may be many days or weeks. If so, the bootstrap is computationally prohibitive. Here we show that the extremely fast "resampling of estimated log likelihoods" or RELL method behaves well under more general circumstances than previously examined. RELL approximates the bootstrap (BP) proportions of trees better that some bootstrap methods that rely on fast heuristics to search the tree space. The BIC approximation of the Bayesian posterior probability (BPP) of trees is made more accurate by including an additional term related to the determinant of the information matrix (which may also be obtained as a product of gradient or score vectors). Such estimates are shown to be very close to MCMC chain values. Our analysis of mammalian mitochondrial amino acid sequences suggest that when model breakdown occurs, as it typically does for sequences separated by more than a few million years, the BPP values are far too peaked and the real fluctuations in the likelihood of the data are many times larger than expected. Accordingly, several ways to incorporate the bootstrap and other types of direct resampling with MCMC procedures are outlined. Genes evolve by a process which involves some sites following a tree close to, but not identical with, the species tree. It is seen that under such a likelihood model BP (bootstrap proportions) and BPP estimates may still be reasonable estimates of the species tree. Since many of the methods studied are very fast computationally, there is no reason to ignore stochastic error even with the slowest ML or likelihood based methods.

Algorithms↗

On statistical methods for comparison of intrasample morphometric variability: Zalavár revisited.

In studies of morphology, methods for comparing amounts of variability are often important. Three different ways of utilizing determinants of covariance matrices for testing for surplus variability in a hypothesis sample compared to a reference sample are presented: an F-test based on standardized generalized variances, a parametric bootstrap based on draws on Wishart matrices, and a nonparametric bootstrap. The F-test based on standardized generalized variances and the Wishart-based bootstrap are applicable when multivariate normality can be assumed. These methods can be applied with only summary data available. However, the nonparametric bootstrap can be applied with multivariate nonnormally distributed data as well as multivariate normally distributed data, and small sample sizes. Therefore, this method is preferable when raw data are available. Three craniometric samples are used to present the methods. A Hungarian Zalavár sample and an Austrian Berg sample are compared to a Norwegian Oslo sample, the latter employed as reference sample. In agreement with a previous study, it is shown that the Zalavár sample does not represent surplus variability, whereas the Berg sample does represent such a surplus variability.

Multivariate Analysis↗

Phylogeny of the avian family Ciconiidae (storks) based on cytochrome b sequences and DNA-DNA hybridization distances.

This study is a phylogenetic analysis of the avian family Ciconiidae, the storks, based on two molecular data sets: 1065 base pairs of sequence from the mitochondrial cytochrome b gene and a complete matrix of single-copy nuclear DNA-DNA hybridization distances. Sixteen of the nineteen stork species were included in the cytochrome b data matrix, and fifteen in the DNA-DNA hybridization matrix. Both matrices included outgroups from the families Cathartidae (New World vultures) and Threskiornithidae (ibises, spoonbills). Optimal trees based on the two data sets were congruent in those nodes with strong bootstrap support. In the best-fit tree based on DNA-DNA hybridization distances, nodes defining relationships among very recently diverged species had low bootstrap support, while nodes defining more distant relationships had strong bootstrap support. In the optimal trees based on the sequence data, nodes defining relationships among recently diverged species had strong bootstrap support, while nodes defining basal relationships in the family had weak support and were incongruent among analyses. A combinable-component consensus of the best-fit DNA-DNA hybridization tree and a consensus tree based on different analyses of the cytochrome b sequences provide the best estimate of relationships among stork species based on the two data sets.

Animals↗

Early evolutionary relationships among known life forms inferred from elongation factor EF-2/EF-G sequences: phylogenetic coherence and structure of the archaeal domain.

Phylogenies were inferred from both the gene and the protein sequences of the translational elongation factor termed EF-2 (for Archaea and Eukarya) and EF-G (for Bacteria). All treeing methods used (distance-matrix, maximum likelihood, and parsimony), including evolutionary parsimony, support the archaeal tree and disprove the "eocyte tree" (i.e., the polyphyly and paraphyly of the Archaea). Distance-matrix trees derived from both the amino acid and the DNA sequence alignments (first and second codon positions) showed the Archaea to be a monophyletic-holophyletic grouping whose deepest bifurcation divides a Sulfolobus branch from a branch comprising Methanococcus, Halobacterium, and Thermoplasma. Bootstrapped distance-matrix treeing confirmed the monophyly-holophyly of Archaea in 100% of the samples and supported the bifurcation of Archaea into a Sulfolobus branch and a methanogen-halophile branch in 97% of the samples. Similar phylogenies were inferred by maximum likelihood and by maximum (protein and DNA) parsimony. DNA parsimony trees essentially identical to those inferred from first and second codon positions were derived from alternative DNA data sets comprising either the first or the second position of each codon. Bootstrapped DNA parsimony supported the monophyly-holophyly of Archaea in 100% of the bootstrap samples and confirmed the division of Archaea into a Sulfolobus branch and a methanogen-halophile branch in 93% of the bootstrap samples. Distance-matrix and maximum likelihood treeing under the constraint that branch lengths must be consistent with a molecular clock placed the root of the universal tree between the Bacteria and the bifurcation of Archaea and Eukarya. The results support the division of Archaea into the kingdoms Crenarchaeota (corresponding to the Sulfolobus branch and Euryarchaeota). This division was not confirmed by evolutionary parsimony, which identified Halobacterium rather than Sulfolobus as the deepest offspring within the Archaea.

Amino Acid Sequence↗

Evolutionary relationships among the eukaryotic crown taxa taking into account site-to-site rate variation in 18S rRNA.

In this study we constructed a bootstrapped distance tree of 500 small subunit ribosomal RNA sequences from organisms belonging to the so-called crown of eukaryote evolution. Taking into account the substitution rate of the individual nucleotides of the rRNA sequence alignment, our results suggest that (1) animals, true fungi, and choanoflagellates share a common origin: The branch joining these taxa is highly supported by bootstrap analysis (bootstrap support [BS] > 90%), (2) stramenopiles and alveolates are sister groups (BS = 75%), (3) within the alveolates, dinoflagellates and apicomplexans share a common ancestor BS > 95%), while in turn they both share a common origin with the ciliates (BS > 80%), and (4) within the stramenopiles, heterokont algae, hyphochytriomycetes, and oomycetes form a monophyletic grouping well supported by bootstrap analysis (BS > 85%), preceded by the well-supported successive divergence of labyrinthulomycetes and bicosoecids. On the other hand, many evolutionary relationships between crown taxa are still obscure on the basis of 18S rRNA. The branching order between the animal-fungal-choanoflagellates clade and the chlorobionts, the alveolates and stramenopiles, red algae, and several smaller groups of organisms remains largely unresolved.When among-site rate variation is not considered, the inferred tree topologies are inferior to those where the substitution rate spectrum for the 18S rRNA is taken into account. This is primarily indicated by the erroneous branching of fast-evolving sequences. Moreover, when different substitution rates among sites are not considered, the animals no longer appear as a monophyletic grouping in most distance trees.

Animals↗

Applying non-parametric statistical methods to the classical measurements of inclusion complex binding constants.

A study on using non-parametric statistical methods was carried out to calculate the binding constant of an inclusion complex and to estimate its associated uncertainty. First, a correct evaluation of the stoichiometry was carried out in order to ensure an accurate determination of the binding constant. For this purpose, the modified Benesi-Hildelbrand method had been previously applied. Then, four statistical methods (three non-parametric methods: two bootstrap approaches, the jackknife method and a parametric one: Fieller's theorem) were employed in order to compute the binding constant. The results obtained from applying these methods and the combination of the methods: jackknife after bootstrap and bootstrap after jackknife were compared. The best results in terms of accuracy were obtained from the application of a bootstrap method: the resampling residuals approach. These procedures were applied to the inclusion complex 2-hydroxil-propyl-beta-cyclodextrin-2,4-dichloro-phenoxyacetic, which shows photochemically-induced fluorescence.

Binding Sites↗