PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Robust analysis of a mixed-effect model for a multicenter clinical trial.

We consider a multicenter clinical trial with treatments as fixed effects and centers and residuals as bivariate and univariate random effects, respectively. There exist situations where it is difficult to justify the conventional normality assumptions for the random components. Following Khatri and Patel (Commun. Stat.-Theory Methods 1992, 21, 21-39), we propose the weighted least-squares (WLS) method and two bootstrap methods, percentile and BCa, that are robust to the departure from normality. Through a simulation study, we compare WLS and bootstrap confidence intervals for the treatment difference. While all three methods give confidence intervals with desired coverage rates, the WLS method gives shorter intervals. We also propose a bootstrap method that is robust to outliers. A numerical example is given to illustrate the methodology.

Computer Simulation↗

Genetic diversity and phylogenetic analysis of human immunodeficiency virus type 1 subtypes circulating in French Guiana.

We investigated the characterization of different HIV-1 subtypes present in French Guiana by use of three different methods. Serological methods were used for the initial screening, which were then confirmed by the heteroduplex mobility assay (HMA). The V3 env region was subsequently sequenced for phylogenetic analysis, to confirm the subtype of the samples, and to assign a subtype to samples that gave results that were difficult to interpret or discordant by serology or HMA. A total of 221 HIV-1 seropositive samples were typed; 110 of them were confirmed by HMA and 16 were sequenced. Of the 221 samples tested 210 patients (95%) were found to be infected with subtype B, 10 (4.5%) were infected with subtype A, and one patient was infected with subtype F. Phylogenetic analysis demonstrated that the strains from French Guiana were closely related to the subtype A and B subtypes, and that one strain was closely related to an F subtype (100% bootstrap value). Four strains from French Guiana clustered in the subtype A (99% bootstrap value) and the other strains were associated with subtype B (100% bootstrap value). The geographic position of French Guiana suggested that HIV-1 was probably introduced into the country via several routes, and thus the pattern of the HIV-1 epidemic might evolve in the near future.

AIDS Serodiagnosis↗

Substantial overestimation of standard errors of relative survival rates of cancer patients.

Relative survival rates are among the most commonly reported outcome measures of cancer patients. They are calculated as ratios of observed survival rates and the expected survival rates in the absence of cancer. Standard errors of relative survival rates are commonly calculated by dividing the standard error for absolute survival rates by the expected survival, without taking possible random variation of the latter into account. The aim of this study was to empirically assess the validity of these commonly reported standard errors. Using data from the nationwide Finnish Cancer Registry, the authors calculated 5- and 10-year absolute, expected, and relative survival rates for patients with 25 common forms of cancer in Finland in 1989. The authors used bootstrap analysis to empirically assess the random error of absolute and relative survival rates and then compared the results with conventionally derived estimates of standard errors. The conventional and bootstrap standard errors were closely similar for all estimates of absolute survival. By contrast, the conventional estimates of standard errors of 5- and 10-year relative survival exceeded the bootstrap estimates by up to 17% and 32%, respectively. The authors conclude that conventional derivation may substantially overestimate standard errors for relative survival.

Age Distribution↗

Prediction error estimation: a comparison of resampling methods.

MOTIVATION: In genomic studies, thousands of features are collected on relatively few samples. One of the goals of these studies is to build classifiers to predict the outcome of future observations. There are three inherent steps to this process: feature selection, model selection and prediction assessment. With a focus on prediction assessment, we compare several methods for estimating the 'true' prediction error of a prediction model in the presence of feature selection. RESULTS: For small studies where features are selected from thousands of candidates, the resubstitution and simple split-sample estimates are seriously biased. In these small samples, leave-one-out cross-validation (LOOCV), 10-fold cross-validation (CV) and the .632+ bootstrap have the smallest bias for diagonal discriminant analysis, nearest neighbor and classification trees. LOOCV and 10-fold CV have the smallest bias for linear discriminant analysis. Additionally, LOOCV, 5- and 10-fold CV, and the .632+ bootstrap have the lowest mean square error. The .632+ bootstrap is quite biased in small sample sizes with strong signal-to-noise ratios. Differences in performance among resampling methods are reduced as the number of specimens available increase. SUPPLEMENTARY INFORMATION: A complete compilation of results and R code for simulations and analyses are available in Molinaro et al. (2005) (http://linus.nci.nih.gov/brb/TechReport.htm).

Algorithms↗

An assessment of accuracy, error, and conflict with support values from genome-scale phylogenetic data.

Despite the importance of molecular phylogenetics, few of its assumptions have been tested with real data. It is commonly assumed that nonparametric bootstrap values are an underestimate of the actual support, Bayesian posterior probabilities are an overestimate of the actual support, and among-gene phylogenetic conflict is low. We directly tested these assumptions by using a well-supported yeast reference tree. We found that bootstrap values were not significantly different from accuracy. Bayesian support values were, however, significant overestimates of accuracy but still had low false-positive error rates (0% to 2.8%) at the highest values (>99%). Although we found evidence for a branch-length bias contributing to conflict, there was little evidence for widespread, strongly supported among-gene conflict from bootstraps. The results demonstrate that caution is warranted concerning conclusions of conflict based on the assumption of underestimation for support values in real data.

Models, Genetic↗

Evaluation of old and new tests of heterogeneity in epidemiologic meta-analysis.

The identification of heterogeneity in effects between studies is a key issue in meta-analyses of observational studies, since it is critical for determining whether it is appropriate to pool the individual results into one summary measure. The result of a hypothesis test is often used as the decision criterion. In this paper, the authors use a large simulation study patterned from the key features of five published epidemiologic meta-analyses to investigate the type I error and statistical power of five previously proposed asymptotic homogeneity tests, a parametric bootstrap version of each of the tests, and tau2-bootstrap, a test proposed by the authors. The results show that the asymptotic DerSimonian and Laird Q statistic and the bootstrap versions of the other tests give the correct type I error under the null hypothesis but that all of the tests considered have low statistical power, especially when the number of studies included in the meta-analysis is small (<20). From the point of view of validity, power, and computational ease, the Q statistic is clearly the best choice. The authors found that the performance of all of the tests considered did not depend appreciably upon the value of the pooled odds ratio, both for size and for power. Because tests for heterogeneity will often be underpowered, random effects models can be used routinely, and heterogeneity can be quantified by means of R(I), the proportion of the total variance of the pooled effect measure due to between-study variance, and CV(B), the between-study coefficient of variation.

Epidemiologic Methods↗

Evolutionary relationships of human populations on a global scale.

Using gene frequency data for 29 polymorphic loci (121 alleles), we conducted a phylogenetic analysis of 26 representative populations from around the world by using the neighbor-joining (NJ) method. We also conducted a separate analysis of 15 populations by using data for 33 polymorphic loci. These analyses have shown that the first major split of the phylogenetic tree separates Africans from non-Africans and that this split occurs with a 100% bootstrap probability. The second split separates Caucasian populations from all other non-African populations, and this split is also supported by bootstrap tests. The third major split occurs between Native American populations and the Greater Asians that include East Asians (mongoloids), Pacific Islanders, and Australopapuans (native Australians and Papua New Guineans), but Australopapuans are genetically quite different from the rest of the Greater Asians. The second and third levels of population splitting are quite different from those of the phylogenetic tree obtained by Cavalli-Sforza et al. (1988), where Caucasians, Northeast Asians, and Ameridians from the Northeurasian supercluster and the rest of non-Africans form the Southeast Asian supercluster. One of the major factors that caused the difference between the two trees is that Cavalli-Sforza et al. used unweighted pair-group method with arithmetic mean (UPGMA) in phylogenetic inference, whereas we used the NJ method in which evolutionary rate is allowed to vary among different populations. Bootstrap tests have shown that the UPGMA tree receives poor statistical support whereas the NJ tree is well supported. Implications that the phylogenetic tree obtained has on the current controversy over the out-of-Africa and the multiregional theories of human origins are discussed.

Alleles↗

Phylogenetic relationships of reverse transcriptase and RNase H sequences and aspects of genome structure in the gypsy group of retrotransposons.

The gypsy group of long-terminal-repeat retrotransposons contains elements having the same order of enzyme domains in the pol gene as do retroviruses. Elements in the gypsy group are now known from yeast, filamentous fungi, plants, insects, and echinoids. Reverse transcriptase and RNase H amino acid sequences from elements in the gypsy group--including the recently described SURL elements, TED, Cft1, and Ulysses,--were aligned and analyzed by using parsimony and bootstrapping methods, with plant caulimoviruses and/or retroviruses as outgroups. Clades supported at the 95% level after bootstrapping include (1) 17.6 with 297 and (2) all of the SURL elements together. Other likely relationships supported at lower bootstrap confidence intervals include (1) SURL elements with mag, (2) 17.6 and 297 with TED, and this collective group with 412 and gypsy, (3) Tf1 with Cft1, (4) IFG7 with Del, and (5) all of the retrotransposons in the gypsy group together, to the exclusion of Ty3. In contrast with an earlier analysis, our results place mag within the gypsy group rather than outside of a cluster that contains gypsy group retrotransposons and plant caulimoviruses. Several features of retrotransposon genomes provide further support for some of the aforementioned relationships. The union of SURL elements with mag is supported by the presence of two RNA binding sites in the nucleocapsid protein. Location of the tRNA primer binding site and the presence of a long open reading frame 3' to the pol gene support the 17.6-297-TED-412-gypsy cluster.

Amino Acid Sequence↗

Estimation of component and parameter distributions in spectral analysis.

A method is presented for estimating the distributions of the components and parameters determined with spectral analysis when it is applied to a single data set. The method uses bootstrap resampling to simulate the effect of noise on the computed spectrum and to correct for possible bias in the estimates. A number of bootstrap procedures are reviewed, and one is selected for application to the kinetic analysis of positron emission tomography dynamic studies. The technique is shown to require minimal assumptions about noise in the measurements, and its small sample properties are established through Monte-Carlo simulations. The advantages and limitations of spectral analysis with bootstrap resampling for deriving inferences for tracer kinetic modeling are illustrated through sample analyses of time-activity curves for [18F]fluorodeoxyglucose and [15O]-labeled water.

Algorithms↗

Impact of changing the statistical methodology on hospital and surgeon ranking: the case of the New York State cardiac surgery report card.

BACKGROUND: Risk adjustment is central to the generation of health outcome report cards. It is unclear, however, whether risk adjustment should be based on standard logistic regression, fixed-effects or random-effects modeling. OBJECTIVE: The objective of this study was to determine how robust the New York State (NYS) Coronary Artery Bypass Graft (CABG) Surgery Report Card is to changes in the underlying statistical methodology. METHODS: Retrospective cohort study based on data from the NYS Cardiac Surgery Reporting System on all patient undergoing isolated CABG surgery in NYS and who were discharged between 1997 and 1999 (51,750 patients). Using the same risk factors as in the NYS models, fixed-effects and random-effects models were fitted to the NYS data. Quality outliers were identified using 1) the ratio of observed-to-expected mortality rates (O/E ratio) and confidence intervals (CIs) calculated using both parametric (Poisson distribution) and nonparametric (bootstrapping) techniques; and 2) shrinkage estimators. RESULTS: At the surgeon level, the standard logistic regression model, the fixed-effects model, and the fixed-effects component of the random-effects model demonstrated near-perfect agreement on the identity of quality outliers using a quality indicator based on the O/E ratio and the Poisson distribution. Shrinkage estimators identified the fewest outliers, whereas the O/E ratios with bootstrap CI identified the greatest number of outliers. The results were similar for hospitals, except that the fixed-effects model identified more outliers than either the NYS model or the fixed-effects component of the random-effects model. CONCLUSION: Shrinkage estimators based on random-effects models are slightly more conservative in identifying quality outliers compared with the traditional approach based on fixed-effects modeling and standard regression. Explicitly modeling surgeon provider effect (fixed-effects and random-effects models) did not significantly alter the distribution of quality outliers when compared with standard logistic regression (which does not model provider effect). Compared with the standard parametric approach, the use of a bootstrap approach to construct 95% confidence interval around the O/E ratio resulted in more providers being identified as quality outliers.

Benchmarking↗

16S rDNA sequence analysis of environmental Bdellovibrio-and-like organisms (BALO) reveals extensive diversity.

Bdellovibrio-and-like organisms (BALO) are Gram-negative, predatory bacteria that inhabit terrestrial, freshwater and salt-water environments. Historically, these organisms have been classified together despite documented genetic differences between isolates. The genetic diversity of these microbes was assessed by sequencing the 16S rRNA gene. Primers that selectively amplify predator 16S rDNA, and not contaminating prey DNA, were utilized to study 17 freshwater and terrestrial and nine salt-water BALO isolates. When the 16S rDNA sequences were compared with representatives of other bacterial classes, 25 of the 26 BALO isolates clustered into two groups. One group, supported 100% by bootstrap analysis, included all of the Bdellovibrio bacteriovorus isolates. Each member of this group was isolated from either a freshwater or terrestrial source. The genetic distance between these isolates was less than 12%. The other group, supported 94% by bootstrap analysis, includes Bacteriovorax starrii, Bacteriovorax stolpii and the salt-water isolates. The salt-water isolates form a subgroup (83% by bootstrap) and differ within the subgroup by less than 110%. This observation implies that the salt-water isolates arose from Bacteriovorax progenitors. The difference between isolates in different clades is over 17%, a quantity similar to differences between bacterial species in different classes. However, both the Bdellovibrio and Bacteriovorax clades were closest to other representatives of the delta-Proteobacteria using maximum-likelihood. One freshwater isolate, James Island, was distinct from all other BALO (> 19%), but differed from Pseudomonas putida, a member of the gamma-Proteobacteria, by only 3%. Thus, by 16S rDNA sequence analysis, the BALO appear to have multiple origins, contrary to the unified taxonomic grouping based on morphology and natural history. These observations are consistent with the need to review and revise the taxonomy of these organisms.

Base Sequence↗

Dissection of phylogenetic relationships among 19 rapidly growing Mycobacterium species by 16S rRNA, hsp65, sodA, recA and rpoB gene sequencing.

The current classification of non-pigmented and late-pigmenting rapidly growing mycobacteria (RGM) capable of producing disease in humans and animals consists primarily of three groups, the Mycobacterium fortuitum group, the Mycobacterium chelonae-abscessus group and the Mycobacterium smegmatis group. Since 1995, eight emerging species have been tentatively assigned to these groups on the basis of their phenotypic characters and 16S rRNA gene sequence, resulting in confusing taxonomy. In order to assess further taxonomic relationships among RGM, complete sequences of the 16S rRNA gene (1483-1489 bp), rpoB (3486-3495 bp) and recA (1041-1056 bp) and partial sequences of hsp65 (420 bp) and sodA (441 bp) were determined in 19 species of RGM. Phylogenetic trees based upon each gene sequence, those based on the combined dataset of the five gene sequences and one based on the combined dataset of the rpoB and recA gene sequences were then compared using the neighbour-joining, maximum-parsimony and maximum-likelihood methods after using the incongruence length difference test. Combined datasets of the five gene sequences comprising nearly 7000 bp and of the rpoB+recA gene sequences comprising nearly 4600 bp distinguished six phylogenetic groups, the M. chelonae-abscessus group, the Mycobacterium mucogenicum group, the M. fortuitum group, the Mycobacterium mageritense group, the Mycobacterium wolinskyi group and the M. smegmatis group, respectively comprising four, three, eight, one, one and two species. The two protein-encoding genes rpoB and recA improved meaningfully the bootstrap values at the nodes of the different groups. The species M. mucogenicum, M. mageritense and M. wolinskyi formed new groups separated from the M. chelonae-abscessus, M. fortuitum and M. smegmatis groups, respectively. The M. mucogenicum group was well delineated, in contrast to the M. mageritense and M. wolinskyi groups. For phylogenetic organizations derived from the hsp65 and sodA gene sequences, the bootstrap values at the nodes of a few clusters were <70 %. In contrast, phylogenetic organizations obtained from the 16S rRNA, rpoB and recA genes were globally similar to that inferred from combined datasets, indicating that the rpoB and recA genes appeared to be useful tools in addition to the 16S rRNA gene for the investigation of evolutionary relationships among RGM species. Moreover, rpoB gene sequence analysis yielded bootstrap values higher than those observed with recA and 16S rRNA genes. Also, molecular signatures in the rpoB and 16S rRNA genes of the M. mucogenicum group showed that it was a sister group of the M. chelonae-abscessus group. In this group, M. mucogenicum ATCC 49650(T) was clearly distinguished from M. mucogenicum ATCC 49649 with regard to analysis of the five gene sequences. This was in agreement with phenotypic and biochemical characteristics and suggested that these strains are representatives of two closely related, albeit distinct species.

Bacterial Proteins↗

Optimal modeling of corneal surfaces with Zernike polynomials.

Zernike polynomials are often used as an expansion of corneal height data and for analysis of optical wavefronts. Accurate modeling of corneal surfaces with Zernike polynomials involves selecting the order of the polynomial expansion based on the measured data. We have compared the efficacy of various classical model order selection techniques that can be utilized for this purpose, and propose an approach based on the bootstrap. First, it is shown in simulations that the bootstrap method outperforms the classical model order selection techniques. Then, it is proved that the bootstrap technique is the most appropriate method in the context of fitting Zernike polynomials to corneal elevation data, allowing objective selection of the optimal number of Zernike terms. The process of optimal fitting of Zernike polynomials to corneal elevation data is discussed and examples are given for normal corneas and for abnormal corneas with significant distortion. The optimal model order varies as a function of the diameter of the cornea.

Astigmatism↗

Robust estimation of bioaffinity assay fluorescence signals.

In this paper, the challenging problem of robust mean-signal estimation of a single-step microparticle bioaffinity assay is investigated. For this purpose, a density estimation-based robust algorithm (DER) was developed. The DER algorithm was comparatively evaluated with four other parameter estimation methods (mean value, median filtering, least square estimation, Welsch robust m-estimator). Two important questions were raised and investigated: 1) Which of the five methods can robustly estimate the mean bioaffinity signal? and 2) How many microparticles need to be measured in order to obtain an accurate estimate of the mean signal value? To answer the questions, bootstrap and coefficient of variation (CV) analyses were performed. In the CV analysis, the DER algorithm gave the best results: The CV ranged from 0.8% to 4.9% when the number of microparticles used for the mean signal estimation varied from 800 to 30. In the bootstrap analysis of the standard error, the DER algorithm had the smallest variance. As a conclusion, it can be underlined that: 1) of all methods tested, the DER algorithm gave the most consistent and reproducible results according to the bootstrap and CV analysis; 2) using the DER algorithm accurate estimates could be calculated based on 80-100 particles, corresponding to a typical assay measurement time of 1 min; and 3) the investigated bioaffinity signals contained a large number of outliers (observations that severely deviate from the majority of data) and therefore robust techniques were necessary for the mean signal estimation tasks.

Algorithms↗

Testing for conditional multiple marginal independence.

Survey respondents are often prompted to pick any number of responses from a set of possible responses. Categorical variables that summarize this kind of data are called pick any/c variables. Counts from surveys that contain a pick any/c variable along with a group variable (r levels) and stratification variable (q levels) can be marginally summarized into an r x c x q contingency table. A question that may naturally arise from this setup is to determine if the group and pick any/c variable are marginally independent given the stratification variable. A test for conditional multiple marginal independence (CMMI) can be used to answer this question. Since subjects may pick any number out of c possible responses, the Cochran (1954, Biometrics 10, 417-451) and Mantel and Haenszel (1959, Journal of the National Cancer Institute 22, 719-748) tests cannot be used directly because they assume that units in the contingency table are independent of each other. Therefore, new testing methods are developed. Cochran's test statistic is extended to r x 2 x q tables, and a modified version of this statistic is proposed to test CMMI. Its sampling distribution can be approximated through bootstrapping. Other CMMI testing methods discussed are bootstrap p-value combination methods and Bonferroni adjustments. Simulation findings suggest that the proposed bootstrap procedures and the Bonferroni adjustments consistently hold the correct size and provide power against various alternatives.

Adult↗

Estimation in Bayesian disease mapping.

Recent work on Bayesian inference of disease mapping models discusses the advantages of the fully Bayesian (FB) approach over its empirical Bayes (EB) counterpart, suggesting that FB posterior standard deviations of small-area relative risks are more reflective of the uncertainty associated with the relative risk estimation than counterparts based on EB inference, since the latter fail to account for the variability in the estimation of the hyperparameters. In this article, an EB bootstrap methodology for relative risk inference with accurate parametric EB confidence intervals is developed, illustrated, and contrasted with the hyperprior Bayes. We elucidate the close connection between the EB bootstrap methodology and hyperprior Bayes, present a comparison between FB inference via hybrid Markov chain Monte Carlo and EB inference via penalized quasi-likelihood, and illustrate the ability of parametric bootstrap procedures to adjust for the undercoverage in the "naive" EB interval estimates. We discuss the important roles that FB and EB methods play in risk inference, map interpretation, and real-life applications. The work is motivated by a recent analysis of small-area infant mortality rates in the province of British Columbia in Canada.

Algorithms↗

The history of a nearctic colonization: molecular phylogenetics and biogeography of the Nearctic toads (Bufo).

Previous hypotheses of phylogenetic relationships among Nearctic toads (Bufonidae) and their congeners suggest contradictory biogeographic histories. These hypotheses argue that the Nearctic Bufo are: (1) a polyphyletic assemblage resulting from multiple colonizations from Africa; (2) a paraphyletic assemblage resulting from a single colonization event from South America with subsequent dispersal into Eurasia; or (3) a monophyletic group derived from the Neotropics. We obtained approximately 2.5 kb of mitochondrial DNA sequence data for the 12S, 16S, and intervening valine tRNA gene from 82 individuals representing 56 species and used parametric bootstrapping to test hypotheses of the biogeographic history of the Nearctic Bufo. We find that the Nearctic species of Bufo are monophyletic and nested within a large clade of New World Bufo to the exclusion of Eurasian and African taxa. This suggests that Nearctic Bufo result from a single colonization from the Neotropics. More generally, we demonstrate the utility of parametric bootstrapping for testing alternative biogeographic hypotheses. Through parametric bootstrapping, we refute several previously published biogeographic hypotheses regarding Bufo. These previous studies may have been influenced by homoplasy in osteological characters. Given the Neotropical origin for Nearctic Bufo, we examine current distributional patterns to assess whether the Nearctic-Neotropical boundary is a broad transition zone or a narrow boundary. We also survey fossil and paleogeographic evidence to examine potential Tertiary and Cretaceous dispersal routes, including the Paleocene Isthmian Link, the Antillean and Aves Ridges, and the current Central American Land Bridge, that may have allowed colonization of the Nearctic.

Animals↗

Quantification of variability and uncertainty using mixture distributions: evaluation of sample size, mixing weights, and separation between components.

Variability is the heterogeneity of values within a population. Uncertainty refers to lack of knowledge regarding the true value of a quantity. Mixture distributions have the potential to improve the goodness of fit to data sets not adequately described by a single parametric distribution. Uncertainty due to random sampling error in statistics of interests can be estimated based upon bootstrap simulation. In order to evaluate the robustness of using mixture distribution as a basis for estimating both variability and uncertainty, 108 synthetic data sets generated from selected population mixture log-normal distributions were investigated, and properties of variability and uncertainty estimates were evaluated with respect to variation in sample size, mixing weight, and separation between components of mixtures. Furthermore, mixture distributions were compared with single-component distributions. Findings include: (1). mixing weight influences the stability of variability and uncertainty estimates; (2). bootstrap simulation results tend to be more stable for larger sample sizes; (3). when two components are well separated, the stability of bootstrap simulation is improved; however, a larger degree of uncertainty arises regarding the percentiles coinciding with the separated region; (4). when two components are not well separated, a single distribution may often be a better choice because it has fewer parameters and better numerical stability; and (5). dependencies exist in sampling distributions of parameters of mixtures and are influenced by the amount of separation between the components. An emission factor case study based upon NO(x) emissions from coal-fired tangential boilers is used to illustrate the application of the approach.

Journal Article↗