PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Phylogenetic position of Eimeria antrozoi, a bat coccidium (Apicomplexa: Eimeriidae) and its relationship to morphologically similar Eimeria spp. from bats and rodents based on nuclear 18S and plastid 23S rDNA sequences.

Partial plastid 23S and nuclear 18S rDNA genes were amplified and sequenced from 2 morphologically similar Eimeria species. E. antrozoi from a bat (Antrozous pallidus) and E. arizonensis from deer mice (Peromyscus spp.), as well as some other Eimeria species from bats and rodents. The phylogenetic trees clearly separated E. antrozoi from E. arizonensis. The phylogenies based on plastid 23S rDNA data and combined data of both plastid and nuclear genes grouped 2 bat Eimeria and 3 morphologically similar Eimeria species from rodents into 2 separate clades with high bootstrap support (100%, 3 rodent Eimeria species; 72-97%, 2 bat Eimeria species), which supports E. antrozoi as a valid species. The rodent Eimeria species did not form a monophyletic group. The 2 bat Eimeria species formed a clade with the 3 morphologically similar rodent Eimeria species (E. arizonensis, E. albigulae, E. onychomysis, all from cricetid rodents) with 100% bootstrap support, whereas 2 other rodent Eimeria species (E. nieschulzi, E. falciformis, from murid rodents) formed a separate clade with 100% bootstrap support. This suggests that the 2 Eimeria species from bats might be derived from rodent Eimeria species and may have arisen as a result of lateral host transfer between rodent and bat hosts.

Animals↗

Population pharmacokinetics of an angiotensin II receptor antagonist, telmisartan, in healthy volunteers and hypertensive patients.

OBJECTIVE: To describe the factors affecting pharmacokinetics of telmisartan, an angiotensin II receptor antagonist, a population pharmacokinetic (PPK) model has been developed based upon the data collected from healthy volunteers and hypertensive patients. METHODS: A total of 1566 plasma samples were collected from 20 healthy volunteers and 129 hypertensive patients, together with the demographic background. The data were analyzed by the NONMEM program using two-compartment model with first-order absorption. The robustness of the obtained PPK model was validated by the bootstrapping resampling method. RESULTS: The oral clearance (CL/F) was found to be associated with age, dose and alcohol consumption, but neither related to serum creatinine nor smoking history. The volume of distribution for the central compartment was related to age and dose, and the volume of distribution for the peripheral compartment was related to body weight and gender. The absorption rate constant (Ka) and the absorption lag time were described as function of dose. The CL/F decreased with advanced age. The CL/F decreased and Ka increased with higher dose, reflecting the super-proportional increase in the plasma levels of telmisartan. The AUC and C(max) values predicted by the present PPK model were well consistent with the observed values. The means of parameter estimates obtained with 200 bootstrap replicates were within 95-111% of the final parameter estimates from the original data set. CONCLUSION: A PPK model for telmisartan developed here well described the individual variability and exposure, and robustness of the model has been validated by the bootstrapping method.

Journal Article↗

Assessing performance of prediction rules in machine learning.

INTRODUCTION: An important goal in machine learning is to assess the degree to which prediction rules are robust and replicable, since these rules are used for decision making and for planning follow-up studies. This requires an estimate of a prediction rule's true error rate, a statistic that can be estimated by resampling data. However, there are many possible approaches depending upon whether we draw observations with or without replacement, or sample once, repeatedly, or not at all, and the pros and cons of each are often unclear. This study illustrates and compares different methods for estimating true error with the aim of providing practical guidance to users of machine learning techniques. METHODS: We conducted Monte Carlo simulation studies using four different error estimators: bootstrap, split sample, resubstitution and a direct estimate of true error. Here, 'split sample' refers to a single random partition of the data into a pair of training and test samples, a popular scheme. We used stochastic gradient boosting as a learning algorithm, and considered data from two studies for which the underlying data mechanism was known to be complex: a library of 6000 tripeptide substrates collected for analysis of proteasome inhibition as part of anticancer drug design, and a cardiovascular study involving 600 subjects receiving antiplatelet treatment for acute coronary syndrome. RESULTS: There were important differences in the performance of the various error estimators examined. Error estimators for split sample and resubstitution, while being the most transparent in action and the simplest to apply, did not quantify the performance of prediction rules as accurately as the bootstrap. This was true for both types of study data, despite their highly different nature. CONCLUSIONS: The robustness and reliability of decisions based on analysis of genomics data could, in many cases, be improved by following best practices for prediction error estimation. For this, techniques such as bootstrap should be considered.

Algorithms↗

Determining the significance of scale values from multidimensional scaling profile analysis using a resampling method.

Although multidimensional scaling (MDS) profile analysis is widely used to study individual differences, there is no objective way to evaluate the statistical significance of the estimated scale values. In the present study, a resampling technique (bootstrapping) was used to construct confidence limits for scale values estimated from MDS profile analysis. These bootstrap confidence limits were used, in turn, to evaluate the significance of marker variables of the profiles. The results from analyses of both simulation data and real data suggest that the bootstrap method may be valid and may be used to evaluate hypotheses about the statistical significance of marker variables of MDS profiles.

Career Choice↗

Molecular epidemiology of serotype O foot-and-mouth disease virus isolated from cattle in Ethiopia between 1979-2001.

Partial 1D gene characterization was used to study phylogenetic relationships between 17 serotype O foot-and-mouth disease (FMD) viruses in Ethiopia as well as with other O-type isolates from Eritrea, Kenya, South and West Africa, the Middle East, Asia and South America. A homologous region of 495 bp corresponding to the C-terminus end of the 1D gene was used for phylogenetic analysis. This study described three lineages, viz. African/Middle East-Asia, Cathay and South American. Within lineage I, three topotypes were defined, viz. East and West Africa and the Middle East-Asia together with the South African isolate. The Ethiopian isolates clustered as part of topotype I, the East African topotype. Two clades (based on < 12 % nucleotide difference) A and B were identified within the East African isolates, with clade A being further classified into three significant branches, A1 (80% bootstrap support), A2 (89% bootstrap support) and A3 (94% bootstrap support). Clade B consisted of two Kenyan isolates. Within topotype I, the 17 Ethiopian isolates showed genetic heterogeneity between themselves with sequence differences ranging from 4.6-14 %. Lineage 2 and 3 could be equated to two significant topotypes, viz. Cathay and South America. Comparison of amino acid variability at the immunodominant sites between the vaccine strain (ETH/19/77) and other Ethiopian outbreak isolates revealed variations within these sites. These results encourage further work towards the reassessment of the type O vaccine strain currently being used in Ethiopia to provide protection against field variants of the virus.

Amino Acid Sequence↗

Mediation in experimental and nonexperimental studies: new procedures and recommendations.

Mediation is said to occur when a causal effect of some variable X on an outcome Y is explained by some intervening variable M. The authors recommend that with small to moderate samples, bootstrap methods (B. Efron & R. Tibshirani, 1993) be used to assess mediation. Bootstrap tests are powerful because they detect that the sampling distribution of the mediated effect is skewed away from 0. They argue that R. M. Baron and D. A. Kenny's (1986) recommendation of first testing the X --> Y association for statistical significance should not be a requirement when there is a priori belief that the effect size is small or suppression is a possibility. Empirical examples and computer setups for bootstrap analyses are provided.

Guidelines as Topic↗

[Molecular variability in the commom shrew Sorex araneus L. from European Russia and Siberia inferred from the length polymorphism of DNA regions flanked by short interspersed elements (Inter-SINE PCR) and the relationships between the Moscow and Seliger chromosome races].

Genetic exchange among chromosomal races of the common shrew Sorex araneus and the problem of reproductive barriers have been extensively studied by means of such molecular markers as mtDNA, microsatellites, and allozymes. In the present study, the interpopulation and interracial polymorphism in the common shrew was derived, using fingerprints generated by amplified DNA regions flanked by short interspersed repeats (SINEs)-interSINE PCR (IS-PCR). We used primers, complementary to consensus sequences of two short retroposons: mammalian element MIR and the SOR element from the genome of Sorex araneus. Genetic differentiation among eleven populations of the common shrew from eight chromosome races was estimated. The NP and MJ analyses, as well as multidimensional scaling showed that all samples examined grouped into two main clusters, corresponding to European Russia and Siberia. The bootstrap support of the European Russia cluster in the NJ and MP analyses was respectively 76 and 61%. The bootstrap index for the Siberian cluster was 100% in both analyses; the Tomsk race, included into this cluster, was separated with the bootstrap support of NJ/MP 92/95%.

Animals↗

Using Bagging classifier to predict protein domain structural class.

Classification and prediction of protein domain structural class is one of the important topics in the molecular biology. We introduce the Bagging (Bootstrap aggregating), one of the bootstrap methods, for classifying and predicting protein structural classes. By a bootstrap aggregating procedure, the Bagging can improve a weak classifier, for instance the random tree method, to a significant step towards optimality. In this research, it is demonstrated that the Bagging performed at least as well as LogitBoost and Support vector machines in predicting the structural classes for a given protein domain dataset by 10 cross-validation test, which indicate that the Bagging method is promising and anticipated that it could be potentially further improved on predicting protein structural classes as well as other bio-macromolecular attributes, if the bagging method and other existing methods can be effectively complemented with each other.

Algorithms↗

Estimation of error rates in discriminant analysis with selection of variables.

Accurate estimation of misclassification rates in discriminant analysis with selection of variables by, for example, a stepwise algorithm, is complicated by the large optimistic bias inherent in standard estimators such as those obtained by the resubstitution method. Application of a bootstrap adjustment can reduce the bias of the resubstitution method; however, the bootstrap technique requires the variable selection procedure to be repeated many times and is therefore difficult to compute. In this paper we propose a smoothed estimator that requires relatively little computation and which, on the basis of a Monte Carlo sampling study, is found to perform generally at least as well as the bootstrap method.

Algorithms↗

Simultaneous small-sample multivariate Bernoulli confidence intervals.

A technique based on the bootstrap is presented for assessing the simultaneous confidence level of k small-sample confidence intervals for multivariate Bernoulli marginal frequencies. The small-sample intervals used are those of Clopper and Pearson (1934, Biometrika 26, 404-413) and require iterative computation. To estimate the simultaneous confidence level, the multivariate Bernoulli vectors are resampled via the bootstrap and the Clopper-Pearson intervals recomputed on each pseudosample. The bootstrap estimate is then the proportion of times (computed via Monte Carlo) that all the k intervals computed by resampling contain the original sample frequencies. The technique is applied to single-sample HLA data.

Analysis of Variance↗

Combination of prostate-specific antigen, clinical stage, and Gleason score to predict pathological stage of localized prostate cancer. A multi-institutional update.

OBJECTIVE: To combine the clinical data from 3 academic institutions that serve as centers of excellence for the surgical treatment of clinically localized prostate cancer and develop a multi-institutional model combining serum prostate-specific antigen (PSA) level, clinical stage, and Gleason score to predict pathological stage for men with clinically localized prostate cancer. DESIGN: In this update, we have combined clinical and pathological data for a group of 4133 men treated by several surgeons from 3 major academic urologic centers within the United States. Multinomial log-linear regression was performed for the simultaneous prediction of organ-confined disease, isolated capsular penetration, seminal vesicle involvement, or pelvic lymph node involvement. Bootstrap estimates of the predicted probabilities were used to develop nomograms to predict pathological stage. Additional bootstrap analyses were then obtained to validate the performance of the nomograms. PATIENTS AND SETTINGS: A total of 4133 men who had undergone radical retropubic prostatectomy for clinically localized prostate cancer at The Johns Hopkins Hospital (n=3116), Baylor College of Medicine (n=782), and the University of Michigan School of Medicine (n=235) were enrolled into this study. None of the patients had received preoperative hormonal or radiation therapy. OUTCOME MEASURES: Simultaneous prediction of organ-confined disease, isolated capsular penetration, seminal vesicle involvement, or pelvic lymph node involvement using updated nomograms. RESULTS: Prostate-specific antigen level, TNM clinical stage, and Gleason score contributed significantly to the prediction of pathological stage (P<.001). Bootstrap estimates of the median and 95% confidence interval of the predicted probabilities are presented in the nomograms. For most cells in the nomograms, there is a greater than 25% probability of qualifying for more than one of the pathological stages. In the validation analyses, 72.4% of the time the nomograms correctly predicted the probability of a pathological stage to within 10% (organ-confined disease, 67.3%; isolated capsular penetration, 59.6%; seminal vesicle involvement, 79.6%; pelvic lymph node involvement, 82.9%). CONCLUSIONS: The data represent a multi-institutional modeling and validation of the clinical utility of combining PSA level measurement, clinical stage, and Gleason score to predict pathological stage for a group of men with localized prostate cancer. Clinicians can use these nomograms when counseling individual patients regarding the probability of their tumor being a specific pathological stage; this will enable patients and physicians to make more informed treatment decisions based on the probability of a pathological stage, as well as risk tolerance and the values they place on various potential outcomes.

Decision Support Techniques↗

POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data.

MOTIVATION: Parent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). RESULTS: We demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. AVAILABILITY AND IMPLEMENTATION: The code for this method in Python is available at https://github.com/bystrogenomics/POISE.

Community Detection↗

Predictors of intussusception in young children.

OBJECTIVE: To identify predictors of intussusception in young children. DESIGN: A retrospective cross-sectional study. SETTING AND PATIENTS: A consecutive sample of children younger than 5 years on whom contrast enemas were performed because of suspected intussusception seen at an urban children's hospital from 1990 to 1995. METHODS: We evaluated historical, clinical, and radiographic variables. Variables documented in 75% or more of the medical records and associated with intussusception (P< or =.20) in the univariate analysis were evaluated in a multiple logistic regression analysis. Variables retaining significance (P< or =.05) in the multivariate analysis were considered independent predictors of intussusception. We used bootstrap resampling techniques to validate the multivariate model. RESULTS: Sixty-eight (59%) of the 115 patients had intussusception. Univariate predictors of intussusception included male sex, age younger than 2 years, history of emesis, rectal bleeding, lethargy, abdominal mass, and a highly suggestive abdominal radiograph. In the multivariate analysis, we identified only 4 independent predictors (adjusted odds ratio; 95% confidence interval): a highly suggestive abdominal radiograph (18.3; 4.0-83.1), rectal bleeding (17.3; 2.9-104.0), male sex (6.2; 1.2-32.3), and a history of emesis (13.4; 1.4-126.0). We identified 3 of these 4 variables (all but emesis) as independent predictors in more than 50% of 1000 bootstrap data samples. CONCLUSIONS: Rectal bleeding, a highly suggestive abdominal radiograph, and male sex are variables independently associated with intussusception in a cohort of children suspected of having this diagnosis. Knowledge of these variables may assist in clinical decision making regarding diagnostic and therapeutic interventions.

California↗

Cost-effectiveness of clozapine compared with conventional antipsychotic medication for patients in state hospitals.

BACKGROUND: An open-label, randomized controlled trial compared clozapine with physicians'-choice medications among long-term state hospital inpatients in Connecticut. The goal was to examine clozapine's cost-effectiveness in routine practice for people experiencing lengthy hospitalizations. METHODS: Long-stay patients with schizophrenia in a state hospital were randomly assigned to begin open-label clozapine (n = 138) or to continue receiving conventional antipsychotic medications (n = 89). We interviewed study participants every 4 months for 2 years to assess psychiatric symptoms and functional status, and we collected continuous measures of prescribed medications, service utilization, and other costs. We used both parametric and nonparametric techniques to examine changes in cost and parametric analyses to examine changes in effectiveness. We used bootstrap techniques to estimate incremental cost-effectiveness ratios and create cost-effectiveness acceptability curves. RESULTS: Both groups incurred similar costs during the 2-year study period, with a trend for clozapine to be less costly than usual care in the second study year. Clozapine was more effective than usual care on many but not all measures. With the use of effectiveness measures that favored clozapine (extrapyramidal side effects, disruptiveness), bootstrap techniques indicated that, even when a payer is unwilling to incur any additional cost for gains in effectiveness, the probability that clozapine is more cost-effective than usual care is at least 0.80. These findings were not as evident when outcomes where clozapine was not clearly superior (psychotic symptoms, weight gain) were examined. CONCLUSION: Clozapine demonstrated cost-effectiveness on some but not all measures of effectiveness when the alternative was a range of conventional antipsychotic medications.

Adult↗

A comparison of statistical methods for clustered data analysis with Gaussian error.

We investigate by simulation the properties of four different estimation procedures under a linear model for correlated data with Gaussian error: maximum likelihood based on the normal mixed linear model; generalized estimating equations; a four-stage method, and a bootstrap method that resamples clusters rather than individuals. We pay special attention to the group randomized trials where the number of independent clusters is small, cluster sizes are big, and the correlation within the cluster is weak. We show that for balanced and near balanced data when the number of independent clusters is small (< or = 10), the bootstrap is superior if analysts do not want to impose strong distribution and covariance structure assumptions. Otherwise, ML and four-stage methods are slightly better. All four methods perform well when the number of independent clusters reaches 50.

Bias↗

Confidence intervals for the log-normal mean .

In this paper we conduct a stimulation study to evaluate coverage error, interval width and relative bias of four main methods for the construction of confidence intervals of log-normal means: the naive method; Cox's method; a conservative method; and a parametric bootstrap method. The simulation study finds that the naive method is inappropriate, that Cox's method has the smallest coverage error for moderate and large sample sizes, and that the bootstrap method has the smallest coverage error for small sample sizes. In addition, Cox's method produces the smallest interval width among the three appropriate methods. We also apply the four methods to a real data set to contrast the differences.

Bias↗

Methods for combining rates from several studies.

When several independent groups have conducted studies to estimate a procedure's success rate, it is often of interest to combine the results of these studies in the hopes of obtaining a better estimate for the true unknown success rate of the procedure. In this paper we present two hierarchical methods for estimating the overall rate of success. Both methods take into account the within-study and between-study variation and assume in the first stage that the number of successes within each study follows a binomial distribution given each study's own success rate. They differ, however, in their second stage assumptions. The first method assumes in the second stage that the rates of success from individual studies form a random sample having a constant expected value and variance. Generalized estimating equations (GEE) are then used to estimate the overall rate of success and its variance. The second method assumes in the second stage that the success rates from different studies follow a beta distribution. Both methods use the maximum likelihood approach to derive an estimate for the overall success rate and to construct the corresponding confidence intervals. We also present a two-stage bootstrap approach to estimating a confidence interval for the success rate when the number of studies is small. We then perform a simulation study to compare the two methods. Finally, we illustrate these two methods and obtain bootstrap confidence intervals in a medical example analysing the effectiveness of hyperdynamic therapy for cerebral vasospasm.

Confidence Intervals↗

Approximating the power of Wilcoxon's rank-sum test against shift alternatives.

Three methods of approximating the power of Wilcoxon's rank-sum test against shift alternatives are studied. They are obtained by using a Gaussian assumption, Edgeworth expansion, or bootstrap. It is assumed that a historical data set is available to use in estimating the shape of the distribution. The methods are compared through simulation across several different distributional types. The results indicate that the bootstrap generally gives the most reliable approximation, however the Edgeworth expansion has the practical advantage that a lower bound on the power can be roughly approximated. The methods are illustrated on muscle strength data from patients with osteogenesis imperfecta. Published in 1999 by John Wiley & Sons, Ltd. This is US Government work and is in the public domain in the United States.

Abdominal Muscles↗