PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Personality traits in captive lion-tailed macaques (Macaca silenus).

Personality influences an individual's perception of a situation and orchestrates behavioral responses. It is an important factor in elucidating variation in behavior both within and between species. The major focus of this research was to test a method that differs from those used in most previous personality studies, while investigating the personality traits of 52 captive lion-tailed macaques from four zoos. In this study, data from behavioral observations, a P-type principal components analysis (PCA), and bootstrapped confidence intervals as criteria for judging the significance of factor loadings were used rather than subjective ratings, R-type factor analyses, and arbitrary rules of thumb to determine significance. We investigated the relationships among individual component scores and sex, hormonal status, and dominance rank (controlling for age and social group) using a multiple regression analysis with bootstrapped confidence intervals. Three personality dimensions emerged from this analysis: Component 1 contained Extraversion-like behaviors related to sociability and affiliativeness. The higher mean Component score for females suggests that they are more "extraverted" than males. Only agonistic behaviors were significantly related to component 2. High-ranking individuals exhibited higher mean Component 2 scores than mid- or low-ranked individuals. Bold and cautious behaviors both loaded positively on Component 3, suggesting a dimension related to curiosity. The mean Component 3 score for females was higher than the mean score for males. The method used in this study should facilitate intraspecific and general interspecific comparisons. Developing a standardized trait term list that is applicable to many species, and collecting trait term data in the same manner and concurrent with behavioral observations (and physiologic measures when feasible) could prove useful in primate research and should be explored.

Animals↗

Reconstructing the evolutionary history of the Lorisidae using morphological, molecular, and geological data.

Major aspects of lorisid phylogeny and systematics remain unresolved, despite several studies (involving morphology, histology, karyology, immunology, and DNA sequencing) aimed at elucidating them. Our study is the first to investigate the evolution of this enigmatic group using molecular and morphological data for all four well-established genera: Arctocebus, Loris, Nycticebus, and Perodicticus. Data sets consisting of 386 bp of 12S rRNA, 535 bp of 16S rRNA, and 36 craniodental characters were analyzed separately and in combination, using maximum parsimony and maximum likelihood. Outgroups, consisting of two galagid taxa (Otolemur and Galagoides) and a lemuroid (Microcebus), were also varied. The morphological data set yielded a paraphyletic lorisid clade with the robust Nycticebus and Perodicticus grouped as sister taxa, and the galagids allied with Arctocebus. All molecular analyses maximum parsimony (MP) or maximum likelihood (ML) which included Microcebus as an outgroup rendered a paraphyletic lorisid clade, with one exception: the 12S + 16S data set analyzed with ML. The position of the galagids in these paraphyletic topologies was inconsistent, however, and bootstrap values were low. Exclusion of Microcebus generated a monophyletic Lorisidae with Asian and African subclades; bootstrap values for all three clades in the total evidence tree were over 90%. We estimated mean genetic distances for lemuroids vs. lorisoids, lorisids vs. galagids, and Asian vs. African lorisids as a guide to relative divergence times. We present information regarding a temporary land bridge that linked the two now widely separated regions inhabited by lorisids that may explain their distribution. Finally, we make taxonomic recommendations based on our results.

Animals↗

Comparison of confidence intervals for adjusted attributable risk estimates under multinomial sampling.

The epidemiologic concept of the adjusted attributable risk is a useful approach to quantitatively describe the importance of risk factors on the population level. It measures the proportional reduction in disease probability when a risk factor is eliminated from the population, accounting for effects of confounding and effect-modification by nuisance variables. The computation of asymptotic variance estimates for estimates of the adjusted attributable risk is often done by applying the delta method. Investigations on the delta method have shown, however, that the delta method generally tends to underestimate the standard error, leading to biased confidence intervals. We compare confidence intervals for the adjusted attributable risk derived by applying computer intensive methods like the bootstrap or jackknife to confidence intervals based on asymptotic variance estimates using an extensive Monte Carlo simulation and within a real data example from a cohort study in cardiovascular disease epidemiology. Our results show that confidence intervals based on bootstrap and jackknife methods outperform intervals based on asymptotic theory. Best variants of computer intensive confidence intervals are indicated for different situations.

Adult↗

Factors underlying variation in spontaneous and clastogen-induced sister chromatid exchanges and chromosome breakage frequencies.

The latent "factors" influencing spontaneous and clastogen-induced genetic damage, measured by rates of sister chromatid exchange (SCE) and chromosome breakage (CB), were investigated in a small sample of 20 unrelated, healthy individuals. The covariation of spontaneous and clastogen-induced (bleomycin [BLM], streptonigrin [SN], mitomycin-C [MMC], 4-nitroquinoline-1-oxide [4NQO]) SCEs and CBs was analyzed by maximum-likelihood factor analysis. A single-factor model resulted in large standardized regression coefficients of measured variables on the factor for spontaneous and BLM- and SN-induced SCE frequencies, and a modest regression coefficient for MMC-induced SCEs. A two-factor model, after varimax rotation, yielded one factor strongly associated with spontaneous and BLM- and SN-induced SCE frequencies, and a second factor associated with spontaneous and BLM- and SN-induced CBs. A bootstrap analysis of this data set indicated the statistical significance of one regression coefficient (i.e., P less than or equal to 0.05) and borderline significance (0.07 less than or equal to P less than or equal to 0.11) of three other regression coefficients on the first factor, to be interpreted as an effector of SCE frequencies. However, for the second factor, none of the bootstrapped regression coefficients was significant (P greater than 0.22). Due to the modest sample utilized in this study, the validity of this model should be further explored using additional, larger data sets.

4-Nitroquinoline-1-oxide↗

Identifying SNPs predictive of phenotype using random forests.

There has been a great interest and a few successes in the identification of complex disease susceptibility genes in recent years. Association studies, where a large number of single-nucleotide polymorphisms (SNPs) are typed in a sample of cases and controls to determine which genes are associated with a specific disease, provide a powerful approach for complex disease gene mapping. Genes of interest in those studies may contain large numbers of SNPs that classical statistical methods cannot handle simultaneously without requiring prohibitively large sample sizes. By contrast, high-dimensional nonparametric methods thrive on large numbers of predictors. This work explores the application of one such method, random forests, to the problem of identifying SNPs predictive of the phenotype in the case-control study design. A random forest is a collection of classification trees grown on bootstrap samples of observations, using a random subset of predictors to define the best split at each node. The observations left out of the bootstrap samples are used to estimate prediction error. The importance of a predictor is quantified by the increase in misclassification occurring when the values of the predictor are randomly permuted. We extend the concept of importance to pairs of predictors, to capture joint effects, and we explore the behavior of importance measures over a range of two-locus disease models in the presence of a varying number of SNPs unassociated with the phenotype. We illustrate the application of random forests with a data set of asthma cases and unaffected controls genotyped at 42 SNPs in ADAM33, a previously identified asthma susceptibility gene. SNPs and SNP pairs highly associated with asthma tend to have the highest importance index value, but predictive importance and association do not always coincide.

Case-Control Studies↗

Reduction of selection bias in genomewide studies by resampling.

The accuracy of gene localization, the reliability of locus-specific effect estimates, and the ability to replicate initial claims of linkage and/or association have emerged as major methodological concerns in genomewide studies of complex diseases and quantitative traits. To address the issue of multiple comparisons inherent in genomewide studies, the use of stringent criteria for assessing statistical significance has been generally acknowledged as a strategy to control type I error. However, the application of genomewide significance criteria does not take account of the selection bias introduced into parameter estimates, e.g., estimates of locus-specific effect size of disease/trait loci. Some have argued that reliable locus-specific parameter estimates can only be obtained in an independent sample. In this report, we examine statistical resampling techniques, including cross-validation and the bootstrap, applied to the initial sample to improve the estimation of locus-specific effects. We compare them with the naive method in which all data are used for both hypothesis testing and parameter estimation, as well as with the split-sample approach in which part of the data are reserved for estimation. Upward bias of the naive estimator and inadequacy of the split-sample approach are derived analytically under a simple quantitative trait model. Simulation studies of the resampling methods are performed for both the simple model and a more realistic genomewide linkage analysis. Our results suggest that cross-validation and bootstrap methods can substantially reduce the estimation bias, especially when the effect size is small or there is no genetic effect.

Algorithms↗

A comparison of methods for intermediate fine mapping.

The arrival of highly dense genetic maps at low cost has geared the focus of linkage analysis studies toward developing methods for placing putative trait loci in narrow regions with high confidence. This shift has led to a new analytic scheme that expands the traditional two-stage protocol of preliminary genome scan followed by fine mapping through inserting a new stage in between the two. The goal of this new "intermediate" fine mapping stage is to isolate disease loci to narrow intervals with high confidence so that association studies can be more focused, efficient, and cost-effective. In this paper, we compared and contrasted five methods that can be used for performing this intermediate step. These methods are: the lod support approach, the generalized estimating equations (GEE) method, the confidence set inference (CSI) procedure, and two bootstrap methods. We compared these methods in terms of the coverage probability and precision of localization of the resulting intervals. Results from a simulation study considering several two-locus models demonstrated that the two bootstrap methods yield intervals with approximately correct coverage. On the other hand, the 1-lod support intervals, and those produced by the GEE method, tend to significantly undercover the trait locus, while the regions obtained by the CSI incline to overcover the gene position. When the observed coverage of the confidence intervals produced by all the methods was held to be the same, those obtained through the CSI procedure displayed a higher ability to localize loci, especially when these loci have a minor contribution to the trait and when the amount of data available for the analysis is relatively small. However, with very large sample sizes, lod support intervals emerged as a winner. Application of the methods to the data from the Arthritis Research Campaign National Repository led to intervals containing the position of a known trait locus for all methods, with the greatest precision achieved by the CSI.

Arthritis, Rheumatoid↗

The analysis of incomplete cost data due to dropout.

Incomplete data due to premature withdrawal (dropout) constitute a serious problem in prospective economic evaluations that has received only little attention to date. The aim of this simulation study was to investigate how standard methods for dealing with incomplete data perform when applied to cost data with various distributions and various types of dropout. Selected methods included the product-limit estimator of Lin et al. the expectation maximisation (EM-) algorithm, several types of multiple imputation (MI) and various simple methods like complete case analysis and mean imputation. Almost all methods were unbiased in the case of dropout completely at random (DCAR), but only the product-limit estimator, the EM-algorithm and the MI approaches provided adequate estimates of the standard error (SE). The best estimates of the mean and SE for dropout at random (DAR) were provided by the bootstrap EM-algorithm, MI regression and MI Monte Carlo Markov chain. These methods were able to deal with skewed cost data in combination with DAR and only became biased when costs also included the costs of expensive events. None of the methods were able to deal adequately with informative dropout. In conclusion, the EM-algorithm with bootstrap, MI regression and MI MCMC are robust to the multivariate normal assumption and are the preferred methods for the analysis of incomplete cost data when the assumption of DCAR is not justified.

Algorithms↗

Choosing between predictors of fractures.

The identification of those at highest risk of osteoporotic fractures is a clinical goal that requires appropriate statistical comparisons of potential predictors of fractures. This article provides a formal approach of comparing individual predictors (e.g., bone mass at one site vs bone mass at another), or sets of predictors (e.g., bone mass vs other risk factors), and contrasts newer methods, such as bootstrapping, to receiver-operating-characteristics (ROC) curves, which have been previously used. The advantages of the bootstrapping approach are illustrated using time-to-fracture data from a published study demonstrating the use of baseline bone mass measurements in the prediction of fractures in 521 subjects with variable lengths of follow-up, extending to 12.5 years. Bone mineral density (BMD) was shown to be significantly better than bone mineral content (BMD) in predicting fractures in free-living subjects, but not in retirement-community subjects. Bone mineral apparent density (BMAD) was also compared with BMC and BMD and shown not to improve fracture prediction in these subjects.

Bone Density↗

Population pharmacokinetic analysis of pegylated human erythropoietin in rats.

The purpose of this study was to model the pharmacokinetics of the pegylated human erythropoietin (PEG-EPO) after single-dose administration in rats, and to evaluate the influence of weight, sex, and pregnancy status on the pharmacokinetic parameters. A total of 436 serum concentrations from 193 Sprague-Dawley rats were obtained from four pharmacokinetic/toxicokinetic studies, in which a single dose of PEG-EPO was administered by the intravenous (i.v.; dose range: 2.5 to 500 microg/kg) and subcutaneous (s.c.; dose range: 12.5 to 500 microg/kg) route. Pharmacokinetic analysis was performed using nonlinear mixed effect modeling (NONMEM V software) to determine the population mean of pharmacokinetic parameters and the variances of the interindividual random effects. The effect of weight, sex, and pregnancy status on the pharmacokinetic parameters was evaluated by forward inclusion and backward elimination process, using the likelihood ratio test. Nonparametric bootstrap analysis was employed as an internal model evaluation technique to qualify the model developed. An open two-compartment model with linear elimination from the central compartment, a first-order absorption with lag time characterized the serum concentration-time profiles of PEG-EPO after i.v. and s.c. administration. For a male rat of 0.24 kg, the average CL, Vc, Q, Vp, Ka, Tlag, and F was estimated to be 0.728 mL/h, 15.8 mL, 0.373 mL/h, 6.99 mL, 0.0618 h(-1), 3.13 h, and 48.8%, respectively. A twofold increase in weight corresponded with a 170 and 238% increase in CL and Vc, respectively. In female rats, Vp was reduced by 11%, whereas F was increased by 15%. No effect of pregnancy status on any of the parameters could be identified. The interindividual variability in CL, Vc, Vp, Ka, and F was estimated at 10.7, 14.7, 16.6, 11.0, and 13.6%, respectively. Nonparametric bootstrap analysis confirmed the accuracy and the precision of the NONMEM parameter estimates. A population pharmacokinetic approach was used to integrate the knowledge gathered from several pharmacokinetic/toxicokinetic studies in rats. The pharmacokinetics of PEG-EPO in the rat was successfully modeled using a two-compartmental model with a linear elimination from the central compartment and a first-order absorption process with lag time. Weight and sex, but not pregnancy status, were identified as covariates of interest during preclinical development. The population pharmacokinetic model developed will be further used for the purpose of interspecies scaling and PK/PD modeling.

Animals↗

Psychological Capital and Perceived Stress in Nurses: The Mediating Role of Need for Recovery and Recovery Experiences.

AIM: To examine whether recovery experiences and need for recovery mediate the association between psychological capital and perceived stress in nurses. DESIGN: Cross-sectional online survey. METHODS: A total of 184 nurses currently practicing in France completed an anonymous online questionnaire administered in January 2023. Participants completed the French versions of the Psychological Capital Questionnaire, the Recovery Experience Scale, the Need for Recovery Scale, and the Perceived Stress Scale. Two mediation models were estimated using ordinary least squares regression with percentile bootstrap inference for the indirect effects. RESULTS: Psychological capital was positively associated with recovery experiences and negatively associated with need for recovery. Need for recovery was strongly and positively associated with perceived stress, while psychological capital showed no direct association with perceived stress. The indirect effect of psychological capital on perceived stress through need for recovery was statistically detectable, whereas a complementary indirect effect through recovery experiences was in the expected direction but did not meet the bootstrap criterion for statistical significance. CONCLUSION: In this cross-sectional sample, psychological capital was associated with lower perceived stress primarily through its association with reduced need for recovery rather than through a direct association with stress. Interventions intended to protect nurse well-being should therefore be conceived as combining individual resource-building with organizational architectures that enable effective recovery. IMPLICATION FOR NURSING PRACTICE: Psychological capital should be seen as one element in integrated interventions rather than a stand-alone solution to nursing stress. Brief recovery-focused interventions for nurses need to be combined with scheduling practices and organizational policies that protect rest and make recovery feasible, a configuration that appears more promising than psychological capital training alone. REPORTING METHOD: The study followed the STROBE Statement for the reporting of cross-sectional studies. NO PATIENT OR PUBLIC CONTRIBUTION: This study focused on nurses as study participants. No patient or public stakeholder was involved in the design, conduct, or interpretation of the research, since the research question concerns occupational psychological resources rather than clinical practice or care delivery.

Humans↗

An approximate Bayesian risk-analysis for the gastro-intestinal safety of ibuprofen.

PURPOSE: Although several studies on ibuprofen and its gastro-intestinal (GI) risk have been reported, the dose-response relationship was not clear due to the lack of information regarding high-dose exposure. Analysis using Bayesian methods is appropriate whenever data are sparse, although such methods are not easily implemented. METHODS: A retrospective cohort study to assess this dose-response relationship was carried out using a record linkage database. A Bayesian risk-analysis was conducted using the Bayesian bootstrap approximation. Risks of GI events at different dose levels were compared using the posterior distributions and the number of events predicted to occur in the future was estimated. Risk factors such as age, gender and co-morbidity were adjusted for in the analysis. This approximation was compared with the full Bayesian approach using the usual but more computer-intensive tool of Markov Chain Monte Carlo simulation. RESULTS: There were 1, 5 and 10 complicated GI events during exposure to high, medium and low dose ibuprofen with 0.2, 1.8 and 7.0 thousand person-years (PY) exposure, respectively. After adjusting for other risk factors the relative risks of high versus low and medium versus low doses were 6.3 (95% CI = 0.21, 24.17) and 2.5 (95% CI = 0.71, 5.85), respectively. Using the approximate Bayesian method prediction of the number of events in a population of females aged 50-59 with no previous medical problems with 1000 PY drug exposure showed that the estimated probability of having more than five events was 0.048 for the medium-dose group and 0.14 for the high-dose group. CONCLUSIONS: High dose ibuprofen appears to have a considerably greater risk of having a larger number of adverse GI events than a medium dose. The approximate Bayesian bootstrap method was demonstrated to be a robust and easily implemented alternative to the full Bayesian approach to risk analysis whenever data are sparse.

Anti-Inflammatory Agents, Non-Steroidal↗

Survival curve estimation with partial non-random exposure information.

The objective of this paper is to estimate survival curves for two different exposure groups when the exposure group is not known for all observations, and the data is subject to left truncation and right censoring. The situation we consider is when the probability that the exposure group is missing may depend on whether the observation is censored or uncensored, in which case the exposure is not missing at random. The problem was motivated by a study of Alzheimer's disease to estimate the distribution of ages at diagnosis for individuals with and without an apolipoprotein E4 allele (the exposure group). Genotyping for this risk factor was incomplete and performed more frequently on the cases of Alzheimer's disease (the uncensored observations) than the censored observations. The survival curves are estimated in discrete time using an EM algorithm. A bootstrapping procedure is proposed that guarantees each bootstrap sample has the same proportion of observations with missing exposure. A simulation is performed to evaluate the bias of the estimators and to investigate design and efficiency issues. The methods are applied to the Alzheimer's disease study.

Adult↗

Improved confidence intervals for the sensitivity at a fixed level of specificity of a continuous-scale diagnostic test.

For a continuous-scale diagnostic test, it is of interest to construct a confidence interval for the sensitivity of the diagnostic test at the cut-off that yields a predetermined level of its specificity (for example, 80, 90 or 95 per cent). In this paper we propose two new intervals for the sensitivity of a continuous-scale diagnostic test at a fixed level of specificity. We then conduct simulation studies to compare the relative performance of these two intervals with the best existing BCa bootstrap interval, proposed by Platt et al. Our simulation results show that the newly proposed intervals are better than the BCa bootstrap interval in terms of coverage accuracy and interval length.

Computer Simulation↗

Parametric modelling of cost data in medical studies.

The cost of medical resources used is often recorded for each patient in clinical studies in order to inform decision-making. Although cost data are generally skewed to the right, interest is in making inferences about the population mean cost. Common methods for non-normal data, such as data transformation, assuming asymptotic normality of the sample mean or non-parametric bootstrapping, are not ideal. This paper describes possible parametric models for analysing cost data. Four example data sets are considered, which have different sample sizes and degrees of skewness. Normal, gamma, log-normal, and log-logistic distributions are fitted, together with three-parameter versions of the latter three distributions. Maximum likelihood estimates of the population mean are found; confidence intervals are derived by a parametric BC(a) bootstrap and checked by MCMC methods. Differences between model fits and inferences are explored.Skewed parametric distributions fit cost data better than the normal distribution, and should in principle be preferred for estimating the population mean cost. However for some data sets, we find that models that fit badly can give similar inferences to those that fit well. Conversely, particularly when sample sizes are not large, different parametric models that fit the data equally well can lead to substantially different inferences. We conclude that inferences are sensitive to choice of statistical model, which itself can remain uncertain unless there is enough data to model the tail of the distribution accurately. Investigating the sensitivity of conclusions to choice of model should thus be an essential component of analysing cost data in practice.

Clinical Trials as Topic↗

A simple significance test for quantile regression.

Where OLS regression seeks to model the mean of a random variable as a function of observed variables, quantile regression seeks to model the quantiles of a random variable as functions of observed variables. Tests for the dependence of the quantiles of a random variable upon observed variables have only been developed through the use of computer resampling or based on asymptotic approximations resting on distributional assumptions. We propose an exceedingly simple but heretofore undocumented likelihood ratio test within a logistic regression framework to test the dependence of a quantile of a random variable upon observed variables. Simulated data sets are used to illustrate the rationale, ease, and utility of the hypothesis test. Simulations have been performed over a variety of situations to estimate the type I error rates and statistical power of the procedure. Results from this procedure are compared to (1) previously proposed asymptotic tests for quantile regression and (2) bootstrap techniques commonly used for quantile regression inference. Results show that this less computationally intense method has appropriate type I error control, which is not true for all competing procedures, comparable power to both previously proposed asymptotic tests and bootstrap techniques, and greater computational ease. We illustrate the approach using data from 779 adolescent boys age 12-18 from the Third National Health and Nutrition Examination Survey (NHANES III) to test hypotheses regarding age, ethnicity, and their interaction upon quantiles of waist circumference.

Adolescent↗

Statistical assessment of mediational effects for logistic mediational models.

The concept of mediation has broad applications in medical health studies. Although the statistical assessment of a mediational effect under the normal assumption has been well established in linear structural equation models (SEM), it has not been extended to the general case where normality is not a usual assumption. In this paper, we propose to extend the definition of mediational effects through causal inference. The new definition is consistent with that in linear SEM and does not rely on the assumption of normality. Here, we focus our attention on the logistic mediation model, where all variables involved are binary. Three approaches to the estimation of mediational effects-Delta method, bootstrap, and Bayesian modelling via Monte Carlo simulation are investigated. Simulation studies are used to examine the behaviour of the three approaches. Measured by 95 per cent confidence interval (CI) coverage rate and root mean square error (RMSE) criteria, it was found that the Bayesian method using a non-informative prior outperformed both bootstrap and the Delta methods, particularly for small sample sizes. Case studies are presented to demonstrate the application of the proposed method to public health research using a nationally representative database. Extending the proposed method to other types of mediational model and to multiple mediators are also discussed.

Adolescent↗

Variance estimation for clustered recurrent event data with a small number of clusters.

Often in biomedical studies, the event of interest is recurrent and within-subject events cannot usually be assumed independent. In semi-parametric estimation of the proportional rates model, a working independence assumption leads to an estimating equation for the regression parameter vector, with within-subject correlation accounted for through a robust (sandwich) variance estimator; these methods have been extended to the case of clustered subjects. We consider variance estimation in the setting where subjects are clustered and the study consists of a small number of moderate-to-large-sized clusters. We demonstrate through simulation that the robust estimator is quite inaccurate in this setting. We propose a corrected version of the robust variance estimator, as well as jackknife and bootstrap estimators. Simulation studies reveal that the corrected variance is considerably more accurate than the robust estimator, and slightly more accurate than the jackknife and bootstrap variance. The proposed methods are used to compare hospitalization rates between Canada and the U.S. in a multi-centre dialysis study.

Canada↗