PubMed HealthSearch

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Statistical properties of bootstrap estimation of phylogenetic variability from nucleotide sequences. I. Four taxa with a molecular clock.

The statistical properties of sample estimation and bootstrap estimation of phylogenetic variability from a sample of nucleotide sequences are studied by using model trees of three taxa with an outgroup and by assuming a constant rate of nucleotide substitution. The maximum-parsimony method of tree reconstruction is used. An analytic formula is derived for estimating the sequence length that is required if P, the probability of obtaining the true tree from the sampled sequences, is to be equal to or higher than a given value. Bootstrap estimation is formulated as a two-step sampling procedure: (1) sampling of sequences from the evolutionary process and (2) resampling of the original sequence sample. The probability that a bootstrap resampling of an original sequence sample will support the true tree is found to depend on the model tree, the sequence length, and the probability that a randomly chosen nucleotide site is an informative site. When a trifurcating tree is used as the model tree, the probability that one of the three bifurcating trees will appear in > or = 95% of the bootstrap replicates is < 5%, even if the number of bootstrap replicates is only 50; therefore, the probability of accepting an erroneous tree as the true tree is < 5% if that tree appears in > or = 95% of the bootstrap replicates and if more than 50 bootstrap replications are conducted. However, if a particular bifurcating tree is observed in, say, < 75% of the bootstrap replicates, then it cannot be claimed to be better than the trifurcating tree even if > or = 1,000 bootstrap replications are conducted. When a bifurcating tree is used as the model tree, the bootstrap approach tends to overestimate P when the sequences are very short, but it tends to underestimate that probability when the sequences are long. Moreover, simulation results show that, if a tree is accepted as the true tree only if it has appeared in > or = 95% of the bootstrap replicates, then the probability of failing to accept any bifurcating tree can be as large as 58% even when P = 95%, i.e., even when 95% of the samples from the evolutionary process will support the true tree. Thus, if the rate-constancy assumption holds, bootstrapping is a conservative approach for estimating the reliability of an inferred phylogeny for four taxa.

Phylogeny

The defibrillation success rate versus energy relationship: Part II--Estimation with the "bootstrap".

Seventy or so defibrillation trials were typically attempted to determine the relationship between defibrillation success rate and energy (DSRE). Clinically, it may be desirable to estimate the DSRE relationship with fewer trials. We used the statistical resampling technique called the "bootstrap" to determine the number of defibrillation trials necessary for an accurate estimation of the DSRE relationship. The bootstrap technique assumes that the observed database is the maximum likelihood sample of the estimated population. The observed database is repeatedly resampled to produce a large bootstrap data-base and the bootstrap best estimate of a statistic is determined. DSRE data were obtained from ten dogs (20.5 +/- 1.5 kg). We bootstrapped our experimental DSRE data by two methods: (1) randomly choosing with replacement a specified number of defibrillation trials per energy; and (2) randomly choosing with replacement a specified number of defibrillation trials per bootstrap replication. For both bootstrap techniques, 100 replications were made. We performed a linear regression analysis on the bootstrap success rates and the observed success rates determined from 71.0 +/- 6.8 defibrillation attempts from each of the ten dogs. We concluded that 28 defibrillation trials are necessary to estimate the observed DSRE relationship with a correlation coefficient of 0.95.

Animals

Exploring the order of odds ratios using the bootstrap.

We show the use of the bootstrap resampling method to examine further the order of a series of odds ratios. Specifically, the bootstrap provides a method for estimating the probabilities that one would find in subsequent independent samples from the same population the observed odds ratio rankings. To illustrate this use of the bootstrap, we modelled the responses of 77 white male physicians to an ethical dilemma involving hypothetical patients. Would the physician report positive HIV status to the health department or would he maintain patient confidentiality? To see if a patient's sex, race, or sexual preference would influence the physicians' decisions, each physician received one of eight randomly selected descriptions of a hypothetical patient. To evaluate the initial order of the patient categories, we constructed 1000 bootstrap samples. Black heterosexual males ranked first or second in 92.2 per cent of the bootstrap samples; black homosexual males ranked first, second or third in 88.6 per cent; and white homosexual females ranked sixth or seventh in 82.9 per cent. Thus we would expect to observe these rankings of the categories in a high percentage of subsequent independent samples.

Black or African American

Bootstrapping on the adaptive landscape.

Different versions of a gene or of a multigenic system may be essentially equivalent so far as the specific function of the structures which they code for or control is concerned, but very different with respect to their amenability to evolution. The structural features which increase evolutionary amenability are a disadvantage to the organism in terms of energy. Nevertheless, they accumulate in the course of evolution as a consequence of hitchhiking along with the desirable traits whose evolution they make possible. This is the bootstrap principle of evolutionary adaptability. In terms of the adaptive landscape bootstrapping corresponds to populations evolving in such a way that they occupy regions of the landscape which are more amenable to evolutionary hill climbing. The bootstrapping idea has implications for structure-function relations in a number of complex biological information processing systems, including biochemical systems, the immune system, and the brain. Bootstrapping is also discussed in connection with the origin of information processing (the origin of life) and in connection with possible designs for macromolecular computing systems.

Adaptation, Biological

Statistical analysis of the extended Hansen method using the bootstrap technique.

In this study, simple bootstrap techniques are combined with the extended Hansen solubility approach to calculate biases, standard errors, and confidence limits of the partial solubility parameters and to obtain bias-corrected values for these solubility parameters. The bootstrap method is rather new in its application to problems in the pharmaceutical sciences and, therefore, is described here in some detail. This method provides measures of the statistical variation of ratios of regression coefficients without making unwarranted assumptions about data variability. The bootstrap can be used in many statistical packages such as MINITAB, SPSS, SAS, BMDP, or GLIM, all of which are widely available, and could be useful in other areas of the pharmaceutical sciences where regression analysis is employed.

Computer Simulation

A bootstrap resampling procedure for model building: application to the Cox regression model.

A common problem in the statistical analysis of clinical studies is the selection of those variables in the framework of a regression model which might influence the outcome variable. Stepwise methods have been available for a long time, but as with many other possible strategies, there is a lot of criticism of their use. Investigations of the stability of a selected model are often called for, but usually are not carried out in a systematic way. Since analytical approaches are extremely difficult, data-dependent methods might be an useful alternative. Based on a bootstrap resampling procedure, Chen and George investigated the stability of a stepwise selection procedure in the framework of the Cox proportional hazard regression model. We extend their proposal and develop a bootstrap-model selection procedure, combining the bootstrap method with existing selection techniques such as stepwise methods. We illustrate the proposed strategy in the process of model building by using data from two cancer clinical trials featuring two different situations commonly arising in clinical research. In a brain tumour study the adjustment for covariates in an overall treatment comparison is of primary interest calling for the selection of even 'mild' effects. In a prostate cancer study we concentrate on the analysis of treatment-covariate interactions demanding that only 'strong' effects should be selected. Both variants of the strategy will be demonstrated analysing the clinical trials with a Cox model, but they can be applied in other types of regression with obvious and straightforward modifications.

Brain Neoplasms

Statistical properties of bootstrap estimation of phylogenetic variability from nucleotide sequences: II. Four taxa without a molecular clock.

The statistical properties of sample estimation and bootstrap estimation of phylogenetic variability from a sample of nucleotide sequences were studied by considering model trees of three taxa with an outgroup. The cases of constant and varying rates of nucleotide substitution were compared. From sequences obtained by simulation, phylogenetic trees were constructed by using the maximum parsimony (MP) and neighbor-joining (NJ) methods. The effectiveness and consistency of the MP method were studied in terms of proportions of informative sites. The results of simulation showed that bootstrap estimation of the confidence level for an inferred phylogeny can be used even under unequal rates of evolution if the rate differences are not large so that the MP method is not misleading. The condition under which the MP method becomes misleading (inconsistent) is more stringent for slowly evolving sequences than for rapidly evolving ones, and it also depends on the length of the internal branch. If the rate differences are large so that the MP method becomes consistently misleading, then bootstrap estimation will reinforce an erroneous conclusion on topology. Similar conclusions apply to the NJ method with uncorrected distances. The NJ method with corrected distances performs poorly when the sequence length is short but can avoid the inconsistency problem if the sequence length is long and if the distances can be estimated accurately.

Base Sequence

Nonparametric assessment of toxicologic assay linearity by bootstrap analysis.

An important aspect in the evaluation of toxicologic assay methodology is the assessment of calibration. This paper presents a new approach for validating calibration using bootstrap analysis. The technique is illustrated with a quantitative assay for benzoylecgonine in urine by gas chromatography/mass spectrometry (GC/MS). Because the bootstrap analysis does not rely on parametric assumptions regarding the distribution of errors about the regression line, it can be applied to situations where the parametric distribution of response errors about the regression line is unknown. This application of bootstrap methodology yields a probabilistic measure of confidence on the linearity of the calibration curve.

Calibration

[Jackknife and bootstrap].

The jackknife and the bootstrap are two non parametric methods which provide estimates- of the bias and the variance of an estimator, without any assumption about its statistical distribution. The jackknife is based on the observation of the estimator for subsamples, generally of size n-1, obtained from the original sample. The bootstrap is based on the observation of the estimator on size n samples drawn from the original sample. The two methods are presented, their principle is illustrated through their application to simple examples and to more complex epidemiological problems.

Bias

Identifying Single-Cell Expression Quantitative Trait Loci Using a Bootstrap Penalized Hurdle Model.

BACKGROUND: Expression quantitative trait loci (eQTL) analysis links genetic variants to gene expression levels, helping to uncover how genetic variation contributes to gene regulation. While traditional eQTL analyses rely on bulk RNA-seq data, recent advances in single-cell RNA sequencing (scRNA-seq) have made it possible to detect cell-type-specific eQTLs. However, the inherent sparsity and heterogeneity of scRNA-seq data present major challenges for standard modeling approaches. METHODS: In this paper, we propose a novel statistical framework, Bootstrap Penalized Hurdle regression model (BPHurdle), designed specifically for scRNA-seq data. BPHurdle employs a hurdle modeling framework, where a logistic component accounts for the excess zeros in single-cell expression data, and a Poisson component jointly evaluates the effects of multiple SNPs on positive gene expression levels. RESULTS: Through simulation studies, we show that BPHurdle achieves high accuracy and robustness in identifying regulatory variants. We further demonstrate its utility on a real dataset through a case study focusing on a subset of differentially expressed genes, where it successfully identifies reliable cell-type-specific eQTLs. CONCLUSIONS: Overall, BPHurdle offers an advanced and flexible approach for single-cell eQTL mapping, providing deeper insight into the genetic regulation of gene expression at cellular resolution.

Quantitative Trait Loci

Bootstrapping: a tool for clinical research.

The use of the bootstrap sampling technique is applied to the type of data found in clinical research. Confidence intervals are computed for simulated values by use of SAS. By applying this approach, clinical researchers are free to explore topics that do not meet the requirements of traditional statistical analytic methods.

Acquired Immunodeficiency Syndrome

What sort of innate structure is needed to "bootstrap" into syntax?

The paper starts from Pinker's theory of the acquisition of phrase structure; it shows that it is possible to drop all the assumptions about innate syntactic structure from this theory. These assumptions can be replaced by assumptions about the basic structure of semantic representation available at the outset of language acquisition, without penalizing the acquisition of basic phrase structure rules. Essentially, the role played by X-bar theory in Pinker's model would be played by the (presumably innate) structure of the language of thought in the revised parallel model. Bootstrapping and semantic assimilation theories are shown to be formally very similar, though making different primitive assumptions. In their primitives, semantic assimilation theories have the advantage that they can offer an account of the origin of syntactic categories instead of postulating them as primitive. Ways of improving on the semantic assimilation version of Pinker's theory are considered, including a way of deriving the NP-VP constituent division that appears to have a better fit than Pinker's to evidence on language variation.

Child, Preschool

Estimating effective population size from samples of sequences: a bootstrap Monte Carlo integration method.

We would like to use maximum likelihood to estimate parameters such as the effective population size N(e) or, if we do not know mutation rates, the product 4N(e) mu of mutation rate per site and effective population size. To compute the likelihood for a sample of unrecombined nucleotide sequences taken from a random-mating population it is necessary to sum over all genealogies that could have led to the sequences, computing for each one the probability that it would have yielded the sequences, and weighting each one by its prior probability. The genealogies vary in tree topology and in branch lengths. Although the likelihood and the prior are straightforward to compute, the summation over all genealogies seems at first sight hopelessly difficult. This paper reports that it is possible to carry out a Monte Carlo integration to evaluate the likelihoods approximately. The method uses bootstrap sampling of sites to create data sets for each of which a maximum likelihood tree is estimated. The resulting trees are assumed to be sampled from a distribution whose height is proportional to the likelihood surface for the full data. That it will be so is dependent on a theorem which is not proven, but seems likely to be true if the sequences are not short. One can use the resulting estimated likelihood curve to make a maximum likelihood estimate of the parameter of interest, N(e) or of 4N(e) mu. The method requires at least 100 times the computational effort required for estimation of a phylogeny by maximum likelihood, but is practical on today's work stations. The method does not at present have any way of dealing with recombination.

Base Sequence

Input evidence regarding the semantic bootstrapping hypothesis.

The input language addressed to 18 language-learning children (MLU 1.00-3.00) was analysed so as to assess the quality of the semantic-syntactic correspondence posited by the semantic bootstrapping hypothesis. The correspondence appears to be quite satisfactory with little variation from the lower to the higher MLUs. All the persons and things referred to in the corpora were labelled by the mothers using nouns. All the actions referred to were labelled using verbs. Most of the attributive information was conveyed by adjectives. Spatial information was expressed through the use of spatial prepositions. As to the functional categories, all agents of actions and causes of events were encoded as subjects of sentences. All patients, themes, sources, goals, locations, and instruments were encoded as objects of sentences (either direct or oblique). This good semantic-syntactic correspondence may make the child's construction of grammatical categories easier.

Child Language

Inference of horizontal genetic transfer from molecular data: an approach using the bootstrap.

Inconsistencies in taxonomic relationships implicit in different sets of nucleic acid sequences potentially result from horizontal transfer of genetic material between genomes. A nonparametric method is proposed to determine whether such inconsistencies are statistically significant. A similarity coefficient is calculated from ranked pairwise identities and evaluated against a distribution of similarity coefficients generated from resampled data. Subsequent analyses of partial data sets, obtained by the elimination of individual taxa, identify particular taxa to which the significance may be attributed, and can sometimes help in distinguishing horizontal genetic transfer from inconsistencies due to convergent evolution or variation in evolutionary rate. The method was successfully applied to data sets that were not found to be significantly different with existing methods that use comparisons of phylogenetic trees. The new statistical framework is also applicable to the inference of horizontal transfer from restriction fragment length polymorphism distributions and protein sequences.

Animals