PubMed Health⌕ Search

PubMed · 15238124

How many do I need? Basic principles of sample size estimation.

Abstract

BACKGROUND: In conducting randomized trials, formal estimations of sample size are required to ensure that the probability of missing an important difference is small, to reduce unnecessary cost and to reduce wastage. Nevertheless, this aspect of research design often causes confusion for the novice researcher. AIM: This paper attempts to demystify the process of sample size estimation by explaining some of the basic concepts and issues to consider in determining appropriate sample sizes. METHOD: Using a hypothetical two group, randomized trial as an example, we examine each of the basic issues that require consideration in estimating appropriate sample sizes. Issues discussed include: the ethics of randomized trials, the randomized trial, the null hypothesis, effect size, probability, significance level and type I error, and power and type II error. The paper concludes with examples of sample size estimations with varying effect size, power and alpha levels. CONCLUSION: Health care researchers should carefully consider each of the aspects inherent in sample size estimations. Such consideration is essential if care is to be based on sound evidence, which has been collected with due consideration of resource use, clinically important differences and the need to avoid, as far as possible, types I and II errors. If the techniques they employ are not appropriate, researchers run the risk of misinterpreting findings due to inappropriate, unrepresentative and biased samples.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Declan Devane, Cecily M Begley, Mike Clarke. 2004. How many do I need? Basic principles of sample size estimation.. https://doi.org/10.1111/j.1365-2648.2004.03093.x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Adjusting for partially missing baseline measurements in randomized trials.

Adjustment for baseline variables in a randomized trial can increase power to detect a treatment effect. However, when baseline data are partly missing, analysis of complete cases is inefficient. We consider various possible improvements in the case of normally distributed baseline and outcome variables. Joint modelling of baseline and outcome is the most efficient method. Mean imputation is an excellent alternative, subject to three conditions. Firstly, if baseline and outcome are correlated more than about 0.6 then weighting should be used to allow for the greater information from complete cases. Secondly, imputation should be carried out in a deterministic way, using other baseline variables if possible, but not using randomized arm or outcome. Thirdly, if baselines are not missing completely at random, then a dummy variable for missingness should be included as a covariate (the missing indicator method). The methods are illustrated in a randomized trial in community psychiatry.

Data Interpretation, Statistical↗

Missed hypothyroidism diagnosis uncovered by linking laboratory and pharmacy data.

BACKGROUND: Although diagnostic errors are important, they have received less attention than medication errors. Timely follow-up of abnormal laboratory test results represents a critical aspect of the diagnostic process, and failures at this step are a cause of delayed or missed diagnosis, resulting in suboptimal clinical outcomes and malpractice litigation. We linked laboratory and pharmacy databases to (1) explore the potential for linking laboratory and pharmacy databases to uncover diagnostic errors, and (2) determine the frequency of failed follow-up of elevated levels of thyroid-stimulating hormone (TSH). METHODS: We downloaded TSH test results for 2 consecutive years from a laboratory database and linked this database with a pharmacy database to screen for patients with TSH levels of 20 mU/mL or higher who were not receiving levothyroxine. Patients with elevated TSH levels lacking prescriptions were followed up by telephone and record review. RESULTS: During the 2-year period, 982 (2.7%) of 36 760 unique patients tested for TSH level had elevated TSH levels. Of these patients, 177 (18.0%) had no recorded levothyroxine prescriptions. We attempted to contact 177 patients with high TSH levels who were not taking thyroid medications and reached 123 (69.5%). Of the 123 patients we were able to reach, 12 in 2000 and 11 in 2001 were unaware of their abnormal test results or a diagnosis of hypothyroidism, representing 2.3% of 982 patients with elevated TSH levels. We were unable to reach another 54 patients (5.5% of the total number of patients with elevated TSH levels) by either telephone or mail. CONCLUSIONS: By linking laboratory and pharmacy databases, we uncovered patients who did not undergo follow-up for abnormal TSH results. Conservatively, there was no follow-up for abnormal TSH results in more than 2% of patients, and another 5% of patients were lost to follow-up and possibly unaware of their results. Uncovering patients with missed diagnosis illustrates a potential use of linking laboratory and pharmacy databases to identify vulnerabilities in the care system and improve patient safety.

Data Interpretation, Statistical↗