PubMed Health⌕ Search

Biomedical subjects

Geoff Cumming

Publications and source records attributed to Geoff Cumming.

11 recordsLinked to original sources

Confidence intervals and replication: where will the next mean fall?

Confidence intervals (CIs) give information about replication, but many researchers have misconceptions about this information. One problem is that the percentage of future replication means captured by a particular CI varies markedly, depending on where in relation to the population mean that CI falls. The authors investigated the distribution of this percentage for varsigma known and unknown, for various sample sizes, and for robust CIs. The distribution has strong negative skew: Most 95% CIs will capture around 90% or more of replication means, but some will capture a much lower proportion. On average, a 95% CI will include just 83.4% of future replication means. The authors present figures designed to assist understanding of what CIs say about replication, and they also extend the discussion to explain how p values give information about replication.

Confidence Intervals↗

Impact of criticism of null-hypothesis significance testing on statistical reporting practices in conservation biology.

Over the last decade, criticisms of null-hypothesis significance testing have grown dramatically, and several alternative practices, such as confidence intervals, information theoretic, and Bayesian methods, have been advocated. Have these calls for change had an impact on the statistical reporting practices in conservation biology? In 2000 and 2001, 92% of sampled articles in Conservation Biology and Biological Conservation reported results of null-hypothesis tests. In 2005 this figure dropped to 78%. There were corresponding increases in the use of confidence intervals, information theoretic, and Bayesian techniques. Of those articles reporting null-hypothesis testing--which still easily constitute the majority--very few report statistical power (8%) and many misinterpret statistical nonsignificance as evidence for no effect (63%). Overall, results of our survey show some improvements in statistical practice, but further efforts are clearly required to move the discipline toward improved practices.

Conservation of Natural Resources↗

Toward improved statistical reporting in the journal of consulting and clinical psychology.

Philip Kendall's (1997) editorial encouraged authors in the Journal of Consulting and Clinical Psychology (JCCP) to report effect sizes and clinical significance. The present authors assessed the influence of that editorial--and other American Psychological Association initiatives to improve statistical practices--by examining 239 JCCP articles published from 1993 to 2001. For analysis of variance, reporting of means and standardized effect sizes increased over that period, but the rate of effect size reporting for other types of analyses surveyed remained low. Confidence interval reporting increased little, reaching 17% in 2001. By 2001, the percentage of articles considering clinical (not only statistical) significance was 40%, compared with 36% in 1996. In a follow-up survey of JCCP authors (N=62), many expressed positive attitudes toward statistical reform. Substantially improving statistical practices may require stricter editorial policies and further guidance for authors on reporting and interpreting measures.

Analysis of Variance↗

Researchers misunderstand confidence intervals and standard error bars.

Little is known about researchers' understanding of confidence intervals (CIs) and standard error (SE) bars. Authors of journal articles in psychology, behavioral neuroscience, and medicine were invited to visit a Web site where they adjusted a figure until they judged 2 means, with error bars, to be just statistically significantly different (p < .05). Results from 473 respondents suggest that many leading researchers have severe misconceptions about how error bars relate to statistical significance, do not adequately distinguish CIs and SE bars, and do not appreciate the importance of whether the 2 means are independent or come from a repeated measures design. Better guidelines for researchers and less ambiguous graphical conventions are needed before the advantages of CIs for research communication can be realized.

Confidence Intervals↗

Cultivating expertise in informal reasoning.

People generally develop some degree of competence in general informal reasoning and argument skills, but how do they go beyond this to attain higher expertise? Ericsson has proposed that high-level expertise in a variety of domains is cultivated through a specific type of practice, referred to as "deliberate practice." Applying this framework yields the empirical hypothesis that high-level expertise in informal reasoning is the outcome of extensive, deliberate practice. This paper reports results from two studies evaluating the hypothesis. University student participants completed 12 weeks of deliberate practice in informal reasoning. Quantity of practice was recorded by computer, and additionally assessed via self-report. The hypothesis was supported: Students in both studies showed a large improvement, and practice, as measured by computer, was related to amount of improvement in informal reasoning. These findings support adopting a deliberate practice approach when attempting to teach or learn expertise in informal reasoning.

Adult↗

Editors can lead researchers to confidence intervals, but can't make them think: statistical reform lessons from medicine.

Since the mid-1980s, confidence intervals (CIs) have been standard in medical journals. We sought lessons for psychology from medicine's experience with statistical reform by investigating two attempts by Kenneth Rothman to change statistical practices. We examined 594 American Journal of Public Health (AJPH) articles published between 1982 and 2000 and 110 Epidemiology articles published in 1990 and 2000. Rothman's editorial instruction to report CIs and not p values was largely effective: In AJPH, sole reliance on p values dropped from 63% to 5%, and CI reporting rose from 10% to 54%; Epidemiology showed even stronger compliance. However, compliance was superficial: Very few authors referred to CIs when discussing results. The results of our survey support what other research has indicated: Editorial policy alone is not a sufficient mechanism for statistical reform. Achieving substantial, desirable change will require further guidance regarding use and interpretation of CIs and appropriate effect size measures. Necessary steps will include studying researchers' understanding of CIs, improving education, and developing empirically justified recommendations for improved statistical practice.

Biomedical Research↗

Reform of statistical inference in psychology: the case of memory & cognition.

Geoffrey Loftus, Editor of Memory & Cognition from 1994 to 1997, strongly encouraged presentation of figures with error bars and avoidance of null hypothesis significance testing (NHST). The authors examined 696 Memory & Cognition articles published before, during, and after the Loftus editorship. Use of figures with bars increased to 47% under Loftus's editorship and then declined. Bars were rarely used for interpretation, and NHST remained almost universal. Analysis of 309 articles in other psychology journals confirmed that Loftus's influence was most evident in the articles he accepted for publication, but was otherwise limited. An e-mail survey of authors of papers accepted by Loftus revealed some support for his policy, but allegiance to traditional practices as well. Reform of psychologists' statistical practices would require more than editorial encouragement.

Bibliometrics↗

A clinical procedure for assessment of severity of knee pain.

We describe a clinical procedure for assessing knee pain: 10 standardised movements of the knee are made (r active and 6 passive), and the subject's behavioural response to each is scored by the assessor on a 0-3 scale. The total score is the Pain Index of the Knee (PIK). We report investigations of repeatability, inter-assessor agreement, and validity of components of the PIK compared with VAS ratings by the subject. We conclude from these results and from our experience of the PIK in practice that it is useful both clinically and for research purposes. It is a simple and efficient procedure for pain measurement in terms of behavioural response, with good reliability and validity, at least at low and medium levels of pain.

Aged↗

Inference by eye: confidence intervals and how to read pictures of data.

Wider use in psychology of confidence intervals (CIs), especially as error bars in figures, is a desirable development. However, psychologists seldom use CIs and may not understand them well. The authors discuss the interpretation of figures with error bars and analyze the relationship between CIs and statistical significance testing. They propose 7 rules of eye to guide the inferential use of figures with error bars. These include general principles: Seek bars that relate directly to effects of interest, be sensitive to experimental design, and interpret the intervals. They also include guidelines for inferential interpretation of the overlap of CIs on independent group means. Wider use of interval estimation in psychology has the potential to improve research communication substantially.

Confidence Intervals↗