Comments on journal guidelines for reporting statistics.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to John Ludbrook.
Explore the source record for details and available documents.
BACKGROUND: Techniques for interim analysis, the statistical analysis of results while they are still accumulating, are highly-developed in the setting of clinical trials. But in the setting of laboratory experiments such analyses are usually conducted secretly and with no provisions for the necessary adjustments of the Type I error-rate. DISCUSSION: Laboratory researchers, from ignorance or by design, often analyse their results before the final number of experimental units (humans, animals, tissues or cells) has been reached. If this is done in an uncontrolled fashion, the pejorative term 'peeking' has been applied. A statistical penalty must be exacted. This is because if enough interim analyses are conducted, and if the outcome of the trial is on the borderline between 'significant' and 'not significant', ultimately one of the analyses will result in the magical P = 0.05. I suggest that Armitage's technique of matched-pairs sequential analysis should be considered. The conditions for using this technique are ideal: almost unlimited opportunity for matched pairing, and a short time between commencement of a study and its completion. Both the Type I and Type II error-rates are controlled. And the maximum number of pairs necessary to achieve an outcome, whether P = 0.05 or P > 0.05, can be estimated in advance. SUMMARY: Laboratory investigators, if they are to be honest, must adjust the critical value of P if they analyse their data repeatedly. I suggest they should consider employing matched-pairs sequential analysis in designing their experiments.
There have been published at least two major sets of contributions to the matter of peer review of manuscripts since my last article on this topic. In one, the merits of truly open peer review, in which the names of authors and their affiliations are revealed to reviewers, and the names of reviewers to authors, are extolled. The other contribution is not so original, in that it exhorts biomedical investigators and authors to consult with professional statisticians. But the vigorous correspondence that followed was interesting. I have come down strongly in favour of open peer review for all biomedical journals. However, I have also warned investigators and authors that statisticians often do not agree and, sometimes, violently disagree. I suggest that it is time a prospective, comparative study of statistical reviewers and their reviews should be carried out.
1. Clinical and experimental pharmacologists and physiologists often wish to compare two methods of measurement, or two measurers. 2. Biostatisticians insist that what should be sought is not agreement between methods or measurers, but disagreement or bias. 3. If measurements have been made on a continuous scale, the main choice is between the Altman-Bland method of differences and least products regression analysis. It is argued that although the former is relatively simple to execute, it does not distinguish adequately between fixed and proportional bias. Least products regression analysis, although more difficult to execute, does achieve this goal. There is almost universal agreement among biostatisticians that the Pearson product-moment correlation coefficient (r) is valueless as a test for bias. 4. If measurements have been made on a categorical scale, unordered or ordered, the most popular method of analysis is to use the kappa statistic. If the categories are unordered, the unweighted kappa statistic (K) is appropriate. If the categories are ordered, as they are in most rating scales in clinical, psychological and epidemiological research, the weighted kappa statistic (K(w)) is preferable. But K(w) corresponds to the intraclass correlation coefficient, which, like r for continuous variables, is incapable of detecting bias. Simple techniques for detecting bias in the case of ordered categorical variables are described and commended to investigators.
BACKGROUND: The Surgical Research Society of Australasia (SRS) was established in 1961. The main intent was to promote surgical research, especially by young surgeons. The 50th Scientific Meeting of the SRS is due to be held in 2002. METHODS: Selected information was extracted from the programme books for the scientific meetings of the SRS, while historical material was retrieved from the Records and Archives Unit of the Royal Australasian College of Surgeons. RESULTS: There have been 48 scientific meetings of the SRS between 1962 and 2000 inclusive, which have included the presentation of 1627 free papers. Forty-one percent of the presenters (325/784) have done so on more than one occasion. The content of the presentations reflects the changing nature of surgical research during the last several decades. CONCLUSIONS: The SRS has achieved its original intent of promoting surgical research.
Analysis of the reproducibility of peer review of manuscripts by means of the kappa statistic is fatally flawed from the point of view of statistical theory. An alternative, simple, method of analysis is proposed. On this basis, agreement among reviewers for the Journal of Clinical Neuroscience is at least as good as that reported recently for other clinical neuroscience journals. Nevertheless, a broad review of peer review processes demonstrates that they are far from satisfactory. Might electronic publishing of scientific articles provide a solution?
1. I recently reviewed, inter alia, methods for comparing two raters who make judgements on an ordered categorical scale, directed principally at the kappa statistic and its weaknesses. 2. The main weakness of the kappa statistic is that it fails to detect the all-important feature of systematic bias between raters. 3. I described various methods for detecting bias between two raters. These included a modified McNemar test and the single binomial test. Others that have been suggested are the symmetry of disagreement index (SD) and the marginal homogeneity test. 4. I now realize that none of the above four tests for bias is satisfactory, because all ignore the extent of agreement. 5. The bias index (BI) does take into account the extent of agreement, but its inventors did not propose how BI could be evaluated. I now describe a method for doing this.