PubMed HealthSearch

SEARCH · PubMed Health

Results for “Observer Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

[Observer variation in clinical practice. A review].

Observer variation denotes the discrepancy between two consecutive observations of the same indicant. Principles and results of observer variation studies are outlined. Observer variation seems to be present wherever it is searched for, in history taking as well as in physical examination. It has been found unanimously to be smaller for intra- than for inter- observer variation while it is still controversial whether clinical experience lessens observer variation. Even fundamental clinical observations and particularly evaluations of borderline cases between normal and abnormal exhibit observer variation. The possibilities of reducing observer variation, and the importance of observer variation to clinicians and patients, and also the influence on research and society are referred to briefly. The future role of observer variation studies is discussed. Observer variation often is evaluated by the Kappa statistics whose statistical advantages and clinical disadvantages are outlined.

Observer Variation

Is pancreatogram interpretation reliable?--a study of observer variation and error.

Observer variation in the interpretation of endoscopic pancreatograms has been examined by asking four experienced observers to assess 40 sets of well-documented radiographs (from 20 patients with cancer and 20 with pancreatitis), both without ("blind") and with clinical details, each on three occasions. Individual consistency for "blind" diagnoses ranged from 61% to 78%, increasing significantly with clinical information. Overall diagnostic accuracy with clinical information varied from 52% to 83% for cancer, and from 87% to 95% for pancreatitis. However, unanimous and correct opinions were given by the four observers for only 53% of all cases, even when clinical details were provided. Clinical information changed the radiographic diagnosis in 43% of assessments, 83% of these changes leading to improved accuracy. ERCP gives direct information about the major pancreatic and biliary duct systems and often provides an accurate diagnosis. Caution must be exercised in relying upon radiological appearances alone.

Diagnosis, Differential

Observer variation in quantification of immunocytochemistry by image analysis.

This paper reports the findings of a study designed to examine observer variation as a source of inaccuracy inherent in the use of computer-assisted image analysis to measure areas of stained tissue. The rat pituitary immunostained for prolactin and galanin was used as an example to estimate patterns of immunoreactivity exhibited by different cell types. Six observers, with differing experience, selected grey level threshold values on 40 fields of images of stained tissue making three repeats of each field. The 40 fields consisted of 20 serial pairs of colocalized fields, one immunostained for prolactin, the other for galanin. The 20 pairs consisted of four pairs from each of five animals. Analysis of observer variation in the selection of threshold values showed large differences in the within- and between-observer variation. Analysis of the components of variance in the estimation of the ratios of stained tissues showed that the major source of variation was the within-observer component. An additional experiment using two observers, where half of the images were compared to the original microscope images before setting threshold levels, showed that the opportunity to make a comparison did not reduce observer variation. It is suggested that any study which uses semi-automatic methods to segment regions of a digital image can benefit from an analysis of this kind so that the sources of variation can be determined to enable maximum discriminating power in future studies.

Animals

Observer variation in recording clinical data from women presenting with breast lesions. Report from the Yorkshire Breast Cancer Group.

The degree of observer variation in recording 11 186 items of clinical data from 242 woman who presented complaining of a lump in the breast to a group of 10 surgeons was studied. Each women was interviewed and examined twice and the findings (of the two clinicians) compared. There was a wide range of variation among the observers. Variation in recording the presence or absence of axillary nodes was considerable (45%), as was that in sizing the primary lesion (55%). In 20% of cases the two assessments of primary lesion size differed by over 2 cm. In other respects the results were more encouraging; much variation could be eliminated by wording the proforma more clearly. Moreover, variation was not person-specific, so that these findings are probably reasonably representative. Any future trial of breast lesions should (a) design specific proformata, (b) define terminology, (c) make these definitions universally available, and (d) conduct observer variation studies before the start of the full trial.

Analysis of Variance

[Observer variation and accuracy in the clinical diagnosis of ascites].

Seventeen observers participated in an observer variation study of the clinical evaluation of ascites. In a blinded design, eight patients with diagnoses of liver disease were examined. Fourteen observers examined the patients twice in order also to estimate the intra-observer variation. The accuracy of the observers' statements was compared with ultrasound findings, by which mean ascites was demonstrated in two patients. Poor correspondance between the observers' gradings and volume estimations, and poor accuracy of the gradings, make qualitative and quantitative ascites estimations useless. The inter-observer agreement was found to be low although the intra-observer agreement was good. The observers' subjective certainty of correctness of their own findings, marked as certain/uncertain, did not reflect the chance of making a correct statement on each particular occasion. The individual patient's general ability of inducing certainty did relate the chance of forming a correct diagnosis. Ultrasonic investigation of the abdomen is recommended in all situations in which demonstration of ascites is essential to diagnosis or therapy.

Ascites

Observer variation in the measurement of Breslow depth and Clark's level in thin cutaneous malignant melanoma.

We have assessed the degree of observer variation of both Breslow depth and Clark's level in a series of 50 thin malignant melanomas. Our findings are similar to those of previous international studies in the Breslow depth is the more reproducible measure. Significant intra- and inter-observer variation exists and in some cases it was up to +/- 0.86 mm. Even small differences will potentially affect patient management at our centre and this was analysed using kappa statistics. Good agreement was found between observers and this could be improved by comparing the mean of two or more measurements. This removes larger errors, but smaller observer errors and differences in subjective interpretation of the deepest malignant cell mean that agreement will never be more than 90 per cent. This is high compared with studies of observer variation in other pathological conditions, e.g., dysplasia of the cervix, but where surgical management is potentially disfiguring it is not high enough. We conclude that Breslow depth and Clark's level should not be the sole basis of wide excision protocols.

Biometry

[Assessment of micturition cystourethrography. Intra- and inter-observer variation].

The reliability of any method of investigation depends upon the accuracy and reproducibility of the results of the investigation. The accuracy of voiding cystoureterography (VCU) which is greatly dependent on the radiographic assessment cannot be assessed because no standard answers exist. The reliability of the method may be assessed in the form of intra- and inter-observer variations. VCU investigations from 24 women with incontinence were assessed by two independent radiologists. Fifteen of the radiological examinations had previously been described by one of the radiologists so that intra-observer variation could be assessed by these. Inter-observer variation was 70% (95% confidence limits 51-89%), calculated from the diagnoses anterior and posterior suspension defects, combined defects and normal conditions. The corresponding intra-observer variation was 53% (95% confidence limits 27-78%). The radiographic criteria for subdivision of suspension defects into anterior and posterior defects are, theoretically, very simple but appear to be difficult to attain in practice. The indications for employing a form of examination where assessment of the result of examination is so obscure should be very weighty.

Adult

Observer variation in detecting the radiologic features associated with bronchiolitis.

Chest radiographs are commonly obtained to assess children for bronchiolitis, both to corroborate the diagnosis and to exclude other diagnostic possibilities. Their utility in this setting has not previously been examined. Using a blinded, randomized study design, we examined the interobserver and intraobserver variation in the detection of the radiologic features of bronchiolitis from the chest radiograph using "weighted kappa" statistics. This observer variation was compared with that found by other authors for other diagnoses. We also determined the reported presence of these radiologic features in radiographs from patients with bronchiolitis as compared with normal controls. Our study showed acceptable interobserver (kappa = 0.40-0.66) and intraobserver agreement (kappa = 0.50-0.78) on the radiologic features of bronchiolitis relative to other diagnoses. We demonstrated a higher reported presence of these accepted radiologic features in patients with bronchiolitis as compared to controls. Although kappa statistics are widely used in studies of observer variation, "weighted kappa" has received little attention in the radiologic literature. This statistical analysis allows observers to equivocate on the presence or absence of a feature and therefore allows the format of observer variation studies to simulate more closely the normal clinical setting.

Bronchiolitis, Viral

Observer variation in the radiographic classification of ankle fractures.

We recorded inter- and intra-observer variations in the classification of ankle fractures by the Lauge Hansen and Weber systems. Radiographs of 94 patients were classified independently by four observers. The observer variation was calculated by kappa statistics, which corrects the obtained values for the agreement expected by chance. There was an acceptable level of agreement for the overall classification into both systems. For the staging of supination-adduction and supination-eversion fractures in the Lauge Hansen system the agreement was poor. The results indicate that future classification systems should be subject to reliability analysis before they are accepted.

Ankle Injuries

Observer variation and depressive phenomenology.

An investigation into observer variation and depressive phenomenology is described. A group of 6 clinicians independently completed a 43-item sheet for 20 depressed patients at the same clinical interview. The coefficient of agreements on the items are given. Patient variables and observer variables did not have any influence on the degree of agreement. There was also an agreement in the subtyping of depression.

Adjustment Disorders

Histologic features and observational variation in cerebellar gliomas in children.

Variation existed in the recognition of histologic features commonly used in the evaluation of cerebellar gliomas of childhood. Some histologic features (e.g., perivascular pseudorosettes, leptomeningeal deposits, and calcification) were more reliably observed than were others (e.g., Rosenthal fibers, cell density, and hypervascularity). Knowledge of which features tend to have greater observational variation may lead to improved definitions, less reliance of these features in clinical decisions, further studies of the potential sources of the variation, and guidelines for minimizing observational variation.

Arachnoid

Observer variation in the assessment of scintigraphy of the thyroid gland.

In order to determine observer variation in the assessment of thyroid scintigrams two specialists in nuclear medicine and two specialists in endocrinology independently evaluated 240 thyroid pertechnetate scintigrams twice, and assessed a number of variables concerning size and isotope uptake. The observed agreement between pairs of observers for the variables ranged from 0.70 to 0.98. By the use of the kappa coefficient the observed agreement was adjusted for change agreement. Kappa can variate from -1 (total disagreement) to +1 (perfect agreement). Kappa values between 0.29 and 0.86 were found. In the intraobserver study the observed agreement ranged from 0.83 to 0.99 resulting in kappa coefficients between 0.53 and 0.96. Thus the level of agreement in the present study was "fair to substantial" for agreement in the interobserver part and "substantial to almost perfect" for agreement in the intraobserver part. No difference was found in the level of agreement between the nuclear specialists and the endocrinologists. Although the treatment of patients is based on knowledge of the case histories and clinical and laboratory findings the high degree of observer variation may lead to misclassification of a number of patients with thyroid disease and subsequently a less optimal choice of treatment.

Adolescent

Observer variation in the clinical assessment of the thyroid gland.

In order to evaluate the reliability of clinical assessment of the thyroid gland, two specialists in endocrinology and two younger doctors independently examined 53 patients twice, and assessed whether they had a diffuse goitre, a multinodular goitre, a solitary nodule or a normal gland. In 30% of the patients all four observers were in agreement, whereas in 47% and 23% of the patients, two and three different diagnoses were given, respectively. Inter-observer variation was determined and kappa values between -0.04 and 0.54 were found. Intra-observer variation was smaller, revealing kappa values between 0.44 and 1.00. The present study suggests that clinical assessment of the thyroid gland may lead to misclassification of the type of thyroid disease, and thereby to a less than optimal choice of therapy.

Adult

Observer variation in interpreting radiographs of the pituitary fossa.

Observer variation in interpreting sellar radiographs in patients suspected or known to have a pituitary tumor has been examined. Two radiologists experienced in interpreting sellar radiographs examined independently, without clinical details, plain films and tomograms of the sella of 101 patients. In most, only minor changes were anticipated. Of the 93 female patients, 67 were under investigation for amenorrhea. Radiographs were examined four times, each radiologist examining each set twice. Appearances were classified as normal, doubtful or abnormal on each occasion. Overall intraobserver agreement was 76%--85%. Neither radiologist changes his opinion by more than one category, e.g. from normal to doubtful. Overall interobserver agreement was 63%--75%. Disagreement between observers concerning 11 (11%) of the patients resulted from differences of opinion about whether minor changes in sellar outline represented an abnormality or merely a normal variation. Kappa analysis suggested that much of the agreement may be ascribed to chance. Agreement rates resemble those for other clinical and radiological investigations.

Adolescent

Clinical assessment of ankylosing spondylitis: a study of observer variation in spinal measurements.

Twenty-two measurements repeated non-sequentially on each of 10 patients by five observers were undertaken to determine their reliability for routine clinical use. Measurements without significant inter-observer variation or with a coefficient of reliability greater than 0.70 were cervical rotation, cervical lateral flexion, tragus to wall distance, fingertip to floor distance on sagittal and lateral flexion, C7 to iliac crest line distraction and modified Schober index. It is concluded that many of the currently used measurements are either statistically unreliable or clinically unhelpful in mild or moderate ankylosing spondylitis. The most clinically useful were cervical rotation using a protractor, cervical lateral flexion using a goniometer, thoracolumbar flexion as the C7 to iliac crest line distraction, thoracolumbar lateral flexion as the fingertip to floor distance and the modified Schober index.

Adult

Intra- and inter-observer variation of optic nerve head measurements in glaucoma suspects using disc-data.

The aim of this study was to determine the intra- and inter-observer variation in the use of a system designed for exact measurements from standard optic nerve head photographs. The commercially available system consisted of a colour CCD Videocamera, a dedicated frame grabber and customized software run on a IBM AT compatible computer. Masked measurements were made 3 times by 2 observers, from stereophotographs of the optic nerve head of 56 eyes from 30 glaucoma suspects. The cup was defined on the basis of contour, not pallor and the disc area was defined as the area inside Elschnig's ring. Intra-observer variances were 0.001 +/- 0.001 mm2 for cup area (mean +/- SD), 0.002 +/- 0.002 mm2 for disc area and 0.002 +/- 0.003 mm2 for rim area. These values for intra-observer variance were comparable with the results obtained using manual planimetric techniques. Intra-observer variance for disc area was significantly larger for the less trained of the two experienced observers. Inter-observer variances were 0.004 +/- 0.009 mm2 for cup area, 0.008 +/- 0.013 mm2 for disc area and 0.009 +/- 0.014 mm2 for rim area. These inter-observer variances were significantly larger than those previously reported for manual planimetry. The absolute differences between the two observers ranged from -0.35 to +0.20 mm2 (-0.08 +/- 0.11 mm2) for cup area, from -0.38 to +0.15 mm2 (-0.08 +/- 0.11 mm2) for disc area and from -0.29 to +0.34 mm2 (-0.06 +/- 0.12 mm2) for rim area.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

Palpatory estimation of liver size. Within- and between-observer variation.

A study based upon 23 patients revealed substantial within- and between-observer variation, when 14 senior and junior surgeons under blindfold conditions were requested to measure the distance from the thoracic cage to the lower border of the liver in the midclavicular and midsternal lines. Clinical decisions should accordingly not rely heavily upon a palpatory estimate of liver size.

Adult

The definition of radiological signs in gastric ulcer and assessment of their validity by inter-observer variation study.

The initial aim was to program a computer with information on the frequency of radiological signs in benign and malignant gastric ulcers in order to obtain a percentage probability of benignancy or malignancy in succeeding ulcers in clinical practice. However, only four of the many signs described in gastric ulcer were confirmed to be of validity (i.e. reliable existence) by an inter-observer variation study using two observers and the films from 69 barium meal examinations. These were projection or non-projection of the in-profile ulcer, presence or absence of adjacent mucosal folds, good or poor definition of the in-face ulcer's edge, and extension of radiating folds to the in-face ulcer's edge. A few more remained unassessed due to insufficient numbers of relevant cases. It is condluced that: as defined in the literature the majority of radiological signs in this field are of uncertain existence; and the four that were found to be valid do not fully describe the important appearances that may be seen in benign and malignant ulcers and would be inadequate to differentiate them to a sufficiently high degree of probability.

Diagnosis, Computer-Assisted