PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Observer Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Reliability of radiographical measurements of spondylolisthesis and extension-flexion radiographs of the lumbar spine.

We studied the reliability of extension-flexion radiography by analysing intra-observer and inter-observer variations in measurements of 30 patients with established L5-S1 spondylolisthesis. The highest intra-observer angular variations among three radiologists were found at the L5-S1 level (mean 1.6 degrees, S.D. 1.6 degrees, max. 9 degrees) and the highest sagittal translation also at the L5-S1 level (mean 0.6 mm, S.D. 0.8 mm, max. 4 mm). The highest angular inter-observer variation was found at the L5-S1 level (mean 2.6 degrees, S.D. 2.3 degrees, max. 11 degrees) while the highest variation in sagittal translatory movement was found at the L4-L5 level (mean 1.4 mm, S.D. 1.2 mm, max. 6 mm). The mean intra-observer variation for L5 olisthesis was 1.0 mm (S.D. 0.9 mm, max. 5 mm) and the corresponding inter-observer variation 1.3 mm (S.D. 1.1 mm, max. 6 mm). Thus, in general, the overall consistency and concordance were good. The variations in intra-observer and inter-observer readings were similar. Variability in readings did not depend on the magnitude or abnormality of the readings. On the other hand, variability may occur because of difficulty in defining the exact landmarks used for the measurements. We point out that in certain cases the diagnostic value of one single reading of an extension-flexion radiography examination may be questionable.

Adolescent↗

Evaluation of a method of assessing faecal loading on plain abdominal radiographs in children.

BACKGROUND: Childhood constipation is common and assessment is often difficult. Plain abdominal radiography is simple and commonly used to assess constipation. The role of radiography with the use of a simple scoring system has not been fully evaluated. OBJECTIVE: To assess the reliability of scoring faecal loading on plain abdominal radiographs in children with intractable constipation. MATERIALS AND METHODS: Plain abdominal radiographs from 33 constipated and 67 control children were independently assessed by three observers on two separate occasions. A scoring system was devised with scores from 0 (no stool) to 5 (gross faecal loading with bowel dilatation) in three areas of the colon, giving a total score of 0-15. RESULTS: There were significant differences between the scores of the constipated and control radiographs for each observer (P = 0.05). There was no intra-observer variation (P = 0.12-0.69), but significant inter-observer variation was demonstrated (P = 0.00). CONCLUSIONS: We have found this scoring system to be a clinically useful and a reproducible tool in assessing childhood constipation. Assessment of faecal loading is subjective and varies between observers, although one observer will consistently score faecal loading on the same radiograph on successive occasions. To limit exposure to ionising radiation, we recommend that radiography be reserved for the investigation of intractable constipation, and its accuracy is improved if all radiographs are scored by the same observer.

Adolescent↗

Regional observer performance variation in the evaluation of gated cardiac blood pool studies.

Receiver operating characteristic (ROC) analysis demonstrated that regional variations of sensitivity exist in the detection of wall motion abnormality in cardiac blood pool imaging studies. The observer response is significantly better in the apex than either the septum or posterolateral wall segments. The observer errors tend to be false-negative in the posterolateral wall segment and false-positive in the other two segments. Image presentation can make a significant difference to the overall sensitivity, and the monochrome cine-sequence performed best in this study.

Gated Blood-Pool Imaging↗

An assessment of disability rating scales used in multiple sclerosis.

Twenty patients with clinically definite, stable multiple sclerosis were examined independently by three of us at the same visit and given scores on the Ambulation Index, Expanded Disability Status Scale, and Kurtzke Functional System scales. Observer error accounted for 12% to 55% of the variation observed between individual Kurtzke Functional System scores, 17.1% of the variation observed between the patients' Expanded Disability Status Scale scores, and only 3.9% of the variation between Ambulation Index scores. The implications of these findings for the choice of scales in clinical trials are described.

Adult↗

Measuring dyspepsia: a new severity index validated in Bologna.

BACKGROUND: Measurement of the severity of dyspepsia symptoms before and after treatment and determining what is a significant change is a major problem in designing dyspepsia treatment studies. OBJECTIVES: To assess the reproducibility, validity and responsiveness to treatment of a dyspepsia questionnaire to be used in clinical and population-based studies. METHODS: Seventy-three dyspeptic patients (35 male, 38 female; mean age 52 years) and 75 healthy volunteers (32 male, 43 female; mean age 52 years) were included. Subjects were interviewed for the presence/absence and severity/frequency of 19 gastrointestinal symptoms. Severity was measured on a 5-point scale. Frequency was also recorded on a 5-point scale. A global symptom index (severity x frequency) was calculated for the eight most severe symptoms; a mean global symptom index (8-MGSI) was considered for the evaluation of the instrument. To evaluate intra-observer variation, one author interviewed subjects (T0) and then repeated the interview 1 week later (T1). For inter-observer variation, two authors interviewed patients. VALIDITY was measured by comparing 8-MGSI of the dyspepsia patients to those of healthy volunteers. Responsiveness was assessed by comparing mean global symptom index before and 1 month after appropriate therapy. RESULTS: Reproducibility: The mean 8-MGSI was 4.5 at T0 and 3.7 at T1 with a correlation coefficient of 0.62. As for inter-observer variation, the average 8-MGSI was 4.8 by the first author and 3.9 by the second with a correlation coefficient of 0.60. VALIDITY: The mean 8-MGSI was, respectively, 1.4 in healthy volunteers and 4.8 in dyspeptic patients (p = 0.001). Responsiveness: After treatment, a significant improvement in 8-MGSI was detected (p = 0.001). CONCLUSIONS: This questionnaire is a reliable, valid and responsive instrument for measuring the presence, severity and frequency of dyspepsia.

Breath Tests↗

Intra-observer and inter-observer agreement of the manual examination of the lumbar spine in chronic low-back pain.

Examination is a cornerstone in the manual procedures leading to mobilisation/manipulation of the low back. The observer variation of the more specific segmental tests remains to be investigated. Two skilled specialists in manual medicine examined the segmental changes in the lumbar spine. The patients were unknown to the examiners and no information of the case history was given. All test results were recorded by an observer present in the room who ensured that no conversation was allowed during the examination. The primary outcome measures were the kappa values for each test. The matching was defined as acceptable (acc) within two neighbouring levels and perfect (per) on the same level. Intra-observer variation (tested in 33 patients and 10 subjects without low-back pain): The agreement between first and second segmental diagnosis examination was 70% (per) and 82% (per + acc). Kappa values were: segmental diagnosis 0.60 (per) and 0.70 (per + acc), multifidus test 0.51 (per) and 0.60 (per + acc), sideflexion 0.57 (per) and 0.69 (per + acc), and ventral flexion 0.31 (per) and 0.45 (per + acc). Inter-observer variation (tested in 60 patients): The agreement for segmental diagnosis between the examiner A and B was 42% (per) and 75% (per + acc). Kappa values were: segmental diagnosis 0.21 (per) and 0.57 (acc), multifidus test 0.12 (per) and 0.48 (acc), sideflexion 0.22 (per) and 0.45 (acc), and ventralflexion 0.22 (per) and 0.44 (acc). By manual tests, skilled examiners seem to be able to diagnose segmental dysfunctions in the low back. The clinical implication of these dysfunctions remains to be clarified.

Adolescent↗

Slit camera focal spot measurement errors in mammography.

Mammography x-ray tube focal spot sizes are routinely measured during acceptance testing and annual performance audits. The National Electrical Manufactures Association (NEMA) recommends the slit camera for this purpose. Investigated were the effect of slit rotational misalignment, tilt misalignment, image film density, film and screen-film image receptors, microscope magnification and reticule accuracy, and observer variation on slit camera focal spot measurements. Our results indicate that small rotational misalignment (< 5 degrees) and tilt misalignment (< 3 degrees) introduce insignificant error. Measured focal spot size increased slightly with image optical density, indicating that for consistent results the image optical density variations should be minimized. Also desirable for accurate field measurements is a high power microscope (25-50x) and a reticule with divisions of < or = 0.02 mm. Screen-film imaging consistently resulted in a slightly smaller measured focal spot size than direct film. The greatest source of error was due to observer variation. Of interest is that reader variability showed a consistent pattern and variation between two measurements by the same observer was much smaller than between observer variation, suggesting that standardized criteria should be established and a method of reader training developed. The length of the focal spot is defined at a reference axis angle specified by the mammography unit manufacturer. Presented is a tabulation of the focal spot geometry and reference axis angles for the majority of mammography units currently and recently marketed in North America.

Biophysical Phenomena↗

[Observer agreement in the measurement of leg length].

The lower limb length measurement is an important element for the diagnosis of mechanical or structural lumbar pain. Also it has been used for referral pain associated with hip or knee osteoarthritis or the groin and suprapubic areas. The aims of the present study were: 1) to measure the intra and inter observers variation; 2) to measure the intra-method variation using two different techniques for lower limb length measurement, one called the "apparent measure" (9) and comparing both with the radiological measurement technique. Two medical doctors, training on the techniques for lower limb measurement, performed the measurements. The exclusion criteria were flexion deformity of the hip or an overweight greater than 20% over the mean weight expected according to age and sex. A correlation coefficient and its 95% confidence interval (CI) were estimated, one tail test (Ho: r = 0.75). Seventeen patients fulfilled the inclusion criteria, 15 females and two males. The mean age was 35.8 years +/- 13.0 (SD). The correlation coefficient for the inter-observers variation using the "apparent measure" was 0.99 (CI = 0.985) and for the difference between legs it was 0.88 (CI = 0.10). The inter-observers variation for lower limb length measurement using the technique of "real measure" showed a correlation coefficient of 0.77 (CI = 0.95) and for the difference in length between legs it was 0.99 (CI = 0.85). The intra-observer correlation coefficient was 0.95 (CI = 0.85). The correlation coefficient for the inter-observer using the X-ray pictures was 0.98 (CI = 0.92).(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

The Glasgow Dyspepsia Severity Score--a tool for the global measurement of dyspepsia.

OBJECTIVE: There is currently no reliable tool for providing a global measurement of the severity of dyspepsia in patients with a variety of upper gastrointestinal disorders. We have designed a questionnaire which records frequency of symptoms, effect on routine activities, time off work, frequency of medical consultations, clinical investigations and use of over-the-counter and prescribed medications. The objective of the paper was to assess this questionnaire with respect to reproducibility, validity, responsiveness and performance time. METHODS AND RESULTS: For intra-observer variation, one author interviewed 50 subjects (25 males) including 20 healthy volunteers and 30 with a variety of upper gastrointestinal pathologies. The interview was repeated one week later by the same author who was blinded to the dyspepsia score for the first interview. The second author, who was blinded to the diagnoses and subject identity, scored all the questionnaires. The mean dyspepsia score was 6.78 on Day 1 and was similar at 6.80 on Day 2. The coefficient of variation between Days 1 and 2 was 2%. For inter-observer variation, 30 patients with non-ulcer dyspepsia (NUD) were interviewed by one author and the interview was repeated on a separate occasion within 24 h by a second author who was blinded to the score from the first interview. The mean dyspepsia score for the first author was 10.7 and for the second author 10.9 with a coefficient of variation between the two authors of 8%. Validity was assessed by comparing the dyspepsia scores in healthy volunteers and patients with upper gastrointestinal diseases. The mean score in 80 healthy volunteers was 1.16 (range: 0-7) and was significantly higher in 70 duodenal ulcer (DU) patients (mean score 11.1, range: 6-16) and 80 NUD patients (mean score 10.5, range: 6-17) (P < 0.001 for both vs. healthy volunteers). Responsiveness was assessed by comparing dyspepsia scores before and one year after eradication of Helicobacter pylori infection in 42 DU patients. The mean dyspepsia score before eradication was 11.4 (range: 6-16) and fell to 1.33 (range: 0-11) one year after eradication (P < 0.001). The mean time taken to complete 150 questionnaires was 4 min (range: 3-5.5 min). CONCLUSION: This new questionnaire for assessing the severity of dyspepsia is highly reproducible and has high validity and responsiveness. In addition, it is simple and rapid to perform. It provides a valuable tool for assessing the response to treatment in patients with dyspepsia.

Adolescent↗

Rational Prescribing in Primary care (RaPP): process evaluation of an intervention to improve prescribing of antihypertensive and cholesterol-lowering drugs.

BACKGROUND: A randomised trial of a multifaceted intervention for improving adherence to clinical practice guidelines for the pharmacological management of hypertension and hypercholesterolemia increased prescribing of thiazides, but detected no impact on the use of cardiovascular risk assessment tools or achievement of treatment targets. We carried out a predominantly quantitative process evaluation to help explain and interpret the trial-findings. METHODS: Several data-sources were used including: questionnaires completed by pharmacists immediately after educational outreach visits, semi-structured interviews with physicians subjected to the intervention, and data extracted from their electronic medical records. Multivariate regression analyses were conducted to explore the association between possible explanatory variables and the observed variation across practices for the three main outcomes. RESULTS: The attendance rate during the educational sessions in each practice was high; few problems were reported, and the physicians were perceived as being largely supportive of the recommendations we promoted, except for some scepticism regarding the use of thiazides as first-line antihypertensive medication. Multivariate regression models could explain only a small part of the observed variation across practices and across trial-outcomes, and key factors that might explain the observed variation in adherence to the recommendations across practices were not identified. CONCLUSION: This study did not provide compelling explanations for the trial results. Possible reasons for this include a lack of statistical power and failure to include potential explanatory variables in our analyses, particularly organisational factors. More use of qualitative research methods in the course of the trial could have improved our understanding.

Journal Article↗

Characterisation of the response of male broiler chickens to diets of various protein and energy contents.

1. The regression of body-weight gain in chicks up to three weeks of age on the linear and quadratic effects of dietary energy and protein contents accounted for approximately 50% of the observed variation. 2. The same model accounted for only 18% of the variation in food consumption, but 68% of the variation in food utilisation (food consumption:body-weight gain ratio). 3. Including initial body weight in the body-weight gain model increased the explained variation from 50 to 65% of the observed variation. 4. The coefficients for the linear and quadratic effects of dietary protein and energy were significantly different from zero for food utilisation, indicating that the response to dietary protein and energy conforms with the 'law of diminishing increments'. 5. Carcass lipid was inversely proportional to carcass water; both showed a linear response to dietary protein concentration.

Animals↗

[Coding of cause of death for mortality statistics--a comparison with results of coding by various statistical offices of West Germany and West Berlin].

1.136 death certificates representing all 1985 Bremen cardiovascular deaths and a 50%-sample of non-cardiovascular deaths in the age group 25-69 years were analyzed for reliability of nosologists' coding according to ICD-coding rules (9th revision). The 1.136 photocopied death certificates were used to assess intra-observer-variation in Bremen and to determine inter-observer-variation among 7 nosologists from 6 different State Statistical Offices and the Federal Statistical Office. Intra-observer-agreement in Bremen was found to be similar to the results presented in a comparable US-study: Bremen: 92.1%; Curb et al. 1983: 94.8%-96.1%; 3-digit-ICD-Code. Inter-observer-agreement was found to be much lower in Germany than in two US-studies: 3 coders agreeing on 3-digit-ICD-Code: Bremen: 67.7% (average, 3 coders out of 7); Curb et al.: 90.2% (3 coders); 3 coders agreeing on 4-digit-ICD-Code: Bremen: 61.5%; NCHS 1980: 90.3%. Agreement-rates were also much lower in Germany than in the USA (Curb et al.) when particular disease groups were analysed: Ischaemic heart disease (ICD 410-414): Bremen: 82.7% (average); USA: 97.2%; cerebrovascular disease (ICD 430-438): Bremen 65.6% (average); USA: 93.2%; neoplasms (ICD 140-239): Bremen: 94.0% (average); USA: 97.8%. We conclude that training, individual characteristics of nosologists, and other factors may cause important artifacts when comparing German mortality statistics on a regional level or during different time intervals.

Berlin↗

Precision of Larsen grading of radiographs in assessing progression of rheumatoid arthritis in individual patients.

A study was designed to evaluate observer variation in the assessment of radiographic deterioration of individual patients using the Larsen grading system. Radiographs of hands and feet of 52 patients were assessed by three observers. Each patient had paired films taken one year apart which were assessed together for change in score. To assess within-observer variation each set of films was read twice by all observers. The average progression was 11.6 (SD 9.0). Analysis of the source of variation showed the single observer replication SD to be 3.7 but that for different observers to be 5.5. This may be interpreted as indicating that to achieve 95% confidence of detecting a true change an increase in Larsen score of 8 is required if the same observer assesses or up to 11 if a different observer assesses.

Arthritis, Rheumatoid↗

The reduction of inter- and intra-observer variability for defining regions of interest in nuclear medicine.

The use of region-of-interest (ROI) techniques to quantify data obtained in radionuclide images is commonplace. However, the reproducibility of quantitation due to inter- and intra-observer variations using particular methods of deriving ROIs is often not appreciated. We examined such variations in the results obtained by four independent observers of varying experience using four methods of depicting a ROI about an organ. The set of image data consisted of renal scans with varying target-to-background ratios, and the ROI facilities included two edge-detection methods. The results indicated that, once observers were experienced with edge-detection methods, a lower inter- and intra-observer variation could be achieved, although the technique of 'shrinking' a ROI about a subjectively chosen display level was reasonably satisfactory. In terms or reproducibility, the least satisfactory method of depicting a ROI was the commonly used manually guided 'bug' around arbitrarily chosen display levels representing the boundary of an organ.

Humans↗

An improved procedure to quantify tumour vascularity using true colour image analysis. Comparison with the manual hot-spot procedure in a human melanoma xenograft model.

In a number of recent papers, the degree of tumour vascularization has been described as a promising new prognostic factor. Methods for the assessment of vascular density involve immunohistochemical staining of the vasculature, followed by counting the number of vessel profiles in the angiogenic hot spot. One of the problems of this procedure is the selection of the angiogenic hot spot, which has been described as being subject to inter-observer variation. In this study, the value of true colour image analysis in reducing inter-observer variation has been assessed. Highly (MV3) and poorly (M14) vascularized human melanoma xenografts were used to evaluate the image analysis procedure, and the image analysis results were compared with results from the conventional manual hot-spot procedure. Assessment by image analysis was performed on measurement fields covering the entire tumour tissue specimens rather than on a single hot-spot field. Also, by selecting the most densely vascularized area from all fields assessed by the semi-automatic procedure, it was possible to objectify the hot spot selection (automated hot-spot procedure). Manual assessment showed a good correlation between two independent observers for MV3 xenografts (r = 0.74, P = 0.014), but a poor correlation for M14 xenographs (r = 0.4, P > 0.05). Automated assessment by different operators showed good correlations for both MV3 xenografts (r = 0.99, P < 0.001) and M14 xenografts (r = 0.80, P = 0.006). It is concluded that although both manual vessel counting and semi-automated image analysis can differentiate between the level of vascularization in the two types of xenograft (P < 0.001 for both methods), the automated method is favourable in that it showed no significant inter-observer effects. In M14 xenografts, the manual hot-spot vessel densities did not correlate well with the automated hot-spot densities (r = 0.27, P > 0.05), indicating that selection of angiogenic hot spots in this tumour type is indeed subject to observer bias. The automated hot-spot vessel densities were a reliable indicator of overall tumour vessel density in both tumour types. Image analysis allows analysis of vessel subclasses based on morphological criteria such as vessel profile area or diameter. In the model system used, the discrimination between MV3 and M14 xenografts was further enhanced by selectively examining vessels with diameters between 6 and 9 microns (P < 0.0005). In conclusion, image analysis appears to offer an objective and more reproducible method to quantify tumour vascularity than manual counting of vessel profiles in the hot spot. Analysis of subclasses of vessels may further enhance the value of vessel density measurements in discriminating between tumour types differing in biological behaviour.

Animals↗

Comparison of predicted scaffold-compatible sequence variation in the triple-hairpin structure of human imunodeficiency virus type 1 gp41 with patient data.

It has been proposed that the ectodomain of human immunodeficiency virus type 1 (HIV-1) gp41 (e-gp41), involved in HIV entry into the target cell, exists in at least two conformations, a pre-hairpin intermediate and a fusion-active hairpin structure. To obtain more information on the structure-sequence relationship in e-gp41, we performed in silico a full single-amino-acid substitution analysis, resulting in a Fold Compatible Database (FCD) for each conformation. The FCD contains for each residue position in a given protein a list of values assessing the energetic compatibility (ECO) of each of the 20 natural amino acids at that position. Our results suggest that FCD predictions are in good agreement with the sequence variation observed for well-validated e-gp41 sequences. The data show that at a minECO threshold value of 5 kcal/mol, about 90% of the observed patient sequence variation is encompassed by the FCD predictions. Some inconsistent FCD predictions at N-helix positions packing against residues of the C helix suggest that packing of both peptides may involve some flexibility and may be attributed to an altered orientation of the C-helical domain versus the N-helical region. The permissiveness of sequence variation in the C helices is in agreement with FCD predictions. Comparison of N-core and triple-hairpin FCDs suggests that the N helices may impose more constraints on sequence variation than the C helices. Although the observed sequences of e-gp41 contain many multiple mutations, our method, which is based on single-point mutations, can predict the natural sequence variability of e-gp41 very well.

Amino Acid Sequence↗

Peri-acetabular radiolucent lines: inter- and intra-observer agreement on post-operative radiographs.

Peri-acetabular radiolucent lines (RLLs) seen on "early" post-operative radiographs have been identified as a potential predictor of long-term implant performance. This study examines the inter- and intra-observer variation encountered when assessing such radiographs. Four consultant orthopaedic surgeons assessed the presence, extent and width of RLLs in 220 radiographs performed on 50 patients taken one to two weeks, six weeks, six months and one year following surgery. Inter-observer agreement was fair at 7-14 days but improved to moderate to good in films at six and 12 months. Intra-observer agreement was moderate to good at 7-10 days but again improved to good at 6 and 12 months. When only the presence or absence of RLLs was considered, both inter-observer and intra-observer agreement improved for both the six-month and one-year radiographs. This experiment shows that caution must be used for the interpretation of RLLs on hip radiographs taken during the very early post-operative period. We recommend that films taken at least six weeks to six months following surgery should be used for assessment to reduce observer variation. For optimum results, a single experienced observer should do the assessment with a simple classification.

Aged↗

Statistical methods in epidemiology. v. Towards an understanding of the kappa coefficient.

PURPOSE: This paper introduces readers to the problem of measuring interrater agreement in observer variation studies. The most usual statistic to quote is the kappa coefficient which measures agreement having corrected for chance. METHOD: The kappa coefficient for measuring agreement between two observers is introduced. Some pointers are given to determining sample size estimation. RESULTS: Some properties of the kappa coefficient are illustrated by taking examples from the author's teaching experiences. CONCLUSION: The kappa coefficient is recommended for measuring agreement in observer variation studies.

Epidemiology↗