PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Observer Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Observer variability in endometrial cytology using kappa statistics.

AIMS: To analyse the current state of endometrial cytology practice by determining the interobserver variability in grading the atypicality of endometrial cells using kappa statistics. METHODS: A series of 70 clusters of benign, hyperplastic, and malignant endometrial cells on cytology specimens were examined by 19 experienced Japanese cytopathologists. They assigned each cluster to one of three diagnostic categories according to increasing cellular atypicality: negative for, suspicious of, and positive for malignancy. Each observer used their subjective judgement in grading without receiving any discriminatory criteria. RESULTS: Most classified the series of 70 clusters of endometrial cells into 33 negative, 20 suspicious, and 17 positive grades. However, observer variation was considerable, kappa statistics showing unfavourable overall agreement (kappa = 0.36). Although there was good agreement on negative and positive (kappa = 0.46 and 0.47, respectively), poor reproducibility occurred in grading suspicious ((kappa = 0.15). CONCLUSIONS: Currently, Japanese cytopathologists are not reliable in grading the atypicality of endometrial cells. This could be attributed mainly to the situation in endometrial cytology practice in which there are no generally accepted criteria for defining cells in which malignancy is suspected.

Adenocarcinoma↗

[Presentation of a grid for computer analysis for compilation of histopathologic lesions in chronic viral hepatitis C. Cooperative study of the METAVIR group].

The METAVIR group, including 10 pathologists with a specialization in liver pathology, has been formed in order to discuss on problems related to chronic viral hepatitis C. Cooperative studies have been planed. At first, observer variation in the assessment of features of chronic viral hepatitis C has been performed. This study has led to the proposal of a coded form for pathological features in liver biopsy of chronic viral hepatitis C. This form will be used to study a large number of biopsies from patients included in clinical trials. In the next future, the prevalence of every of these elementary features will be assessed, as well as correlations between clinical, biological and pathological data.

Biopsy↗

Quantitative characterization of color Doppler images: reproducibility, accuracy, and limitations.

A computer-based quantitative analysis for color Doppler images of complex vascular formations is presented. The red-green-blue-signal from an Acuson XP10 is frame-grabbed and digitized. By matching each image pixel with the color bar, color pixels are identified and assigned to the corresponding flow velocity (color value). Data analysis consists of delineation of a region of interest and calculation of the relative number of color pixels in this region (color pixel density) as well as the mean color value. The mean color value was compared to flow velocities in a flow phantom. The thyroid and carotid artery in a volunteer were repeatedly examined by a single examiner to assess intra-observer variability. The thyroids in five healthy controls were examined by three experienced physicians to assess the extent of inter-observer variability and observer bias. The correlation between the mean color value and flow velocity ranged from 0.94 to 0.96 for a range of velocities determined by pulse repetition frequency. The average deviation of the mean color value from the flow velocity was 22% to 41%, depending on the selected pulse repetition frequency (range of deviations, -46% to +66%). Flow velocity was underestimated with inadequately low pulse repetition frequency, or inadequately high reject threshold. An overestimation occurred with inadequately high pulse repetition frequency. The highest intra-observer variability was 22% (relative standard deviation) for the color pixel density, and 9.1% for the mean color value. The inter-observer variation was approximately 30% for the color pixel density, and 20% for the mean color value. In conclusion, computer assisted image analysis permits an objective description of color Doppler images. However, the user must be aware that image acquisition under in vivo conditions as well as physical and instrumental factors may considerably influence the results.

Algorithms↗

Comparative efficiencies of randomized concentration- and dose-controlled clinical trials.

OBJECTIVE: To compare efficiencies of randomized dose- and concentration-controlled trials (RDCT and RCCT) for estimating the parameters of concentration-effect relationships. RATIONALE: In 1991 Sanathanan and Peck (Controlled Clin Trials 1991;12:780-94) suggested that estimation by RDCT is biased and much less efficient than analysis by RCCT. Their conclusion was based on a pharmacodynamic model that characterizes the effect of theophylline in subjects with asthma, in which the response was related linearly to a limited range of concentrations and independent of concentration otherwise. Therefore it was intended to explore whether the conclusion of Sanathanan and Peck applied to other pharmacodynamic models. RESULTS: The results of Sanathanan and Peck were confirmed for the restricted linear, baseline-plateau model: with large pharmacokinetic and no pharmacodynamic variability, RCCT was 3.1 times more efficient than RDCT. However, under the same conditions, the efficiency of RCCT exceeded that of RDCT only 1.5 and 1.2 times when response was related, without restrictions, to concentration and log concentration, respectively. Moreover, in the presence of even moderate pharmacodynamic variability, the ratio of RCCT/RDCT efficiencies did not exceed 1.30 and 1.08, respectively. The parameters estimated by RDCT with these two models were not biased. Finally, in the presence of interindividual variability of the median effective concentration (EC50), pharmacokinetic variability did not affect the observed variation of the parameters in the log-linear pharmacodynamic relationship. CONCLUSIONS: RCCT generally estimates pharmacodynamic parameters with an efficiency that is not much higher than, or even similar to, those yielded by RDCT. Therefore statistical benefits often do not call for the application of RCCT. However, sometimes its use should be seriously considered, particularly for drugs having small therapeutic indexes or when the baseline and plateau of the response occur near the therapeutic region of concentrations.

Dose-Response Relationship, Drug↗

Interoperator test for anatomical annotation of earprints.

As part of the Forensic Ear Identification (FearID) research project, which aims to obtain estimators for the strength of evidence of earmarks found on crime scenes, a large database of earprints (over 1200 donors) has been collected. Starting from a knowledge-based approach where experts add anatomical annotations of minutiae and landmarks present in prints, comparison of pairs of prints is done using the method of Vector Template Matching (VTM). As the annotation process is subjective, a validation experiment was performed to study its stability. Comparing prints on the basis of VTM, it appears that there are interoperator effects, individual operators yielding significantly more consistent results when annotating prints than different operators. The operators being well trained and educated, the observed variation on both clicking frequency and choice of annotation points suggests that implementation of the above is not the best way to go about objectifying earprint comparison. Processes like the above are relevant for any forensic science dealing with identification (e.g., of glass, tool marks, fibers, faces, fingers, handwriting, speakers) where manual (nonautomated) processes play a role. In these cases, results may be operator dependent and the dependencies need to be studied.

Data Collection↗

Plasma cell quantification in bone marrow by computer-assisted image analysis.

BACKGROUND: Minor and major criteria for the diagnosis of multiple meloma according to the definition of the WHO classification include different categories of the bone marrow plasma cell count: a shift from the 10-30% group to the > 30% group equals a shift from a minor to a major criterium, while the < 10% group does not contribute to the diagnosis. Plasma cell fraction in the bone marrow is therefore critical for the classification and optimal clinical management of patients with plasma cell dyscrasias. The aim of this study was (i) to establish a digital image analysis system able to quantify bone marrow plasma cells and (ii) to evaluate two quantification techniques in bone marrow trephines i.e. computer-assisted digital image analysis and conventional light-microscopic evaluation. The results were compared regarding inter-observer variation of the obtained results. MATERIAL AND METHODS: Eighty-seven patients, 28 with multiple myeloma, 29 with monoclonal gammopathy of undetermined significance, and 30 with reactive plasmocytosis were included in the study. Plasma cells in H&E- and CD138-stained slides were quantified by two investigators using light-microscopic estimation and computer-assisted digital analysis. The sets of results were correlated with rank correlation coefficients. Patients were categorized according to WHO criteria addressing the plasma cell content of the bone marrow (group 1: 0-10%, group 2: 11-30%, group 3: > 30%), and the results compared by kappa statistics. RESULTS: The degree of agreement in CD138-stained slides was higher for results obtained using the computer-assisted image analysis system compared to light microscopic evaluation (corr.coeff. = 0.782), as was seen in the intra- (corr.coeff. = 0.960) and inter-individual results correlations (corr.coeff. = 0.899). Inter-observer agreement for categorized results (SM/PW: kappa 0.833) was in a high range. CONCLUSIONS: Computer-assisted image analysis demonstrated a higher reproducibility of bone marrow plasma cell quantification. This might be of critical importance for diagnosis, clinical management and prognostics when plasma cell numbers are low, which makes exact quantifications difficult.

Algorithms↗

Prospective evaluation of 2 acute graft-versus-host (GVHD) grading systems: a joint Société Française de Greffe de Moëlle et Thérapie Cellulaire (SFGM-TC), Dana Farber Cancer Institute (DFCI), and International Bone Marrow Transplant Registry (IBMTR) prospective study.

The most commonly used grading system for acute graft-versus-host disease (aGVHD) was introduced 30 years ago by Glucksberg; a revised system was developed by the International Bone Marrow Transplant Registry (IBMTR) in 1997. To prospectively compare the 2 classifications and to evaluate the effect of duration and severity of aGVHD on survival, we conducted a multicenter study of 607 patients receiving T-cell-replete allografts, scored weekly for aGVHD in 18 transplantation centers. Sixty-nine percent of donors were HLA-identical siblings and 28% were unrelated donors. The conditioning regimen included total body irradiation in 442 (73%) patients. The 2 classifications performed similarly in explaining variability in survival by aGVHD grade, although the Glucksberg classification predicted early survival better. There was less physician bias or error in assigning grades with the IBMTR scoring system. With either system, only the maximum observed grade had prognostic significance for survival; neither time of onset nor progression from an initially lower grade of aGVHD was associated with survival once maximum grade was considered. Regardless of scoring system, aGVHD severity accounted for only a small percentage of observed variation in survival. Validity of these results in populations receiving peripheral blood transplants or nonmyeloablative conditioning regimens remains to be tested.

Acute Disease↗

Volume estimation of extensor muscles of the lower leg based on MR imaging.

Magnetic resonance imaging can be used to measure the muscle volume of a given muscle or muscle group. The purpose of this study was to determine both the intra- and inter-observer variation of the manually outlined volume of the extensor muscles (tibialis anterior, extensor digitorum longus and extensor hallucis longus), to estimate the minimum number of slices needed for these calculations and to compare estimates of volume based on an assumed conic shape of the muscles with that of an assumed cylindrical shape, the calculation in both cases based on the Cavalieri principle. Eleven young and healthy subjects (4 women and 7 men, age range 24-40 years) participated. Magnetic resonance imaging of the left leg was obtained on a 1.5-T MR system using a knee coil (receive only). A total of 50 consecutive slices were obtained beginning 10 cm below the caput fibula sin. and proceeding distally with a slice thickness of 1.5 mm without gap. The intra-class correlation coefficient (ICC) was used to calculate the relative reliability (interval from 0 to 1.0). A high reliability for both intra- and inter-reliability was observed (ICC 0.98 and 1.0). The difference was only 0.004% between calculations based on measurement of all 50 slices with respect to 8 slices equally distributed along the muscle group. No difference was found between the two different volumetric assumptions in the Cavalieri principle. The manually outlining of extensor muscles volumes was reliable and only 8 slices of the calf were needed. No difference was seen between the two used mathematical calculations.

Adult↗

Expert system support using a Bayesian belief network for the classification of endometrial hyperplasia.

Accurate morphological classification of endometrial hyperplasia is crucial as treatments vary widely between the different categories of hyperplasia and are dependent, in part, on the histological diagnosis. However, previous studies have shown considerable inter-observer variation in the classification of endometrial hyperplasias. The aim of this study was to develop a decision support system (DSS) for the classification of endometrial hyperplasias. The system used a Bayesian belief network to distinguish proliferative endometrium, simple hyperplasia, complex hyperplasia, atypical hyperplasia and grade 1 endometrioid adenocarcinoma. These diagnostic outcomes were held in the decision node. Four morphological features were selected as diagnostic clues used routinely in the discrimination of endometrial hyperplasias. These represented the evidence nodes and were linked to the decision node by conditional probability matrices. The system was designed with a computer user interface (CytoInform) where reference images for a given clue were displayed to assist the pathologist in entering evidence into the network. Reproducibility of diagnostic classification was tested on 50 cases chosen by a gynaecological pathologist. These comprised ten cases each of proliferative endometrium, simple hyperplasia, complex hyperplasia, atypical hyperplasia and grade 1 endometrioid adenocarcinoma. The DSS was tested by two consultant pathologists, two junior pathologists and two medical students. Intra- and inter-observer agreement was calculated following conventional histological examination of the slides on two occasions by the consultants and junior pathologists without the use of the DSS. All six participants then assessed the slides using the expert system on two occasions, enabling inter- and intra-observer agreement to be calculated. Using unaided conventional diagnosis, weighted kappa values for intra-observer agreement ranged from 0.645 to 0.901. Using the DSS, the results for the four pathologists ranged from 0.650 to 0.845. Both consultant pathologists had slightly worse weighted kappa values using the DSS, while both junior pathologists achieved slightly better values using the system. The grading of morphological features and the cumulative probability curve provided a quantitative record of the decision route for each case. This allowed a more precise comparison of individuals and identified why discordant diagnoses were made. Taking the original diagnoses of the consultant gynaecological pathologist as the 'gold standard', there was excellent or moderate to good inter-observer agreement between the 'gold standard' and the results obtained by the four pathologists using the expert system, with weighted kappa values of 0.586-0.872. The two medical students using the expert system achieved weighted kappa values of 0.771 (excellent) and 0.560 (moderate to good) compared to the 'gold standard'. This study illustrates the potential of expert systems in the classification of endometrial hyperplasias.

Bayes Theorem↗

A comparative study of radionuclide venography and contrast venography in the diagnosis of deep venous thrombosis.

BACKGROUND: The value of the radionuclide blood pool venogram in detecting deep venous thrombosis (DVT) has to date been inadequately evaluated. This is despite its lower complication rate than the gold standard of contrast X-ray venography. AIMS: To compare the relative accuracy and inter observer variability of radionuclide blood pool and X-ray contrast venography as well as evaluate previous literature on radionuclide venography. METHODS: Prospective comparison of radionuclide and contrast venography was performed in 39 patients. Sensitivity and specificity of radionuclide venography were compared to contrast venography and confidence intervals were measured using standard error calculations. A meta-analysis of previous studies was also performed. RESULTS: Significant inter observer variation in reports was present in both radionuclide (37%) and contrast (22%) venograms. Using consensus reports sensitivity of radionuclide venography was 87% compared to contrast venography and specificity was 83%. These results are similar to those obtained in previous studies. Furthermore, sensitivity in specificity in the proximal veins were 90% and 92% respectively which were superior to sensitivity and specificity in the distal veins where it was 74% and 90% respectively. CONCLUSION: The radionuclide venogram appears accurate in the proximal veins and in excluding but not diagnosing distal venous thrombosis.

Adult↗

Daily morbidity records: recall and reliability.

Methodological issues concerning the collection and analysis of daily morbidity data in community studies in developing countries are discussed. The effects of recall period and inter-observer variation on symptom prevalence are considered in the context of a longitudinal study in The Gambia, in which prevalence fell by about half over 1-week's recall. In the same study, many infant-days were recorded separately on two occasions, allowing an assessment of reliability in this type of morbidity diary data. The implications of these findings both in terms of data quality and cost-effectiveness are discussed, with the conclusion that weekly interviews examining the previous week's morbidity on a day-by-day basis are operationally optimal.

Adverse Drug Reaction Reporting Systems↗

Use of monitoring software to improve the measurement of carotid wall thickness by B-mode imaging.

METHODOLOGY: High-resolution B-mode imaging is a reliable, easily performed and non-invasive means of studying atherosclerosis in superficial blood vessels. Recently it has been used for in vivo studies on the thickness of the common carotid artery wall. It is very sensitive, although the results of practical investigations are highly dependent on both the operator and the direction and angle of ultrasound beams directed towards the vessel. PROTOCOL: We have assessed inter- and intra-observer reproducibility of the measurement of common carotid artery wall thickness in 13 subjects, using two procedures. The first was a standard echographical investigation. In the second procedure, the principal parameters recorded from the first investigation were used to reposition the beam with the same incident angle. RESULTS: Intra-observer variability (correlation coefficient, r = 0.61 for procedure 1 and r = 0.77 for procedure 2) and inter-observer variation (r = 0.58 for procedure 1 and r = 0.71 for procedure 2) were reduced when the second investigation was assisted by reproducibility software. CONCLUSIONS: The proposed method is a reliable and reproducible way of assessing combined intimal and medial wall thickness in the common carotid artery. It may be possible to improve reproducibility using specific software to aid the operator. Since the intimal and medial thickness of the common carotid artery appears to be a sensitive marker of vascular risk, the proposed standardized method of measuring these parameter may allow early detection and assessment of changes.

Aged↗

The value of MR angiography techniques in the detection of head and neck paragangliomas.

OBJECTIVE: The objective of this study was to compare three-dimensional phase-contrast angiography (3D PCA), 2D time-of-flight (2D TOF), and 3D TOF magnetic resonance (MR) angiography and a proton density weighted technique in terms of their ability to detect head and neck paragangliomas. MATERIALS AND METHODS: 14 patients with 29 paragangliomas were examined at 1.5 T. Three MR angiography sequences (3D PCA, 2D TOF, and multi-slab 3D TOF) and a proton density (PD) weighted sequence were reviewed by four neuroradiologists. The gold standard was digital subtraction angiography. Presence of tumor was assessed in five grades of confidence. Sensitivity and specificity were calculated after dichotomizing the results. Data was analyzed using the logistic regression method. RESULTS: Mean sensitivity and specificity for the four observers were for PD: 72%/97%, for 3D PCA: 75%/90%, for 2D TOF: 66%/93%, and for 3D TOF: 90%/92%. Sensitivity was significantly better for 3D TOF MRA (P < 0.001). No substantial between-observer variation for tumor detection was present. CONCLUSION: Our results demonstrate that, using 3D TOF MRA, paragangliomas in the head and neck region can be detected with high sensitivity and specificity. Further investigation is necessary to judge the value of 3D TOF MR angiography against fat suppressed contrast enhanced T1 weighted and fat suppressed T2 weighted MR sequences to find the optimal imaging sequence for paragangliomas.

Adult↗

Biostatistical aspects of outcome evaluation using TISS-28.

Quantification of therapeutic activities using the Therapeutic Intervention Scoring System (TISS) is an alternative approach to evaluate outcome of patients in intensive care. The reason for using cumulative TISS points is to integrate various adverse events (except mortality) according to the amount of therapeutic effort that they require. The reduced version of TISS with 28 items (TISS-28) allows a reliable assessment of therapeutic activities with limited observer variation, provided that an exact description of all items is given. Measurements can be validated by correlations with established severity of -disease classification systems such as APACHE II. Cumulative TISS-28 values correlate well with length of ICU stay (r = 0.98). On average, 27.2 points/day can be expected in an unselected mixed surgical ICU. Those who die can be included in non-parametric analyses of cumulative TISS values by allocation of arbitrary high values. Quantification of therapeutic interventions is a sensitive measure of outcome in patients who require intensive care but have a low risk of mortality. The usefulness of economic analysis further supports its clinical application.

APACHE↗

Precision in assessment of osteoporosis from spine radiographs.

Inter- and intra-observer variation in spine radiographs of 100 osteoporotic women and longitudinal change in roentgenologic status after 1 year of antiosteoporotic treatment were assessed. The method applied was naked eye inspection, and a score system estimating severity of fractures - vertebral deformation score (VDS). Agreement was assessed by the Kappa coefficient corrected for agreement by change. The results showed a satisfactory inter- and intraobserver agreement for wedge (Kappa = 0.72 and 0.90) and compression fractures (Kappa = 0.60 and 0.92). The method proved less reliable for endplate fractures (Kappa = 0.39 and 0.73). We think that the method of investigation is well suited for monitoring treatment effects in longitudinal studies.

Aged↗

Reliability of the reported size of removed colorectal polyps.

BACKGROUND: The size of colorectal polyps is important in the clinical management of these lesions. AIM: To audit the accuracy in calculating the size of "polyps" by various specialists. MATERIALS AND METHODS: Eighteen pathologists and four surgeons measured, with a conventional millimetre ruler, the largest diameter of 12 polyp phantoms. The results of two independent measurements (two weeks apart) were compared with the gold standard-size assessed at The Royal Institute of Technology, Sweden. RESULTS: Thirty-one percent (83/264-trial 1) and 33% (88/264-trial 2) of the measurements underestimated or overestimated the gold standard size by >1 mm. Of the 22 experienced participants, 95% (21/22-trial 1) and 91% (20/22-trial 2) misjudged by >1 mm the size of one or more polyps. Values given by 13 participants (4.9%) in trial 1 and by 15 participants (5.7%) in trial 2, differed by > or = +/-4 mm from the gold standard size. In addition, a big difference between the highest and the lowest values was recorded in some polyps (up to 11.4 mm). Those disparate values were regarded as a human error in reading the scale on the ruler. CONCLUSION: Using a conventional ruler (the tool of pathologists worldwide) unacceptably high intra-observer and inter-observer variations in assessing the size of polyp-phantoms was found. The volume and the shape of devices, as well as human error in reading the scale of the ruler were confounding factors in size assessment. In praxis, the size is crucial in the management of colorectal polyps. Considering the clinical implications of the results obtained, the possibility of developing a method that will allow assessment of the true size of removed clinical polyps is being explored.

Colonic Diseases↗

Computerized quantitative pathology for the grading of dysplasia in surveillance biopsies of Barrett's oesophagus.

Decision models for surveillance of Barrett's oesophagus (BO) are governed by the grade of dysplasia on endoscopic biopsy, but subjective grading is prone to observer variation. Computerized morphometry and immunoquantitation can objectively discriminate between different grades of dysplasia in oesophagectomy specimens with BO. The present study evaluated the feasibility of such quantitative analysis on surveillance biopsies of BO. Biopsy criteria for quantitative analysis were defined, excluding 101 (21%) of 472 archival BO surveillance biopsies. In the remaining haematoxylin and eosin (H&E) sections, 105 areas that distinctively displayed no dysplasia (ND), low-grade dysplasia (LGD) or high-grade dysplasia (HGD) were demarcated. Agreement on double-blind examination by two experienced pathologists was reached in 66 areas (63%; kappa: 0.44). For 21 ND/LGD and 11 LGD/HGD disagreement areas, corresponding sections for p53 and Ki67 immunohistochemistry were available. The best combination of two discriminating features was stratification index (SI) with p53 area % for ND versus LGD (89% correct classification), and SI with Ki67 area % for LGD versus HGD (91% correct classification). Fifteen of the 21 ND/LGD disagreement areas could be classified uniquely as either ND or LGD by SI and p53, and eight of the 11 LGD/HGD disagreement areas as either LGD or HGD by SI and Ki67. Correlation coefficients for repeated measurements of SI, Ki67, and p53 by the same observer were 0.94, 0.92, and 0.86, and by two independent observers 0.86, 0.93, and 0.92, respectively. Computerized quantitative pathology on BO surveillance biopsies is feasible provided that well-defined biopsy criteria are used. Using a combination of features associated with cellular differentiation and proliferation, such as SI, p53, and Ki67, quantitative pathological analysis assists in reducing diagnostic variability in the grading of dysplasia during surveillance of BO.

Barrett Esophagus↗

A comparison and conversion table of 'the House-Brackmann facial nerve grading system' and 'the Yanagihara grading system'.

A comparison between the House-Brackmann facial nerve grading system (the HB-system) and the Yanagihara grading system (the Y-system) was studied with 199 evaluations of 62 cases of postoperative unilateral acoustic neuroma. In the beginning, an original draft of the conversion table was formulated according to the 199 evaluations, in which, 0-6, 8-14, 16-20, 22-28, 30-38, and 40 points in the Y-system were matched with grade VI, V, IV, III, II, and I in the HB-system respectively. The result of the present study for prediction of sequelae showed that it was not necessary to consider the sequelae in a conversion table. And more, the study of an inter-observer variation showed that the lower and upper limits of the scores in the Y-system may shift within about a 2-point range in the draft table. From these aspects, a newly revised conversion table was proposed as a revised conversion table, in which, 0-6, 8-14, 16-22, 24-30, 32-38, and 40 points in the Y-system were matched with grade VI, V, IV, III, II, and I in the HB-system. This revised conversion table is much easier to remember for clinical use, because the score of the lower limit of each grade is simply in multiples of 8, that is, 0, 8, 16, 24, 32 and 40.

Facial Nerve↗