PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Observer Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Resting energy expenditure should be measured in patients with cirrhosis, not predicted.

Measurements of resting energy expenditure (REE) can be used to determine energy requirements. Prediction formulae can be used to estimate REE but have not been validated in cirrhotic patients. REE was measured, by indirect calorimetry, in 100 cirrhotic patients and 41 comparable healthy volunteers, and the results compared with estimates predicted using the Harris-Benedict, Schofield, Mifflin, Cunningham, and Owen formulae, and the disease-specific Müller formula. The mean (+/- 1 SD) measured REE in the healthy volunteers (1,590 +/- 306 kcal/24 h) was significantly greater than the mean Harris-Benedict, Mifflin, Cunningham, and Owen predictions but comparable with the mean Schofield prediction; individual predicted values varied widely from measured values (95% limits of agreement, -460 to +424 kcal). The mean measured REE in the cirrhotic patients was significantly greater than in the healthy volunteers (23.2 +/- 3. 8 cf 21.9 +/- 2.9 kcal/kg/24 h; P <.05). The mean measured REE in the cirrhotic patients (1,660 +/- 337 kcal/24 h) was significantly different from mean predicted values (Harris-Benedict, 1,532 +/- 252 kcal/24 h, P <.0001; Schofield, 1,575 +/- 254 kcal/24 h, P <.0005; Mifflin, 1,460 +/- 254 kcal/24 h, P <.0001; Cunningham, 1,713 +/- 252 kcal/24 h, P <.05; Owen, 1,521 +/- 281 kcal/24 h, P <.0001; Müller, 1,783 +/- 204 kcal/24 h, P <.0001); individual predicted values varied widely from measured values (95% limits of agreement, -632 to +573 kcal). Simple regression analysis showed that fat-free mass (FFM) was the strongest predictor of measured REE in the cirrhotic patients, accounting for 52% of the variation observed. However, a population-specific prediction equation, derived using stepwise regression analysis, which incorporated FFM, age, and Pugh's score, accounted for only 61% of the observed variation in measured REE. REE should, therefore, be measured in cirrhotic patients, not predicted.

Adult↗

Decision making in radiographic imaging.

In 1987, the U.S. Department of Health and Human Services issued guidelines for prescription of dental radiographic examinations, and although these recommendations have been reprinted in several widely circulated publications, it seems that the adoption of these guidelines is far from common, even among U.S. dental schools. The recommended criteria are founded on the existing knowledge of prevalence and progression of the most common dental diseases and on the fact that occult diseases within the jaws are uncommon. There are, however, other factors that may influence the decision on the time and extent of a radiographic examination, which may lead to deviations from the suggested guidelines. These factors include: education, peer influence, patient's preference, legal considerations, the dentist's field of interest or specialty, the training of the staff, and practice routine. The diagnostic interpretation of radiographs is far from a completely objective process, even if it is a question as simple as the presence and/or extent of a carious lesion. Numerous studies have shown a large variation among observers, both with regard to the occurrence and extent of carious lesions, in bite-wing radiographs. Caries diagnosis is only one example of many situations where significant observer variation is found. The more complex the diagnostic task, the more variation can be expected. The effect of observer variation on treatment decisions regarding carious lesions is used as an example of the problems encountered daily in the dental practice.

American Dental Association↗

The use of X-ray for lymph node determination in the axillary dissection specimen.

A new and simple method by X-ray is described for lymph node determination in the axillary specimen of breast cancer patients. X-rays were performed of the axillary specimens of 49 women with breast cancer. The number of lymph nodes visible on the X-rays were assessed by two radiologists (A and B). The number of nodes identified in the axillary specimens was reported by the pathologist independently. The method described shows a clear correlation between the mean numbers of nodes counted on the X-rays of the specimens (radiologist A 18.3, B 16.1 nodes) and the mean numbers of nodes recovered by the pathologist (18.4). No intra-observer variation was observed and only a small inter-observer variation (2.2 nodes). This method of X-ray determination of lymph nodes can be used in auditing the surgeon's accuracy in performing complete axillary dissection as well as in auditing the number of lymph nodes found by the pathologist.

Journal Article↗

Evaluation and comparison of sources of variability in the measurement of corneal thickness with ultrasonic and optical pachymeters.

Two studies were carried out to determine and compare the effects of several sources of variation on the measurement of corneal thickness using the standard optical pachymeter and three ultrasonic pachymeters. Sources of variation included: intra- and inter-session variation, inter-observer variation, left/right eye variation, and variations due to alternate settings of ultrasonic sound frequencies. It was found that the optical pachymeter had a) two to three times as much intra-session variation as that of the ultrasound pachymeters, b) significant inter-observer variation (P = 0.015), and c) significant differences between left and right eye thickness determinations (P less than 0.005). On the other hand, ultrasonic pachymeters demonstrated a) high reproducibility, b) no inter-observer variation, and c) no left/right eye variation. These results have implications for the use of pachymetry in measuring corneal thickness for radial keratotomy and other refractive surgery.

Cornea↗

Evaluation of the genetic structure of Xylella fastidiosa populations from different Citrus sinensis varieties.

Xylella fastidiosa was isolated from sweet orange plants (Citrus sinensis) grown in two orchards in the northwest region of the Brazilian state of São Paulo. One orchard was part of a germ plasm field plot used for studies of citrus variegated chlorosis resistance, while the other was an orchard of C. sinensis cv. Pêra clones. These two collections of strains were genotypically characterized by using random amplified polymorphic DNA (RAPD) and variable number of tandem repeat (VNTR) markers. The genetic diversity (H(T)) values of X. fastidiosa were similar for both sets of strains; however, H(T)(RAPD) values were substantially lower than H(T)(VNTR) values. The analysis of six strains per plant allowed us to identify up to three RAPD and five VNTR multilocus haplotypes colonizing one plant. Molecular analysis of variance was used to determine the extent to which population structure explained the genetic variation observed. The genetic variation observed in the X. fastidiosa strains was not related to or dependent on the different sweet orange varieties from which they had been obtained. A significant amount of the observed genetic variation could be explained by the variation between strains from different plants within the orchards and by the variation between strains within each plant. It appears, therefore, that the existence of different sweet orange varieties does not play a role in the population structure of X. fastidiosa. The consequences of these results for the management of sweet orange breeding strategies for citrus variegate chlorosis resistance are also discussed.

Bacterial Typing Techniques↗

Ultrasound assessment of tendons in asymptomatic volunteers: a study of reproducibility.

The purpose was to evaluate the inter-visit, inter-observer and intra-observer variation of quantitative and qualitative tendon examinations in vivo for a cohort of asymptomatic volunteers. Eleven healthy male subjects were recruited. The following tendons were assessed by ultrasonography: Achilles tendon, patellar tendon, triceps tendon, extensor pollicis longus, flexor carpi radialis and supraspinatus. For each tendon a quantitative measurement of tendon size was made at a predefined anatomical location. Two experienced sonologists, blind to one another's findings, evaluated each of the tendons independently. Each tendon was evaluated on two occasions 1 week apart. No difference was found to be attributed to variation in tendon size between visits. Inter-observer variation was a source of error with intra-subject, inter-visit measurements proving more reproducible. There was some significant variation between observers. This variation was more marked with some tendon measures than others. Inter-observer variation for triceps, flexor carpi radialis and supraspinatus was most marked. Minimum detectable change in tendons varied from 13 to 57% depending on the plane of scanning and the tendon being examined. Good reproducibility of quantitative tendon measurements can be achieved within a study using two observers by following a defined scanning protocol. However, it is recommended that the same observer perform serial assessments. The data allow minimal detectable changes in tendon size to be calculated.

Adult↗

Evaluation of the limits of visual detection of image misregistration in a brain fluorine-18 fluorodeoxyglucose PET-MRI study.

In routine clinical work, registration accuracy is assessed by visual inspection. However, the accuracy of visual assessment of registration has not been evaluated. This study establishes the limits of visual detection of misregistration in a registered brain fluorine-18 fluorodeoxyglucose positron emission tomography to magnetic resonance image volume. The "best" registered image volume was obtained by automatic registration using mutual information optimization. Translational movements by 1 mm, 2 mm, 3 mm and 4 mm, and rotational movements by 1 degrees , 2 degrees , 3 degrees and 4 degrees in the positive and negative directions in the x- (lateral), y- (anterior-posterior) and z- (axial) axes were introduced to this standard. These 48 images plus six "best" registered images were presented in random sequence to five observers for visual categorization of registration accuracy. No observer detected a definite misregistration in the "best" registered image. Evaluation for inter-observer variation using observer pairings showed a high percentage of agreement in assigned categories for both translational and rotational misregistrations. Assessment of the limits of detection of misregistration showed that a 2-mm translational misregistration was detectable by all observers in the x- and y-axes and 3-mm translational misregistration in the z-axis. With rotational misregistrations, rotation around the z-axis was detectable by all at 2 degrees rotation whereas rotation around the y-axis was detected at 3-4 degrees . Rotation around the x-axis was not symmetric with a positive rotation being identified at 2 degrees whereas negative rotation was detected by all only at 4 degrees. Therefore, visual analysis appears to be a sensitive and practical means to assess image misregistration accuracy. The awareness of the limits of visual detection of misregistration will lead to increase care when evaluating registration quality in both research and clinical settings.

Brain↗

The influence of observer calibration in temporomandibular joint magnetic resonance imaging diagnosis.

OBJECTIVE: This study was undertaken to investigate the potential of reducing observer variation through a calibration program. STUDY DESIGN: The study was based on three sets of randomly selected temporomandibular joint magnetic resonance images. Each set consisted of bilateral images from 20 consecutive patients with temporomandibular disorders. As a baseline, three well-experienced noncalibrated investigators interpreted the images individually for disk position and disk configuration. After the initial interpretation, interobserver agreement was calculated as a kappa index and presented to the examiners. On the same occasion, the investigators analyzed agreement between them on the criteria to be used. RESULTS: Overall data in this study showed an increase in the frequency of interobserver agreement with regard to disk position after the calibration trials were instituted. With regard to disk configuration, substantial interindividual variations were observed even after the observers reached consensus as to the criteria to be used. CONCLUSIONS: These data suggest that after calibration trials, it is possible for three examiners to obtain reliable and reproducible results in reporting temporomandibular joint disk position on magnetic resonance images.

Adult↗

Classification of trochanteric fracture of the proximal femur: a study of the reliability of current systems.

Five observers using the Jensen modification of the Evans classification and the AO classification (with and without subgroups) classified the radiographs of 88 trochanteric hip fractures. Each observer classified the radiographs independently on two occasions 3 months apart. Kappa statistical analysis was used for determination of intra- and inter-observer variation. For the Jensen classification, the mean kappa value was 0.52 (range: 0.44-0.60) for intra-observer variation and 0.34 (range: 0.17-0.38) for inter-observer variation. For the AO system with subgroups, the mean kappa value was 0.42 (range: 0.20-0.65) for intra-observer variation and 0.33 (range: 0.14-0.48) for inter-observer variation. For the AO classification system without subgroups, the mean kappa value was 0.71 (range: 0.60-0.81) for intra-observer variation and 0.62 (range: 0.50-0.71) for inter-observer variation. We recommend classifying trochanteric fractures into three groups as that of the AO system without the subgroups. For ease of use, these three groups may be termed stable trochanteric, unstable trochanteric and trans-trochanteric. Neither the Jensen classification nor the AO classification with subgroups is an acceptable classification system for trochanteric hip fractures.

Femur Head↗

Assessment of analysis-of-variance-based methods to quantify the random variations of observers in medical imaging measurements: guidelines to the investigator.

The random variations of observers in medical imaging measurements negatively affect the outcome of cancer treatment, and should be taken into account during treatment by the application of safety margins that are derived from estimates of the random variations. Analysis-of-variance- (ANOVA-) based methods are the most preferable techniques to assess the true individual random variations of observers, but the number of observers and the number of cases must be taken into account to achieve meaningful results. Our aim in this study is twofold. First, to evaluate three representative ANOVA-based methods for typical numbers of observers and typical numbers of cases. Second, to establish guidelines to the investigator to determine which method, how many observers, and which number of cases are required to obtain the a priori chosen performance. The ANOVA-based methods evaluated in this study are an established technique (pairwise differences method: PWD), a new approach providing additional statistics (residuals method: RES), and a generic technique that uses restricted maximum likelihood (REML) estimation. Monte Carlo simulations were performed to assess the performance of the ANOVA-based methods, which is expressed by their accuracy (closeness of the estimates to the truth), their precision (standard error of the estimates), and the reliability of their statistical test for the significance of a difference in the random variation of an observer between two groups of cases. The highest accuracy is achieved using REML estimation, but for datasets of at least 50 cases or arrangements with 6 or more observers, the differences between the methods are negligible, with deviations from the truth well below +/-3%. For datasets up to 100 cases, it is most beneficial to increase the number of cases to improve the precision of the estimated random variations, whereas for datasets over 100 cases, an improvement in precision is most efficiently achieved by increasing the number of observers. For datasets of at least 50 cases, the standard error ranges between 30% or less with 3 observers down to 10% or less with 8 observers, and the differences in precision between the methods are negligible. The F test (PWD) is very anticonservative and should not be used, while the t test (RES) is reliable for datasets of at least 2 x 50 cases evaluated by 4 or more observers. The likelihood-ratio-test (REML estimation) consistently indicates the significance of a difference in the random variation of an observer between two groups of cases, regardless of the number of cases, and regardless of the number of observers. If a statistical package to perform REML estimation is available, and the investigator feels confident using it, this is the preferred method for studies that involve less than 50 cases evaluated by less than 6 observers. Otherwise, the RES method is an excellent alternative, because of its straightforward implementation, its completeness with respect to the provided statistics, and its overall sufficient accuracy, precision, and reliability of the provided statistical test. If neither the RES method nor REML estimation can provide sufficient performance, either more observers or more cases must be included.

Algorithms↗

The effect of SNP marker density on the efficacy of haplotype tagging SNPs--a warning.

We investigate here the efficacy of selecting haplotype tagging SNPs at different marker densities (2kb-10kb). Our results are based on publicly available data on 5324 markers with a median spacing of 1kb from chromosome 20. We find that whatever density of SNPs is used, htSNP analysis indicates in most cases that at least 80% of the variation can be captured using a subset of SNPs. However, as marker density decreases these htSNPs become increasingly unreliable. In this dataset htSNPs were selected to capture at least 80% of the variation at every observed SNP. At an observed SNP density of 2kb, htSNP analysis suggests that the htSNPs capture on average 95% of the observed variation, when in fact they capture 88% of the unobserved variation. At a density of 10kb, htSNP analysis suggests that 93% of the observed variation was captured, when in fact they capture on average only 78%. Our results indicate that htSNP analysis is only reliable when markers are dense--a spacing of even 2kb shows a considerable loss of information. Such findings are important both for individual studies utilising htSNPs to reduce costs, and for projects such as HapMap which try to characterise human genomic variation using htSNPs.

Chromosomes, Human, Pair 20↗

Normal fetal growth evaluated by longitudinal ultrasound examinations.

Fetal weight estimation was evaluated using the equations of Warsof, Shepard and Hadlock in 192 patients, less than 3 days before delivery. Warsof's and Hadlock's equations resulted in significantly better weight estimates compared to Shepard's equation. No systematic error was found below 2500 g by use of Warsof's equation, whereas Shepard's and Hadlocks's equations resulted in significant over-estimation in the low weight group. In a study of 5 fetuses, of 27-38 weeks gestational age, the intra-observer variation was calculated to 4.6%, whereas the coefficient of variation among observer means was 2.9%. The mixed intra- and inter-observer coefficient of variation was 6.5%. Thirty-five low-risk, uncomplicated pregnancies with reliable last menstrual dates were investigated longitudinally with ultrasound measurements of fetal weight. Population growth curves of fetal weight, fetal femur length, abdominal circumference and biparietal diameter were constructed by weighted polynomial regression. After 27 weeks of gestational age the weight growth curve showed only insignificant non-linearity. Compared to a Danish growth curve based on birth weights, significant higher mean weight was found, especially before 31 weeks of gestational age. The 10th and 90th percentiles for the individual percentage deviation change was +/- 4.4% per 28 days.

Body Weight↗

Quantitative studies of the gastrin-producing cells of the human antrum. A methodological study.

The antral gastrin-producing cells (G-cells) have been identified by the indirect immunoperoxidase technique in two antrum preparations removed due to a recurrent duodenal and gastric ulcer. Morphometric principles were applied to the G-cells with determination of their volume density, numerical density, and mean cell volume. The study showed that within-observer variation, between-observer variation and within-patient variation were negligible, provided at least 200 G-cells were counted. A biopsy material can be used, as well as larger tissue blocks, when this minimum sample size is respected. A method for estimating the total G-cell population and the total G-cell volume in the antrum was developed. In the antrum removed due to a gastric ulcer the number of G-cells was 190 x 10(6) and their total volume 176 mm3.

Cell Count↗

Trachoma: evaluation of a new grading scheme in the United Republic of Tanzania.

A new simplified grading system for trachoma, which is based on the presence or absence of five selected key signs, has been assessed. The level of inter-observer variation and of variation for individual observers (intra-observer variation) showed that the system had good reproducibility following a training period that included interactive clinical teaching. The grading scheme was quickly learned by experienced ophthalmologists and auxiliary health personnel (ophthalmic nurses). The scheme should therefore be suitable for widespread application in field surveys of trachoma.

Allied Health Personnel↗

MRI measurement of brain tumor response: comparison of visual metric and automatic segmentation.

An automatic magnetic resonance imaging (MRI) multispectral segmentation method and a visual metric are compared for their effectiveness to measure tumor response to therapy. Automatic response measurements are important for multicenter clinical trials. A visual metric such as the product of the largest diameter and the largest perpendicular diameter of the tumor is a standard approach, and is currently used in the Radiation Treatment Oncology Group (RTOG) and the Eastern Cooperative Oncology Group (EGOG) clinical trials. In the standard approach, the tumor response is based on the percentage change in the visual metric and is categorized into cure, partial response, stable disease, or progression. Both visual and automatic methods are applied to six brain tumor cases (gliomas) of varying levels of segmentation difficulty. The analyzed data were serial multispectral MR images, collected using MR contrast enhancement. A fully automatic knowledge guided method (KG) was applied to the MRI multispectral data, while the visual metric was taken from the MRI films using the T1 gadolinium enhanced image, with repeat measurements done by two radiologists and two residents. Tumor measurements from both visual and automatic methods are compared to "ground truth," (GT) i.e., manually segmented tumor. The KG method was found to slightly overestimate tumor volume, but in a consistent manner, and the estimated tumor response compared very well to hand-drawn ground truth with a correlation coefficient of 0.96. In contrast, the visually estimated metric had a large variation between observers, particularly for difficult cases, where the tumor margins are not well delineated. The inter-observer variation for the measurement of the visual metric was only 16%, i.e., observers generally agreed on the lengths of the diameters. However, in 30% of the studied cases no consensus was found for the categorical tumor response measurement, indicating that the categories are very sensitive to variations in the diameter measurements. Moreover, the method failed to correctly identify the response in half of the cases. The data demonstrate that automatic 3D methods are clearly necessary for objective and clinically meaningful assessment of tumor volume in single or multicenter clinical trials.

Adult↗

Measurements of tooth length in panoramic radiographs. 2: Observer performance.

Observer performance in tooth-length measurements in panoramic radiography has been examined. Sixty-four teeth, evenly distributed between maxillary first molars, second premolars and mandibular first and second premolars, were fixed in plastic moulds. Each cast was radiographed with an Orthopantomograph, twice with steel balls indicating the cusps and apices of the teeth and once without them. One observer measured the radiographic tooth length twice in both radiographs with indicators in order to estimate the true radiographic tooth length and the repositioning error, and twice in the radiographs without indicators. Seven other observers measured the tooth length on the radiographs without indicators, twice applying their own criteria and twice using defined criteria. The accuracy of the tooth-length measurements was calculated by comparing the true radiographic tooth length with the measurements of the seven observers. The intra-observer variation was defined as the difference between two measurements on the same radiograph. The accuracy and the precision of the tooth-length measurements were highly influenced by the observer performance. The mean radiographic tooth length of the seven observers was closer to the true radiographic tooth length when the observers applied the defined criteria. The inter- and intra-observer variation was higher for the measurements of the maxillary teeth (0.3-1.9 mm) compared with the mandibular teeth (0.3-0.9 mm) for most observers. The highest intra-observer variation was found for the palatal root of the maxillary first molar. The intra-observer variation of four observers decreased somewhat when the defined criteria were used; for the other three observers, however, the variation increased.(ABSTRACT TRUNCATED AT 250 WORDS)

Humans↗

Intraobserver and interobserver variations in sonographic measurements of kidney size in adult volunteers. A comparison of linear measurements and volumetric estimates.

Estimation of renal size by sonography can be performed by measuring renal length, volume, cortical volume or cortical thickness. Observer variation in these measurements is an important factor, especially when repeated measurements are compared. This study was performed to examine the magnitude of intraobserver and interobserver variations for each of the above-mentioned measurements, and to find the measurement with the lowest observer variation. Sonographic measurements were performed by 3 observers on 18 adult volunteers. The standard deviation of the difference (SDD) between any 2 pairs of measurements was used as the indicator of the magnitude of the observer variation. Renal length measurement showed the lowest observer variation with a relative SDD of 4 to 5%. Measurement of cortical thickness showed the poorest reproducibility with a relative SDD of 18 to 23%, while volumetric estimations had a relative SDD of 14 to 17%. Renal length measurement should be preferred to renal volume estimation, especially when comparing repeated measurements.

Adult↗