PubMed Health⌕ Search

Biomedical subjects

Andriy I Bandos

Publications and source records attributed to Andriy I Bandos.

7 recordsLinked to original sources

The prevalence effect in a laboratory environment: Changing the confidence ratings.

RATIONALE AND OBJECTIVES: We sought to assess whether or not prevalence levels affected the confidence ratings of readers during the interpretation of cases in a laboratory receiver operating characteristic-type observer performance study. MATERIALS AND METHODS: We reanalyzed a previously conducted observer performance study that included 14 readers and 5 different levels of prevalence. The previous study yielded the observation that in the laboratory we could not detect a "prevalence effect" in terms of differences in areas under the receiver operating characteristic curves. The detection ratings (for presence or absence) of lung nodules, interstitial disease, and pneumothorax for the five prevalence levels were compared, and a test for trend in averaged ratings as a function of abnormality prevalence was performed within a mixed-model setting that accounts for different sources of variability and correlations induced by the study design. RESULTS: The ratings of the cases in terms of confidence that the specific abnormality in question is present tend, on average, to be larger when actual disease prevalence is lower. The rate of the increase of the average confidence ratings with the decreasing prevalence of a specific abnormality is very similar for actually positive and actually negative cases for every considered abnormality. The observed trend in the changes of the average confidence ratings as a function of prevalence levels was statistically significant (p < 0.01). CONCLUSION: Expectations of disease prevalence in the case mix during a laboratory observer performance study may systematically affect the behavior of observers in terms of their actual confidence ratings.

Humans↗

A permutation test for comparing ROC curves in multireader studies a multi-reader ROC, permutation test.

RATIONALE AND OBJECTIVES: The aim of the study is to develop a permutation test to compare receiver operating characteristic (ROC) curves of two diagnostic modalities in a multireader paired design. MATERIALS AND METHODS: A statistical test for comparing two diagnostic modalities is developed based on all possible exchanges of the set of reader-ratings between the two modalities. An exact permutation test is formed by determining the frequency of the most extreme values of the statistic estimating the average difference in the areas under the ROC curves (AUCs). An asymptotic version of the test is constructed by obtaining the exact permutation variance and appealing to the asymptotic normality of the nonparametric estimator of the average difference in areas. Computer simulations were conducted to validate the type I error for small sample sizes. RESULTS: The new test provides a permutation approach for comparing ROC curves in a multireader paired-design setting in which effects of the readers are considered to be fixed. The type I error of the asymptotic test is close to the true value, even for samples as small as 20 normal and 20 abnormal cases. The test is designed to be sensitive to alternatives in which the AUCs of the two diagnostic modalities differ. CONCLUSIONS: The proposed test provides a powerful method for comparing two diagnostic modalities in a multireader paired-study design when the primary interest is to detect difference in average AUCs.

Analysis of Variance↗

Reader variance in ROC studies--generalizability to reader population at high and low performance levels.

RATIONALE AND OBJECTIVES: To investigate the variability between discriminative performances of readers as a function of average performance levels during receiver operating characteristic (ROC) studies. MATERIALS AND METHODS: Four subsets of cases from previously ascertained ROC rating data by 12 observers when detecting interstitial disease and pneumothorax on posteroanterior chest films were selected for each abnormality and reanalyzed to assess changes in "reader" variance component. The subsets were selected based on a prestudy subjective assessment of the subtleness of depicted abnormality (positive cases) and the difficulty in determining its absence (negative cases). Reader variance component was estimated using a bootstrap approach for each subset and the results were used to assess a general relationship between variability and average performance level. RESULTS: The reader variance component decreased substantially (from 0.007704 to 0.000426), as expected, when the areas under the ROC curves (AUC) for detecting pneumothoraces increased from 84% to 97%. On the other hand, reader variance component increased substantially (from 0.000890 to 0.005181) when AUC for detecting interstitial disease increased from 59% to 87%. The large magnitude of and changes in the reader variance component resulted in a consistent nonmonotone relationship as a function of AUC when other related variance components were included in addition to the reader component. CONCLUSION: Among several factors affecting generalizability of ROC results to the population of readers, the reader variance component depended nonmonotonically on the average diagnostic performance and is lowest at both very high and very low levels of performance.

Clinical Competence↗

A permutation test sensitive to differences in areas for comparing ROC curves from a paired design.

The area under the receiver operating characteristic (ROC) curve (AUC) is a widely accepted summary index of the overall performance of diagnostic procedures and the difference between AUCs is often used when comparing two diagnostic systems. We developed an exact non-parametric statistical procedure for comparing two ROC curves in paired design settings. The test which is based on all permutations of the subject specific rank ratings is formally a test for equality of ROC curves that is sensitive to the alternatives of AUC difference. The operating characteristics of the proposed test were evaluated using extensive simulations over a wide range of parameters. The proposed procedure can be easily implemented in experimental ROC data sets. For small samples and for underlying parameters that are common in experimental studies in diagnostic imaging the test possesses good operating characteristics and is more powerful than the conventional non-parametric procedure for AUC comparisons. We also derived an asymptotic version of the test which uses an exact estimate of the variance in the permutation space and provides a good approximation even when the sample sizes are small. This asymptotic procedure is a simple and precise approximation to the exact test and is useful for large sample sizes where the exact test may be computationally burdensome.

Biometry↗

A conditional nonparametric test for comparing two areas under the ROC curves from a paired design.

RATIONALE AND OBJECTIVES: To develop a conditional nonparametric procedure for comparing two correlated areas under receiver operating characteristic (ROC) curves (AUC). MATERIALS AND METHODS: A nonparametric conditional test to compare areas under two ROC curves was developed using the distribution of the elements of the nonparametric AUC estimators in a permutation space. The conditioning is made on the observed discordances between the relative orderings of ratings of the normal and abnormal cases for the two modalities taken over all possible pairs. The type I error of the procedure was verified using computer simulations. The power of the test was compared with an existing unconditional procedure on simulated datasets from binormal distributions as well as from a mixture of binormal distributions of ratings. RESULTS: The proposed test is conservative for low sample sizes, large AUC, and high correlation between modalities. It possesses a reasonable type I error for sample sizes as low as 20 actually positive and 20 actually negative cases. In plausible situations in which the sample in observer performance studies can not be monotonically transformed into a binormal distribution, this approach may have modest power advantages over the conventional nonparametric test. CONCLUSION: The conditional nonparametric test presented here is an alternative approach to existing unconditional procedures and may offer advantages in certain types of observer performance studies.

Area Under Curve↗

Incorporating utility-weights when comparing two diagnostic systems: a preliminary assessment.

RATIONALE AND OBJECTIVES: We sought to develop a new index that incorporates utility-weights when assessing the overall performance of a diagnostic system and to provide a statistical test for comparing two indices in a paired study design. MATERIALS AND METHODS: The area under the receiver operating characteristic (ROC) curve (AUC) was used as the basis for constructing a new index. The index we propose represents a weighted average of class-specific AUCs each of which relates to a class of pairs of actually negative (normal) and actually positive (abnormal) cases with a specific predetermined utility (or clinical importance). For each pair of normal-abnormal cases, the utility is defined a priori and based on external (covariate) information. In the proposed approach utility-weights represent the relative importance (utility) of discriminating between different types of normal and abnormal cases (pairs of the same type are combined in the classes termed utility-classes). We also describe a simple nonparametric procedure for comparing the proposed indices as computed from paired data. Computer simulations were conducted to evaluate the behavior of the type I error of the proposed test in the simple albeit important instance of two utility-classes. RESULTS: The new index provides an extension of the commonly used area under the ROC curve. It allows for incorporation of utility-weights into the analysis and reduces to the conventional AUC index when all assigned utility-weights are equal to unity. Computer simulations indicate that in the considered scenario of two utility-classes, the type I error of the proposed test is comparable to that of the conventional nonparametric test for equality of AUC indices. CONCLUSIONS: The proposed index and the statistical test provide a practical approach of incorporating utilities when comparing diagnostic systems.

Algorithms↗

Variability in observer performance studies experimental observations.

RATIONALE AND OBJECTIVES: The aim of the study is to assess variance components in observer performance studies and the possible impact on study results and conclusions. MATERIALS AND METHODS: Two previously performed retrospective receiver operating characteristic-type observer performance studies to evaluate the performance of seven radiologists in detecting interstitial disease on conventional posteroanterior chest films and nine radiologists in detecting interstitial disease on a high-resolution workstation were reanalyzed by using the Beiden, Wagner, and Campbell nine-component model to estimate the different variance components. We estimated case-, reader-, and mode-related components of the variance for the group as a whole and after excluding (round robin) each reader. Overall variance was evaluated, and the effect of individual readers on overall study conclusions was assessed. RESULTS: Overall results and conclusions of the reanalysis agreed with the original one in that, as a group, radiologists performed significantly better when using conventional films (P < .05) in both studies. Reader variability was large compared with all other components, and in one study, it was substantially larger for the workstation reading mode. Reader variability was affected substantially by one observer in each study, and in one study, reader-by-mode variability was affected by another reader who performed better on the workstation. CONCLUSION: Estimates of variance components can shed light on the appropriateness of study design, as well as the sensitivity of results to the inclusion (or exclusion) of individual observers.

Humans↗