Toward a strategy for consensus development on a quantitative approach to medical imaging.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Robert F Wagner.
Explore the source record for details and available documents.
This paper considers binary classification. We assess a classifier in terms of the Area Under the ROC Curve (AUC). We estimate three important parameters, the conditional AUC (conditional on a particular training set) and the mean and variance of this AUC. We derive, as well, a closed form expression of the variance of the estimator of the AUC. This expression exhibits several components of variance that facilitate an understanding for the sources of uncertainty of that estimate. In addition, we estimate this variance, i.e., the variance of the conditional AUC estimator. Our approach is nonparametric and based on general methods from U-statistics; it addresses the case where the data distribution is neither known nor modeled and where there are only two available data sets, the training and testing sets. Finally, we illustrate some simulation results for these estimators.
PURPOSE: This study evaluated the effects of combined use of two nonionic surfactants on the characteristics (i.e., appearance, emulsification time, and particle size) of oil-in-water microemulsions generated from flurbiprofen-loaded preconcentrates. METHODS: Three phase diagrams were constructed using Capmul PG8 (propylene glycol monocaprylate) as the oil, Tween 20 (polysorbate 20) and/or Cremophor EL (polyoxyl 35 castor oil) as surfactants. A number of preconcentrates were selected based on phase diagrams: O20T80 (20% Capmul PG8, 80% Tween 20), O20C80 (20% Capmul PG8, 80% Cremophor EL), O20T40C40 (20% Capmul PG8, 40% Tween 20, 40% Cremophor EL). Flurbiprofen loading in preconcentrates was tested at 0%, 1%, 2.5%, and 5% (w/w). The resulting mixtures of these preconcentrates upon dilution 100-fold with aqueous medium were characterized by visual and microscopic observation, HPLC and photon correlation spectroscopy. RESULTS: (a) For preconcentrates using single surfactant, either O20T80 or O20C80, the dilution generated emulsions with visible cloudiness. The particle size increased as the drug loading increased; (b) for preconcentrates using two surfactants O20T40C40, the dilution generated clear microemulsions with small particle sizes (10-11nm), and the increased drug loading seemed to have little effect on the particle size. The microemulsions from preconcentrate O20T40C40 was also found to be stable at ambient temperature over 20 days without significant change in particle size at different drug loadings. Dilution with different aqueous medium, either water, or simulated gastric fluid or simulated intestinal fluid, also did not change the particle sizes of the microemulsions. CONCLUSIONS: The combined use of surfactants in preconcentrate showed the promise in generating desired self-emulsifying microemulsions with small particle size, increased drug loading, and improved physical stability. This will have significant implications in future dosage development for poorly water-soluble drugs in using self-emulsifying microemulsions drug delivery system (SMEDDS).
The author traces some critical moments in the history of Image Science in the last half century from first-hand or once-removed experience. The Image Science used in the field of medical imaging today had its origins in the analysis of photon detection developed for modern television, conventional photography, and the human visual system. Almost all "model observers" used in image assessment today converge to the model originally used by Albert Rose in his analysis of those classic photo-detectors. A more general statistical analysis of the various "defects" of conventional and unconventional photon-imaging technologies was provided by Shaw. A number of investigators in medical imaging elaborated the work of these pioneers into a synthesis with the general theory of signal detectability and extended this work to the various forms of CT, energy-spectral-dependent imaging, and the further complication of anatomical-background-noise limited imaging. The author calls for further extensions of this work to the problem of under-sampled and thus artefact-limited imaging that will be important issues for high-speed CT and MRI.
RATIONALE AND OBJECTIVES: Several statistical methods have been developed for analyzing multireader, multicase (MRMC) receiver operating characteristic (ROC) studies. The objective of this article is to increase awareness of these methods and determine if their results are concordant for published datasets. MATERIALS AND METHODS: Data from three previously published studies were reanalyzed using five MRMC methods. For each method the 95% confidence intervals (CIs) for the mean of the readers' ROC areas for each diagnostic test, the P value for the comparison of the diagnostic tests' mean accuracies, and the 95% CIs for the mean difference in ROC areas of the diagnostic tests were reported. RESULTS: Important differences in P values and CIs were seen when using parametric versus nonparametric estimates of accuracy, and there were the expected differences for random-reader versus fixed-reader models. Controlling for these differences, the Dorfman-Berbaum-Metz (DBM), Obuchowski-Rockette, Beiden-Wagner-Campbell, and Song's multivariate Wilcoxon-Mann-Whitney (WMW) methods gave almost identical results for the fixed-reader model. For the random-reader model, the DBM, Obuchowski-Rockette, and Beiden-Wagner-Campbell methods yielded approximately the same inferences, but the CIs for the Beiden-Wagner-Campbell method tend to be broader. Ishwaran's hierarchical ROC sometimes yielded significance not found with other methods. Song's modification of DBM's jack-knifing algorithm sometimes led to different conclusions than the original DBM algorithm. CONCLUSION: In choosing and applying MRMC methods, it is important to recognize: (1) the distinction between random-reader and fixed-reader models, the uncertainties accounted for by each, and thus the level of generalizeability expected from each; (2) assumptions made by the various MRMC methods; and (3) limitations of a five- or six-reader study when the reader variability is great.
Explore the source record for details and available documents.
Cancer of the lung and bronchus is the leading fatal malignancy in the United States. Five-year survival is low, but treatment of early stage disease considerably improves chances of survival. Advances in multidetector-row computed tomography technology provide detection of smaller lung nodules and offer a potentially effective screening tool. The large number of images per exam, however, requires considerable radiologist time for interpretation and is an impediment to clinical throughput. Thus, computer-aided diagnosis (CAD) methods are needed to assist radiologists with their decision making. To promote the development of CAD methods, the National Cancer Institute formed the Lung Image Database Consortium (LIDC). The LIDC is charged with developing the consensus and standards necessary to create an image database of multidetector-row computed tomography lung images as a resource for CAD researchers. To develop such a prospective database, its potential uses must be anticipated. The ultimate applications will influence the information that must be included along with the images, the relevant measures of algorithm performance, and the number of required images. In this article we outline assessment methodologies and statistical issues as they relate to several potential uses of the LIDC database. We review methods for performance assessment and discuss issues of defining "truth" as well as the complications that arise when truth information is not available. We also discuss issues about sizing and populating a database.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
RATIONALE AND OBJECTIVES: The authors analyzed two methods for arranging the temporal sequencing of unaided versus computer-assisted reading in multiple-reader, multiple-case receiver operating characteristic studies of detection of solitary pulmonary nodules on chest radiographs. In the "independent" mode the readings are separated by about I month; in the "sequential" mode, the computer-assisted reading immediately follows the unassisted reading. MATERIALS AND METHODS: The authors used the method of Beiden, Wagner, and Campbell (BWC) to decompose the components of variance of receiver operating characteristic accuracy measures into those that are correlated and those that are uncorrelated across reading conditions. Only the latter contribute to uncertainty in estimates of the difference in accuracy measures across reading conditions (unaided vs aided). This method was used to analyze data from two independent studies of the detection of solitary pulmonary nodules on chest radiographs. RESULTS: In the sequential reading mode the components that were correlated across reading conditions increased compared to the independent reading mode, as might be expected. What was not anticipated was the fact that the total reader variance was approximately the same for the two reading modes. The results were remarkably similar across the two independent studies analyzed. CONCLUSION: The sequential reading mode may thus be the more sensitive probe of the difference between unassisted and computer-assisted reading, if the mean effect is unperturbed (as here). It is also the least demanding on the logistics and investment of reader time.
In the last 2 decades major advances have been made in the field of assessment methods for medical imaging and computer-assist systems through the use of the paradigm of the receiver operating characteristic (ROC) curve. In the most recent decade this methodology was extended to embrace the complication of reader variability through advances in the multiple-reader, multiple-case (MRMC) ROC measurement and analysis paradigm. Although this approach has been widely adopted by the imaging research community, some investigators appear averse to it, possibly from concern that it could place a greater burden on the scarce resources of patient cases and readers compared to the requirements of alternative methods. The present communication argues, however, that the MRMC ROC approach to assessment in the context of reader variability may be the most resource-efficient approach available. Moreover, alternative approaches may also be statistically uninterpretable with regard to estimated summary measures of performance and their uncertainties. The authors propose that the MRMC ROC approach be considered even more widely by the larger community with responsibilities for the introduction and dissemination of medical imaging technologies to society. General principles of study design are reviewed, and important contemporary clinical trials are used as examples.
The multiple-reader, multiple-case (MRMC) approach to receiver operating characteristic (ROC) analysis is becoming the dominant assessment paradigm in medical imaging. Its most common version involves having many readers read every patient case in the study, a critical feature since differences among competing imaging modalities are often dominated by differences in reader performance. The present authors have carried out MRMC ROC analysis on a uniquely large data set for mammography. The analysis quantifies the great range of observed reader skill in that data set. It also demonstrates that the sample sizes are sufficiently large that the conclusions generalize to the populations sampled here with little uncertainty from the finite sample size. A schematic approach to bracketing the utility matrix is then used to study trends in the resulting expected utility functions that correspond to the range of observed ROC curves. This is done for both the screening and the diagnostic context. The results raise 2 hypotheses for further investigation. First, it is possible that the present ambiguity surrounding the effectiveness of mammography is due in part to the observed range of reader skills and corresponding expected utility functions. Second, it is possible that computer-assisted modalities for mammography may lead to improvements in the expected utility function not only for screening but also in the diagnostic context, especially for the lower performing readers.