PubMed Health⌕ Search

Biomedical subjects

Edward A Sickles

Publications and source records attributed to Edward A Sickles.

At least 19 recordsLinked to original sources

Correlation of radiologist rank as a measure of skill in screening and diagnostic interpretation of mammograms.

PURPOSE: To determine whether skill in the interpretation of screening mammograms is correlated with skill in the interpretation of diagnostic mammograms. MATERIALS AND METHODS: The institutional review board of the University of South Florida approved this study. This study was determined to be exempt from informed consent requirements because of the retrospective use of images and was conducted before HIPPA requirements were implemented. A total of 59 radiologists interpreted screening and diagnostic performance test sets of mammograms with a 1-year interval. Interpretations were recorded with modifications of the Breast Imaging and Reporting Data System. Radiologist skill was measured as the radiologist's ranking among his or her cohort in each of several measures of performance (ie, performance test receiver operating characteristic curve area, performance test screening sensitivity, performance test diagnostic sensitivity, and associated specificities). Correlations between radiologist rank in screening and rank in the diagnostic performance test measures were analyzed with the Spearman rank correlation statistical test. RESULTS: Radiologist rank in screening interpretations and in diagnostic interpretations was found to be significantly correlated in all measurements (P < .05). However, only two measurments (ie, receiver operating characteristic curve area rank correlation of 0.327 and sensitivity rank correlation of 0.402) remained significant after adjusting for multiple testing. The correlation between ranked screening specificity and ranked diagnostic specificity (0.296) was significant at only the .05 level. CONCLUSION: The interpretive performance of radiologists among their peers is moderately correlated between screening and diagnostic interpretations. Thus, proficiency in one area does not guarantee proficiency in the other area for some radiologists.

Clinical Competence↗

Improving the concordance of mammography assessment and management recommendations.

PURPOSE: To retrospectively compare the concordance of initial and final assessment categories for mammograms with management recommendations made before and after the final rules of the Mammography Quality Standards Act (MQSA) were in effect for screening and diagnostic mammography. MATERIALS AND METHODS: The study included mammograms from 1996 to 2001 from the seven mammography registries of the Breast Cancer Surveillance Consortium (BCSC). The authors defined the pre-MQSA period as January 1, 1996-April 27, 1999, and the post-MQSA period as April 28, 1999-December 31, 2001 (2470151 screening and 194199 diagnostic mammograms). Assessment was cross-classified according to management recommendation. Changes in concordance between assessment and recommendation were evaluated by year and by period (before and after MQSA) for computer-linked data and for all data by using Pearson chi(2) test to evaluate differences. Mantel-Haenszel chi(2) test was used to measure change in concordance over time. Each registry and the BCSC Statistical Coordinating Center had a Federal Certificate of Confidentiality and approval from each institution's review board for protection of human subjects to collect and send data to coordinating center and conduct research with these data. Active consent was required at only one site in this HIPAA-compliant study. RESULTS: Concordance increased significantly in the post-MQSA period for Breast Imaging Reporting and Data System categories 3-5 assessments at both screening and diagnostic mammography. The most substantial improvements were in the use of the management recommendation for "additional imaging," which decreased from 41% in 1996 to 15% in 2001 for screening mammograms with an initial assessment of category 4 (P < .001). Recommendation for short-interval follow-up in women with screening mammograms with a category 3 final assessment increased from 51% in 1996 to 76% in 2001 (P < .001). Concordance for diagnostic mammograms assigned category 0 improved from 65% in the pre-MQSA period to 81% in the post-MQSA period (P < .001). CONCLUSION: This analysis demonstrates that over a relatively short period of time, major improvement in radiology reporting has occurred.

Breast Neoplasms↗

Performance benchmarks for screening mammography.

PURPOSE: To retrospectively evaluate the range of performance outcomes of the radiologist in an audit of screening mammography by using a representative sample of U.S. radiologists to allow development of performance benchmarks for screening mammography. MATERIALS AND METHODS: Institutional review board approval was obtained, and study was HIPAA compliant. Informed consent was or was not obtained according to institutional review board guidelines. Data from 188 mammographic facilities and 807 radiologists obtained between 1996 and 2002 were analyzed from six registries from the Breast Cancer Surveillance Consortium (BCSC). Contributed data included demographic information, clinical findings, mammographic interpretation, and biopsy results. Measurements calculated were positive predictive values (PPVs) from screening mammography (PPV(1)), biopsy recommendation (PPV(2)), biopsy performed (PPV(3)), recall rate, cancer detection rate, mean cancer size, and cancer stage. Radiologist performance data are presented as 50th (median), 10th, 25th, 75th, and 90th percentiles and as graphic presentations by using smoothed curves. RESULTS: There were 2 580 151 screening mammographic studies from 1 117 390 women (age range, <30 to >/=80 years). The respective means and ranges of performance outcomes for the middle 50% of radiologists were as follows: recall rate, 9.8% and 6.4%-13.3%; PPV(1), 4.8% and 3.4%-6.2%; and PPV(2), 24.6% and 18.8%-32.0%. Mean cancer detection rate was 4.7 per 1000, and the median [corrected] mean size of invasive cancers was 13 mm. The range of performance outcomes for the middle 80% of radiologists also was presented. CONCLUSION: Community screening mammographic performance measurements of cancer outcomes for the majority of radiologists in the BCSC surpass performance recommendations. Recall rate for almost half of radiologists, however, is higher than the recommended rate.

Adult↗

Reality check: perceived versus actual performance of community mammographers.

OBJECTIVE: Federal regulations mandate that radiologists receive regular albeit limited feedback regarding their interpretive accuracy in mammography. We sought to determine whether radiologists who regularly receive more extensive feedback can report their actual performance in screening mammography accurately. SUBJECTS AND METHODS: Radiologists (n = 105) who routinely interpret screening mammograms in three states (Washington, Colorado, and New Hampshire) completed a mailed survey in 2001. Radiologists were asked to estimate how frequently they recommended additional diagnostic testing after screening mammography and the positive predictive value of their recommendations for biopsy (PPV2). We then used outcomes from 336,128 screening mammography examinations interpreted by the radiologists from 1998 to 2001 to ascertain their true rates of recommendations for diagnostic testing and PPV2. RESULTS: Radiologists' self-reported rate of recommending immediate additional imaging (11.1%) exceeded their actual rate (9.1%) (mean difference, 1.9%; 95% confidence interval [CI], 0.9-3.0%). The mean self-reported rate of recommending short-interval follow-up was 6.2%; the true rate was 1.8% (mean difference, 4.3%; 95% CI, 3.6-5.1%). Similarly, the mean self-reported and true rates of recommending immediate biopsy or surgical evaluation were 3.2% and 0.6%, respectively (mean difference, 2.6%; 95% CI, 1.8-3.4%). Conversely, radiologists' mean self-reported PPV2 (18.3%) was significantly less than their mean true PPV2 (27.6%) (mean difference, -9.3%; 95% CI, -12.4% to -6.2%). CONCLUSION: Despite regular performance feedback, community radiologists may overestimate their true rates of recommending further evaluation after screening mammography and underestimate their true positive predictive value.

Biopsy↗

Physician predictors of mammographic accuracy.

BACKGROUND: The association between physician experience and the accuracy of screening mammography in community practice is not well studied. We identified characteristics of U.S. physicians associated with the accuracy of screening mammography. METHODS: Data were obtained from the Breast Cancer Surveillance Consortium and the American Medical Association Master File. Unadjusted mammography sensitivity and specificity were calculated according to physician characteristics. We modeled mammography sensitivity and specificity by multivariable logistic regression as a function of patient and physician characteristics. All statistical tests were two-sided. RESULTS: We studied 209 physicians who interpreted 1,220,046 screening mammograms from January 1, 1995, through December 31, 2000, of which 7143 (5.9 per 1000 mammograms) were associated with breast cancer within 12 months of screening. Each physician interpreted a mean of 6011 screening mammograms (95% confidence interval [CI] = 4998 to 6677), including a mean of 34 (95% CI = 28 to 40) from women diagnosed with breast cancer. The mean sensitivity was 77% (range = 29%-97%), and the mean false-positive rate was 10% (range = 1%-29%). After adjustment for the patient characteristics of those whose mammograms they interpreted, physician characteristics were strongly associated with specificity. Higher specificity was associated with at least 25 years (versus less than 10 years) since receipt of a medical degree (for physicians practicing for 25-29 years, odds ratio [OR] = 1.54, 95% CI = 1.14 to 2.08; P = .006), interpretation of 2500-4000 (versus 481-750) screening mammograms annually (OR = 1.30, 95% CI = 1.06 to 1.59; P = .011) and a high focus on screening mammography compared with diagnostic mammography (OR = 1.59, 95% CI = 1.37 to 1.82; P<.001). Higher overall accuracy was associated with more experience and with a higher focus on screening mammography. Compared with physicians who interpret 481-750 mammograms annually and had a low screening focus, physicians who interpret 2500-4000 mammograms annually and had a high screening focus had approximately 50% fewer false-positive examinations and detected a few less cancers. CONCLUSION: Raising the annual volume requirements in the Mammography Quality Standards Act might improve the overall quality of screening mammography in the United States.

Adult↗

Breast cancer yield for screening mammographic examinations with recommendation for short-interval follow-up.

PURPOSE: To compare cancer yield for screening examinations with recommendation for short-interval follow-up after diagnostic imaging work-up versus after screening mammography only. MATERIALS AND METHODS: From January 1996 to December 1999, Breast Imaging Reporting and Data System assessments and recommendations were collected prospectively for 1,171,792 screening examinations in 758,015 women aged 40-89 years at seven mammography registries in Breast Cancer Surveillance Consortium. Registries obtained waiver of signed consent or collected signed consent in accordance with institutional review boards at each location. Diagnosis of invasive cancer or ductal carcinoma in situ within 24 months of screening examination and tumor stage and size for invasive cancer were determined through linkage to pathology database or tumor registry. chi2 test was used to determine significant differences between groups. RESULTS: Overall, 5.2% of first and 1.7% of subsequent screens included recommendation for short-interval follow-up, which was similar to likelihood of recommendation for diagnostic evaluation (first screens, 4.6%; subsequent, 2.6%). Most recommendations for short-interval follow-up were based on screening mammography alone (86.2% of first screens, 77.5% of subsequent). Yield of cancer for screening examinations with probably benign finding (PBF) and recommendation for short-interval follow-up based on screening mammography alone tended to be lower than in those with PBF and recommendation for short-interval follow-up after additional work-up (first screens: 0.54% vs 0.96%, P=.10; subsequent: 1.50% vs 1.73%, P=.26). Proportion of stage II and higher disease tended to be higher for examinations with PBF and recommendation for short-interval follow-up based on screening mammography alone compared with those recommended for short-interval follow-up after additional work-up (first screens: 34.7% vs 24.4%, P=.43; subsequent: 27.5% vs 19.2%, P=.13). CONCLUSION: Many first screening examinations include recommendation for short-interval follow-up based on screening mammography alone. Cancer yield for these examinations is low and is lower than that with diagnostic work-up prior to short-interval follow-up recommendation. Absence of diagnostic work-up prior to short-interval follow-up recommendation may result in periodic surveillance of a high proportion of benign lesions.

Adult↗

Performance benchmarks for diagnostic mammography.

PURPOSE: To evaluate a range of performance parameters pertinent to the comprehensive auditing of diagnostic mammography examinations, and to derive performance benchmarks therefrom, by pooling data collected from large numbers of patients and radiologists that are likely to be representative of mammography practice in the United States. MATERIALS AND METHODS: Institutional review board approval was met, informed consent was not required, and this study was Health Insurance Portability and Accountability Act compliant. Six mammography registries contributed data to the Breast Cancer Surveillance Consortium (BCSC), providing patient demographic and clinical information, mammogram interpretation data, and biopsy results from defined population-based catchment areas. The study involved 151 mammography facilities and 646 interpreting radiologists. The study population included women 18 years of age or older who underwent at least one diagnostic mammography examination between 1996 and 2001. Collected data were used to derive mean performance parameter values, including abnormal interpretation rate, positive predictive value (for abnormal interpretation, biopsy recommended, and biopsy performed), cancer diagnosis rate, invasive cancer size, and the percentages of minimal cancers, axillary node-negative invasive cancers, and stage 0 and I cancers. Additional benchmarks were derived for these performance parameters, including 10th, 25th, 50th (median), 75th, and 90th percentile values. RESULTS: The study involved 332,926 diagnostic mammography examinations. Mean performance parameter values were abnormal interpretation rate, 8.0%; positive predictive value for abnormal interpretation, 31.4%; positive predictive value for biopsy recommended, 31.5%; positive predictive value for biopsy performed, 39.5%; cancer diagnosis rate, 25.3 per 1000 examinations; invasive cancer size, 20.2 mm; percentage of minimal cancers, 42.0%; percentage of axillary node-negative invasive cancers, 73.6%; and percentage of stage 0 and I cancers, 62.4%. CONCLUSION: The presented BCSC outcomes data and performance benchmarks may be used by mammography facilities and individual radiologists to evaluate their own performance for diagnostic mammography as determined by means of periodic comprehensive audits.

Adult↗

Follow-up of palpable circumscribed noncalcified solid breast masses at mammography and US: can biopsy be averted?

PURPOSE: To determine whether palpable noncalcified solid breast masses with benign morphology at mammography and ultrasonography (US) can be managed similarly to nonpalpable probably benign lesions (Breast Imaging Reporting and Data System [BI-RADS] category 3)-that is, with periodic imaging surveillance-and to determine whether biopsy can be averted in these lesions. MATERIALS AND METHODS: No institutional review board approval or patient consent was required. This retrospective analysis, based on final imaging reports, included 152 patients (age range, 28-77 years; mean age, 48.3 years) with 157 palpable noncalcified solid masses that were classified as probably benign at initial mammography and US. Of 152 patients, 108 underwent follow-up with mammography and US (6-month intervals for 2 years, then 12-month intervals). The remaining 44 patients underwent surgical or needle biopsy after initial imaging. Lesions were analyzed at initial and follow-up examinations. Statistical analysis included Student t test and corresponding exact 95% confidence intervals. RESULTS: In 108 patients who underwent follow-up only, 112 lesions were palpable. In 102 (94.4%) of 108 patients, masses remained stable during follow-up. Lesions were followed for at least 2 years (mean, 4.1 years; range, 2-7 years). In six (5.6%) patients, palpable lesions increased in size during follow-up; these lesions were benign at subsequent open biopsy. No breast carcinoma was diagnosed in the 44 patients with 45 palpable lesions who underwent biopsy after initial imaging. Of 157 lesions, no malignant tumors were observed (exact one-sided 95% confidence interval: 0%, 1.95%). CONCLUSION: The data strongly suggest that palpable noncalcified solid breast masses with benign morphology at mammography and US can be managed similarly to nonpalpable BI-RADS category 3 lesions, with short-term follow-up (6-month intervals for 2 years). More data, based on a larger series, are required to determine whether this conclusion is correct.

Adult↗

Computer-aided detection output on 172 subtle findings on normal mammograms previously obtained in women with breast cancer detected at follow-up screening mammography.

PURPOSE: To evaluate, by using a computer-aided detection (CAD) program, the nonspecific findings on normal screening mammograms obtained in women in whom breast cancer was later detected at follow-up screening mammography. MATERIALS AND METHODS: Four hundred ninety-three mammogram pairs-an initial negative screening mammogram and a subsequently obtained screening mammogram showing cancer-were collected. The mean interval between examinations was 14.6 months. In 169 cases, in which 172 cancers were later depicted, findings on the initial mammogram were subtle enough that either none or only one or two of five blinded radiologists recommended screening recall. On the initial negative mammograms, of the 172 areas where cancer later developed, 137 (80%) had subtle nonspecific findings and were retrospectively judged as having a benign or normal appearance. The mammograms with these subtle findings were evaluated with a commercially available CAD program, and the numbers of CAD marks on these nonspecific findings were analyzed. RESULTS: Of the 172 cancers, 129 (75%) were invasive and 43 (25%) were ductal carcinoma in situ. The CAD program marked 72 (42%) of the 172 findings that subsequently developed into cancer: 24 (29%) of 82 findings recalled by none, 25 (49%) of 51 findings recalled by one, and 23 (59%) of 39 findings recalled by two of the five radiologists. Among the 137 areas with nonspecific normal or benign findings, 41 (30%) areas where cancer subsequently developed were marked by the CAD program. CONCLUSION: A subset of cancers have perceptible but nonspecific mammographic findings that may be marked by a CAD program, even when the findings do not warrant recall as judged at blinded and unblinded radiologist review. The authors believe failure to act on such nonspecific but CAD-marked findings prospectively does not constitute interpretation below a reasonable standard of care.

Adult↗

A probabilistic expert system that provides automated mammographic-histologic correlation: initial experience.

OBJECTIVE: We sought to determine whether a probabilistic expert system can provide accurate automated imaging-histologic correlations to aid radiologists in assessing the concordance of mammographic findings with the results of imaging-guided breast biopsies. MATERIALS AND METHODS: We created a Bayesian network in which Breast Imaging Reporting and Data System (BI-RADS) descriptors are used to convey the level of suspicion of mammographic abnormalities. Our system is a computer model that links BI-RADS descriptors with diseases of the breast using probabilities derived from the literature. Mammographic findings are used to update pretest probabilities (prevalence of disease) into posttest probabilities applying Bayes' theorem. We evaluated the histologic results of 92 consecutive imaging-guided breast biopsies for concordance with the mammographic findings during radiology-pathology review sessions. First, radiologists with no knowledge of the biopsy results chose BI-RADS descriptors for the mammographic findings. After the histologic diagnosis was revealed, the radiologists assessed concordance between the pathologic results and the mammographic findings. We then input the information gathered from these sessions into the Bayesian network to produce an automated mammographic-histologic correlation. RESULTS: We had a sampling error rate of 1.1% (1/92 biopsies). Our expert system was able to integrate pathologic diagnoses and mammographic findings to obtain probabilities of sampling error, thereby enabling us to identify the incorrect pathologic diagnosis with 100% sensitivity while maintaining a specificity of 91%. CONCLUSION: Our probabilistic expert system has the potential to help radiologists in identifying breast biopsy results that are discordant with mammographic findings and discovering cases in which biopsy sampling errors may have occurred.

Adult↗

Comparison of screening mammography in the United States and the United kingdom.

CONTEXT: Screening mammography differs between the United States and the United Kingdom; a direct comparison may suggest methods to improve the practice. OBJECTIVE: To compare screening mammography performance between the United States and the United Kingdom among similar-aged women. DESIGN, SETTING, AND PARTICIPANTS: Women aged 50 years or older were identified who underwent 5.5 million mammograms from January 1, 1996, to December 31, 1999, within 3 large-scale mammography registries or screening programs: the Breast Cancer Surveillance Consortium (BCSC, n = 978 591) and National Breast and Cervical Cancer Early Detection Program (NBCCEDP, n = 613 388) in the United States; and the National Health Service Breast Screening Program (NHSBSP, n = 3.94 million) in the United Kingdom. A total of 27 612 women were diagnosed with breast cancer (invasive or ductal carcinoma in situ) within 12 months of screening among the 3 groups. MAIN OUTCOME MEASURES: Recall rates (recommendation for further evaluation including diagnostic imaging, ultrasound, clinical examination, or biopsy) and cancer detection rates were calculated for first and subsequent mammograms, and within 5-year age groups. RESULTS: Recall rates were approximately twice as high in the United States than in the United Kingdom for all age groups; however, cancer rates were similar. Among women aged 50 to 54 years who underwent a first screening mammogram, 14.4% in the BCSC and 12.5% in the NBCCEDP were recalled for further evaluation vs only 7.6% in the NHSBSP. Cancer detection rates per 1000 mammogram screens were 5.8, 5.9, and 6.3, in the BCSC, NBCCEDP, and NHSBSP, respectively. Recall rates were lower for subsequent examinations in all 3 settings but remained twice as high in the United States. A similar percentage of women underwent biopsy in each setting, but rates of percutaneous biopsy were lower and open surgical biopsy higher in the United States. Open surgical biopsies not resulting in a diagnosis of cancer (negative biopsies) were twice as high in the United States than in the United Kingdom. Based on a 10-year period of screening 1000 women aged 50 to 59 years, 477, 433, and 175 women in the BCSC, NBCCEDP, and NHSBSP, respectively, would be recalled; and for women aged 60 to 69 years, 396, 334, and 133 women, respectively. The estimated cancer detection rates per 1000 women aged 50 to 59 years were 24.5, 23.8, and 19.4, respectively, and for women aged 60 to 69 years, 31.5, 26.6, and 27.9, respectively. CONCLUSIONS: Recall and negative open surgical biopsy rates are twice as high in US settings than in the United Kingdom but cancer detection rates are similar. Efforts to improve US mammographic screening should target lowering the recall rate without reducing the cancer detection rate.

Aged↗

Association of volume and volume-independent factors with accuracy in screening mammogram interpretation.

BACKGROUND: Early detection of breast cancer is associated with the accurate reading of screening mammograms, but factors that influence reading accuracy are not well understood. We thus investigated whether reading volume and other factors were independently associated with accuracy in reading screening mammograms in a population of U.S. radiologists. METHODS: A random selection of 110 of 292 radiologists who agreed to participate, if selected, interpreted screening mammograms from 148 randomly selected women. Original index mammograms (i.e., mediolateral oblique and craniocaudal views of each breast) were used; comparison original mammograms were provided when available. Radiologist-level and facility-level factors were surveyed. Two standard metrics of screening accuracy, both based on receiver operating characteristic curves, were analyzed. The influence of volume on accuracy after controlling for other factors was assessed with multiple regression analysis. RESULTS: Current reading volume was not statistically significantly associated with interpretive accuracy. More recently trained radiologists interpreted mammograms more accurately than those trained earlier (-0.76% [95% confidence interval (CI) = -1.75% to -0.02%] reduction in sensitivity per year since residency). Facility-level factors that were statistically significantly and independently associated with better accuracy were the number of diagnostic breast imaging examinations and image-guided breast interventional procedures performed (0.55% [95% CI = 0.11% to 2.40%] increase in accuracy per examination or procedure offered), being classified as a comprehensive breast diagnostic and/or screening center or freestanding mammography center (1.39% [95% CI = 0.15% to 3.82%] higher than a hospital radiology department or multispecialty medical clinic), and being a facility that practiced double reading (1.61% [95% CI = 1.99% to 11.65%]) higher than in a facility without such practice). CONCLUSIONS: Individual radiologists' current reading volume was not statistically significantly associated with accuracy in reading screening mammograms, but several other factors were. Expertise reflects a complex multifactorial process that needs further clarification.

Adult↗

Analysis of 172 subtle findings on prior normal mammograms in women with breast cancer detected at follow-up screening.

PURPOSE: To retrospectively review nonspecific findings on prior screening mammograms to determine what features were most often deemed normal or benign despite the development of breast cancer in the same location detected at follow-up screening. MATERIALS AND METHODS: Four hundred ninety-three pairs of consecutive mammographic findings were collected from 13 institutions, consisting of initial normal screening findings and a subsequent finding of cancer at screening (mean interval between examinations, 14.6 months). One designated radiologist reviewed each pair of mammograms and determined that 286 findings were judged visible at prior examination in locations where cancer later developed. Five blinded radiologists independently reviewed the prior findings in these 286 cases, identifying 169 mammograms (172 cancers) with findings so subtle that none or only one or two of the five radiologists recommended screening recall. Two unblinded radiologists reviewed the initial and subsequent findings and recorded descriptors and assessments for each finding and subjective factors influencing why, although the lesion was perceptible, it might have been undetected or not recalled. RESULTS: Of 172 cancers, 129 (75%) were invasive (112 T1 tumors and 17 T2 tumors or higher; median diameter, 10 mm), and 43 (25%) were ductal carcinoma in situ (median size, 10 mm). On the prior mammograms, 80% (137 of 172) of these cancers had subtle nonspecific findings where cancer later developed, and most were assessed as being normal or benign in appearance. CONCLUSION: There is a subset of cancers that display perceptible but nonspecific mammographic findings that do not warrant recall, as judged by both a majority of blinded radiologists and by unblinded reviewers. We believe failure to act on these nonspecific findings prospectively does not necessarily constitute interpretation below a reasonable standard of care.

Adult↗

Evaluation of proscriptive health care policy implementation in screening mammography.

PURPOSE: To evaluate the potential effect of proscriptive health care policies directed toward improving screening mammogram interpretation in the United States. MATERIALS AND METHODS: Percentiles of accuracy based on a random sample of 110 U.S. radiologists were used to examine the number of radiologists who would need to be restricted from providing mammographic interpretation to increase median accuracy from 66% to 67%, 71%, and 76%. In addition, reading volume data recorded for the sampled readers were used to project the percentage reduction in service volume (mammograms per year) that would result from restriction. Characteristics of participating radiologists were compared with those of nonparticipating radiologists by using chi2 testing and analysis of variance to assess the external validity of the results. RESULTS: To increase median accuracy by 1% (from 66% to 67%) would require prohibiting about 2,200 U.S. radiologists (ie, the 11% in the lowest quantile for accuracy) from performing mammographic interpretation and would result in a reduction of yearly service volume of approximately 10%. An increase in median accuracy of 5% (to 71%) would require prohibiting about 6,000 U.S. radiologists (ie, 30%) from performing this service, with an accompanying volume reduction of 25%. An increase in median accuracy of 10% (to 77%) would require prohibiting about 11,400 practicing U.S. radiologists (ie, 57%) from performing this service and would diminish the national service capacity by 50%. CONCLUSION: These data show that implementation of proscriptive health care policies based on accuracy would diminish the service capacity of screening mammography in the United States.

Adult↗

Factors affecting radiologist inconsistency in screening mammography.

RATIONALE AND OBJECTIVES: Although research has successfully documented variability in radiologists' interpretation of mammograms, it has failed to determine the relative contributions of case-specific features and reader inconsistency. Training interventions to improve consistency will be ineffectual if they do not target the principal determinants of disagreement among radiologists. The current study assessed the relative contributions of the case and the interpreter to the problem of inconsistent interpretation. MATERIALS AND METHODS: One hundred ten radiologists independently interpreted mammograms from the same 148 screening cases (43% with biopsy-proved cancers) and reported the presence or absence of calcifications, mass, architectural distortion, and asymmetric density in each of 296 breasts. The radiologists were blinded to disease status (established at biopsy or follow-up). RESULTS: Case-related differences accounted for a greater proportion of interpretation disagreement than did differences between interpreters. The presence of cancer was associated with increased disagreement, perhaps because of the multiplicity of findings. Patient age was also associated with increased disagreement in the reporting of calcifications. CONCLUSION: For screening mammography, increased consistency between radiologists in their recognition and reporting of clinically important findings will best be achieved by reducing disagreement in difficult cases. Current training in the United States addresses difficult cases only as they have been defined intuitively or experientially. The authors' population-based method provides an objective metric to measure case difficulty and basis from which to identify difficult cases for targeted training.

Adult↗