PubMed Health⌕ Search

Biomedical subjects

Mia K Markey

Publications and source records attributed to Mia K Markey.

13 recordsLinked to original sources

Empirical comparison of tests for differential expression on time-series microarray experiments.

Methods for identifying differentially expressed genes were compared on time-series microarray data simulated from artificial gene networks. Select methods were further analyzed on existing immune response data of Boldrick et al. (2002, Proc. Natl. Acad. Sci. USA 99, 972-977). Based on the simulations, we recommend the ANOVA variants of Cui and Churchill. Efron and Tibshirani's empirical Bayes Wilcoxon rank sum test is recommended when the background cannot be effectively corrected. Our proposed GSVD-based differential expression method was shown to detect subtle changes. ANOVA combined with GSVD was consistent on background-normalized simulation data. GSVD with empirical Bayes was consistent without background correction. Based on the Boldrick et al. data, ANOVA is best suited to detect changes in temporal data, while GSVD and empirical Bayes effectively detect individual spikes or overall shifts, respectively. For methods tested on simulation data, lowess after background correction improved results. On simulation data without background correction, lowess decreased performance compared to median centering.

Gene Expression Profiling↗

Objective assessment of aesthetic outcomes of breast cancer treatment: measuring ptosis from clinical photographs.

The aesthetic outcome of breast cancer treatment is an important factor in breast cancer survivors' quality of life. We investigated new quantitative, objective measurements of breast ptosis based on ratios of distances between fiducial points manually identified in oblique and lateral clinical photographs. Ptosis refers to the extent to which the nipple is lower than the inframammary fold. The new objective measures were compared to ratings made using an existing subjective scale. The variability in the objective measurements due to intra- and inter-observer variability in marking fiducial points was shown to be equivalent to less than one point on the subjective ptosis scale.

Adult↗

Impact of missing data in evaluating artificial neural networks trained on complete data.

This study investigated the impact of missing data in the evaluation of artificial neural network (ANN) models trained on complete data for the task of predicting whether breast lesions are benign or malignant from their mammographic Breast Imaging and Reporting Data System (BI-RADS) descriptors. A feed-forward, back-propagation ANN was tested with three methods for estimating the missing values. Similar results were achieved with a constraint satisfaction ANN, which can accommodate missing values without a separate estimation step. This empirical study highlights the need for additional research on developing robust clinical decision support systems for realistic environments in which key information may be unknown or inaccessible.

Area Under Curve↗

A machine learning perspective on the development of clinical decision support systems utilizing mass spectra of blood samples.

Currently, the best way to reduce the mortality of cancer is to detect and treat it in the earliest stages. Technological advances in genomics and proteomics have opened a new realm of methods for early detection that show potential to overcome the drawbacks of current strategies. In particular, pattern analysis of mass spectra of blood samples has attracted attention as an approach to early detection of cancer. Mass spectrometry provides rapid and precise measurements of the sizes and relative abundances of the proteins present in a complex biological/chemical mixture. This article presents a review of the development of clinical decision support systems using mass spectrometry from a machine learning perspective. The literature is reviewed in an explicit machine learning framework, the components of which are preprocessing, feature extraction, feature selection, classifier training, and evaluation.

Artificial Intelligence↗

Breast cancer CADx based on BI-RAds descriptors from two mammographic views.

In this study we compared the performance of computer aided diagnosis (CADx) algorithms based on Breast Imaging Reporting And Data System (BI-RADS) descriptors from one or two views. To select cases for the study with different mediolateral (MLO) and craniocaudal (CC) view descriptors, we assessed the agreement in BI-RADS lesion descriptors, BI-RADS assessment, and subtlety ratings for 1626 cases from the Digital Database for Screening Mammogrpahy (DDSM) using kappa statistics. We used 115 mass caseswith different descriptors for the two views to design linear discriminant analysis (LDA) based CADx algorithms. The CADx algorithms used BI-RADS descriptors and patient age as features. Thealgorithms based on BI-RADS descriptors from both the views performed marginally betterthan algorithms based on BI-RADS descriptors from a single view. A system that averaged theresults of two classifiers trained separately on the MLO and CC views displayed the best performance (Az=0.920 +/- 0.027). Thus, some improvement in performance of BI-RADS based CADx algorithms may be achieved by combining information from two mammographic views.

Breast↗

Correspondence in texture features between two mammographic views.

It is well established that radiologists are better able to interpret mammograms when two mammographic views are available. Consequently, two mammographic projections are standard: mediolateral oblique (MLO) and craniocaudal (CC). Computer-aided diagnosis algorithms have been investigated for assisting in the detection and diagnosis of breast lesions in digitized/digital mammograms. A few previous studies suggest that computer-aided systems may also benefit from combining evidence from the two views. Intuitively, we expect that there would only be value in merging data from two views if they provide complementary information. A measure of the similarity of information is the correlation coefficient between corresponding features from the MLO and CC views. The purpose of this study was to investigate the correspondence in Haralick's texture features between the MLO and CC mammographic views of breast lesions. Features were ranked on the basis of correlation values and the two-view correlation of features for subgroups of data including masses versus calcification and benign versus malignant lesions were compared. All experiments were performed on a subset of mammography cases from the Digital Database for Screening Mammography (DDSM). It was observed that the texture features from the MLO and CC views were less strongly correlated for calcification lesions than for mass lesions. Similarly, texture features from the two views were less strongly correlated for benign lesions than for malignant lesions. These differences were statistically significant. The results suggest that the inclusion of texture features from multiple mammographic views in a CADx algorithm may impact the accuracy of diagnosis of calcification lesions and benign lesions.

Algorithms↗

Towards quantifying the aesthetic outcomes of breast cancer treatment: assessment of surgical scars.

Our long-term goal is to develop decision aids that will improve breast cancer treatment by explicitly taking aesthetics in the consideration. Essentially all breast cancer treatment involves surgery, which inevitably leaves scars. However, the extent and type of scarring is not the same for different surgeries (e.g., different forms of reconstruction.) We present our preliminary experiences in using image processing techniques to quantify scar characteristics in clinical photographs.

Breast Neoplasms↗

Guilt-By-Association feature selection applied to simulated proteomic data.

We propose a new feature selection algorithm, Guilt-By-Association (GBA), which uses hierarchical clustering based on feature correlations to eliminate redundant features. GBA can be used in conjunction with other algorithms to produce a feature selection routine that explicitly considers both the similarities between features and their individual discriminatory powers. In this preliminary study, a simple form of GBA was investigated on simulated proteomic data.

Algorithms↗

Decision tree classification of proteins identified by mass spectrometry of blood serum samples from people with and without lung cancer.

A classification and regression tree (CART) model was trained to classify 41 clinical specimens as disease/nondisease based on 26 variables computed from the mass-to-charge ratio (m/z) and peak heights of proteins identified by mass spectroscopy. The CART model built on all of the specimens (no cross-validation) had an error rate of 4/41 = 10%. The CART model suggests that mass spectra peaks in the 8000-10,000, 20,000-30,000, 45,000-60, 000, and >125,000 m/z ranges may be valuable in distinguishing between the disease/nondisease specimens. The area under the receiver operating characteristics curve was 0.80 +/- 0.07 for leave-one-out cross-validation.

Blood Proteins↗

Self-organizing map for cluster analysis of a breast cancer database.

The purpose of this study was to identify and characterize clusters in a heterogeneous breast cancer computer-aided diagnosis database. Identification of subgroups within the database could help elucidate clinical trends and facilitate future model building. A self-organizing map (SOM) was used to identify clusters in a large (2258 cases), heterogeneous computer-aided diagnosis database based on mammographic findings (BI-RADS) and patient age. The resulting clusters were then characterized by their prototypes determined using a constraint satisfaction neural network (CSNN). The clusters showed logical separation of clinical subtypes such as architectural distortions, masses, and calcifications. Moreover, the broad categories of masses and calcifications were stratified into several clusters (seven for masses and three for calcifications). The percent of the cases that were malignant was notably different among the clusters (ranging from 6 to 83%). A feed-forward back-propagation artificial neural network (BP-ANN) was used to identify likely benign lesions that may be candidates for follow up rather than biopsy. The performance of the BP-ANN varied considerably across the clusters identified by the SOM. In particular, a cluster (#6) of mass cases (6% malignant) was identified that accounted for 79% of the recommendations for follow up that would have been made by the BP-ANN. A classification rule based on the profile of cluster #6 performed comparably to the BP-ANN, providing approximately 25% specificity at 98% sensitivity. This performance was demonstrated to generalize to a large (2177) set of cases held-out for model validation.

Adult↗

Perceptron error surface analysis: a case study in breast cancer diagnosis.

Perceptrons are typically trained to minimize mean square error (MSE). In computer-aided diagnosis (CAD), model performance is usually evaluated according to other more clinically relevant measures. The purpose of this study was to investigate the relationship between MSE and the area (A(z)) under the receiver operating characteristic (ROC) curve and the high-sensitivity partial ROC area ((0.90)A'(z)). A perceptron was used to predict lesion malignancy based on two mammographic findings and patient age. For each performance measure, the error surface in weight space was visualized. Comparison of the surfaces indicated that minimizing MSE tended to maximize A(z), but not (0.90)A'(z).

Breast↗

Differences between computer-aided diagnosis of breast masses and that of calcifications.

PURPOSE: To compare the performance of a computer-aided diagnosis (CAD) system for diagnosis of previously detected lesions, based on radiologist-extracted findings on masses and calcifications. MATERIALS AND METHODS: A feed-forward, back-propagation artificial neural network (BP-ANN) was trained in a round-robin (leave-one-out) manner to predict biopsy outcome from mammographic findings (according to the Breast Imaging Reporting and Data System) and patient age. The BP-ANN was trained by using a large (>1,000 cases) heterogeneous data set containing masses and microcalcifications. The performances of the BP-ANN on masses and microcalcifications were compared with use of receiver operating characteristic analysis and a z test for uncorrelated samples. RESULTS: The BP-ANN performed significantly better on masses than microcalcifications in terms of both the area under the receiver operating characteristic curve and the partial receiver operating characteristic area index. A similar difference in performance was observed with a second model (linear discriminant analysis) and also with a second data set from a similar institution. CONCLUSION: Masses and calcifications should be considered separately when evaluating CAD systems for breast cancer diagnosis.

Adult↗

Cross-institutional evaluation of BI-RADS predictive model for mammographic diagnosis of breast cancer.

OBJECTIVE: Given a predictive model for identifying very likely benign breast lesions on the basis of Breast Imaging Reporting and Data System (BI-RADS) mammographic findings, this study evaluated the model's ability to generalize to a patient data set from a different institution. MATERIALS AND METHODS: The artificial neural network model underwent three trials: it was optimized over 500 biopsy-proven lesions from Duke University Medical Center or "Duke," evaluated on 1,000 similar cases from the University of Pennsylvania Health System or "Penn," and reoptimized for Penn. RESULTS: Trial A's Duke-only model yielded 98% sensitivity, 36% specificity, area index (A(z)) of 0.86, and partial A(z) of 0.51. The cross-institutional trial B yielded 96% sensitivity, 28% specificity, A(z) of 0.79, and partial A(z) of 0.28. The decreases were significant for both A(z) (p = 0.017) and partial A(z) (p < 0.001). In trial C, the model reoptimized for the Penn data yielded 96% sensitivity, 35% specificity, A(z) of 0.83, and partial A(z) of 0.32. There were no significant differences compared with trial B for specificity (p = 0.44) or partial A(z) (p = 0.46), suggesting that the Penn data were inherently more difficult to characterize. CONCLUSION: The BI-RADS lexicon facilitated the cross-institutional test of a breast cancer prediction model. The model generalized reasonably well, but there were significant performance decreases. The cross-institutional performance was encouraging because it was not significantly different from that of a reoptimized model using the second data set at high sensitivities. This study indicates the need for further work to collect more data and to improve the robustness of the model.

Breast Neoplasms↗