PubMed Health⌕ Search

PubMed · 10535635

Feature selection with limited datasets.

Abstract

Computer-aided diagnosis has the potential of increasing diagnostic accuracy by providing a second reading to radiologists. In many computerized schemes, numerous features can be extracted to describe suspect image regions. A subset of these features is then employed in a data classifier to determine whether the suspect region is abnormal or normal. Different subsets of features will, in general, result in different classification performances. A feature selection method is often used to determine an "optimal" subset of features to use with a particular classifier. A classifier performance measure (such as the area under the receiver operating characteristic curve) must be incorporated into this feature selection process. With limited datasets, however, there is a distribution in the classifier performance measure for a given classifier and subset of features. In this paper, we investigate the variation in the selected subset of "optimal" features as compared with the true optimal subset of features caused by this distribution of classifier performance. We consider examples in which the probability that the optimal subset of features is selected can be analytically computed. We show the dependence of this probability on the dataset sample size, the total number of features from which to select, the number of features selected, and the performance of the true optimal subset. Once a subset of features has been selected, the parameters of the data classifier must be determined. We show that, with limited datasets and/or a large number of features from which to choose, bias is introduced if the classifier parameters are determined using the same data that were employed to select the "optimal" subset of features.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M A Kupinski, M L Giger. 1999. Feature selection with limited datasets.. https://doi.org/10.1118/1.598821

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Assessment of blinding in pharmacotherapy and noninvasive neuromodulation randomized controlled trials for neuropathic pain in adults.

In randomized controlled trials (RCTs), study participants and research personnel are often blinded to minimize biases related to knowing treatment allocation. To determine if blinding was effective, participants may be asked which treatment they believe they received ("treatment guess"). This descriptive review characterized blinding assessment (BA) reporting in pharmacotherapy and neuromodulation neuropathic pain RCTs. Of 288 papers, 36 (12.5%) reported a BA. One paper reported the results of 2 studies, so in total 37 studies with a BA were assessed. Of these, 19 were crossover, 17 parallel, and 1 partial crossover in design. All 37 studies assessed participant blinding, and 10 also assessed investigator blinding. Approximately 27% included an "unsure" answer option for treatment guess, and 38% asked the reason for the guess. There were no clear patterns in BA reporting across time nor based on treatment type. Seventeen trials provided sufficient data to calculate Bang Blinding Index (BI) to determine blinding success. Participants remained blinded (BI = 0 &#xb1; 0.2) in 10/17 placebo and 10/17 treatment arms, 6 placebo and 5 treatment arms had a BI > 0.2 suggesting possible unblinding, whereas 1 placebo and 2 treatment arms had a BI < -0.2 suggesting misinformed guessing. Overall, we found that BAs are done in a minority of published neuropathic pain trials and with variable methodology. Given the importance of minimizing risk of bias because of treatment unblinding, future studies should consider including BAs, and further consensus building is necessary to determine if and how BAs should be conducted and interpreted in analgesic clinical trials.

Bias↗

Country estimates of maternal mortality: an alternative model.

Ever since the publication of country level estimates of maternal mortality for 1990 by WHO and UNICEF, there has been some degree of controversy about these estimates. The recent publication of a 1995 revision, based on the modification of the multivariate model used for 1990, has not managed to put this controversy to rest. Countries with national estimates of their own have generally protested against the higher figures resulting from the multivariate modelling approach used by WHO and UNICEF, but some experts have also objected to the model itself. As a result of earlier discussions with the WHO/UNICEF team, some adjustments were incorporated into their model, notably the age standardization of maternal mortality ratios (MMRs) and proportions maternal among deaths of females of reproductive age (PMDF) of demographic and health surveys (DHS) direct sisterhood data, as the use of unstandardized values was shown to cause systematic biases. However, a model feature that continued to be controversial was the use of the PMDF as the dependent variable. As will be shown in this paper, the use of this dependent variable has a number of conceptual and practical disadvantages, such as its dependence on non-maternal deaths and the need for separate projections of births and deaths of women of reproductive age, in order to convert the estimated PMDF into a more conventional MMR. The latter greatly increases the uncertainty of the resulting MMR estimates, even though this additional variance is ignored in the WHO/UNICEF estimates of confidence intervals. On balance, the MMR, while also subject to some legitimate objections, is still considered preferable as an independent variable. This paper therefore derives alternative country estimates for 1995 based on a multivariate model of the MMR. The model is shown to lead to smaller root mean square relative errors of the MMR estimates. While the overall number of maternal deaths estimated worldwide is very similar to the number reached by WHO/UNICEF, there are major disagreements with respect to particular countries. Finally, a discussion is included on the appropriate way to incorporate the DHS direct sisterhood data, as this affects the results substantially.

Bias↗

Evaluating the safety of medicines, with particular reference to contraception.

Toxicological studies and clinical trials cannot be expected to predict all important adverse effects of medicines and contraceptives. Post-marketing surveillance is essentially an epidemiological task that involves detecting associations between drugs and events. The first alerts about drug safety problems have often come from case reports, but epidemiological studies are needed to confirm adverse (or beneficial) effects and to provide quantitative information. This article illustrates methodological principles by considering three examples from the field of contraceptive safety: oral contraceptives and breast cancer, intrauterine contraception and pelvic inflammatory disease, and newer oral contraceptives and venous thromboembolism. Key issues that emerge include bias and confounding, the place of subgroup analyses, random error, and the use of computerized databases. In research on contraceptive and drug safety, conclusions usually need to be based on careful assessment of multiple observational studies.

Bias↗