PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Formulation of radiometric feasibility measures for feature selection and planning in visual servoing.

Feature selection and planning are integral parts of visual servoing systems. Because many irrelevant and nonreliable image features usually exist, higher accuracy and robustness can be expected by selecting and planning good features. Assumption of perfect radiometric conditions is common in visual servoing. The following paper discusses the issue of radiometric constraints for feature selection in the context of visual servoing. Here, radiometric constraints are presented and measures are formulated to select the optimal features (in a radiometric sense) from a set of candidate features. Simulation and experimental results verify the effectiveness of the proposed measures.

Journal Article↗

Contributions of Ih to feature selectivity in layer II stellate cells of the entorhinal cortex.

Abstract Stellate cells (SCs) of the entorhinal cortex generate prominent subthreshold oscillations that are believed to be important contributors to the hippocampal theta rhythm. The slow inward rectifier Ih is expressed prominently in SCs and has been suggested to be a dominant factor in their integrative properties. We studied the input-output relationships in stellate cells (SCs) of the entorhinal cortex, both in control conditions and in the presence of the Ih antagonist ZD7288. Our results show that Ih is responsible for SCs' subthreshold resonance, and contributes to enhanced spiking reliability to theta-rich stimuli. However, SCs still exhibit other traits of rhythmicity, such as subthreshold oscillations, under Ih blockade. To clarify the effects of Ih on SC spiking, we used a generalized form of principal component analysis to show that SCs select particular features with relevant temporal signatures from stimuli. The spike-selected mix of those features varies with the frequency content of the stimulus, emphasizing the inherent nonlinearity of SC responses. A number of controls confirmed that this selectivity represents a stimulus-induced change in the cellular input-output relationship rather than an artifact of the analysis technique. Sensitivity to slow features remained statistically significant in ZD7288. However, with Ih blocked, slow stimulus features were less predictive of spikes and spikes conveyed less information about the stimulus over long time scales. Together, these results suggest that Ih is an important contributor to the input-output relationships expressed by SCs, but that other factors in SCs also contribute to subthreshold oscillations and nonlinear selectivity to slow features.

Action Potentials↗

Many are called, but few are chosen. Feature selection and error estimation in high dimensional spaces.

We address the problems of feature selection and error estimation when the number of possible feature candidates is large and the number of training samples is limited. A Monte Carlo study has been performed to illustrate the problems when using stepwise feature selection and discriminant analysis. The simulations demonstrate that in order to find the correct features, the necessary ratio of number of training samples to feature candidates is not a constant. It depends on the number of feature candidates, training samples and the Mahalanobis distance between the classes. Moreover, the leave-one-out error estimate may be a highly biased error estimate when feature selection is performed on the same data as the error estimation. It may even indicate complete separation of the classes, while no real difference between the classes exists. However, if feature selection and leave-one-out error estimation are performed in one process, an unbiased error estimate is achieved, but with high variance. The holdout error estimate gives a reliable estimate with low variance, depending on the size of the test set.

Algorithms↗

Application of the mutual information criterion for feature selection in computer-aided diagnosis.

The purpose of this study was to investigate an information theoretic approach to feature selection for computer-aided diagnosis (CAD). The approach is based on the mutual information (MI) concept. MI measures the general dependence of random variables without making any assumptions about the nature of their underlying relationships. Consequently, MI can potentially offer some advantages over feature selection techniques that focus only on the linear relationships of variables. This study was based on a database of statistical texture features extracted from perfusion lung scans. The ultimate goal was to select the optimal subset of features for the computer-aided diagnosis of acute pulmonary embolism (PE). Initially, the study addressed issues regarding the approximation of MI in a limited dataset as it is often the case in CAD applications. The MI selected features were compared to those features selected using stepwise linear discriminant analysis and genetic algorithms for the same PE database. Linear and nonlinear decision models were implemented to merge the selected features into a final diagnosis. Results showed that the MI is an effective feature selection criterion for nonlinear CAD models overcoming some of the well-known limitations and computational complexities of other popular feature selection techniques in the field.

Diagnosis, Computer-Assisted↗

Feature Selection for Classification of SELDI-TOF-MS Proteomic Profiles.

BACKGROUND: Proteomic peptide profiling is an emerging technology harbouring great expectations to enable early detection, enhance diagnosis and more clearly define prognosis of many diseases. Although previous research work has illustrated the ability of proteomic data to discriminate between cases and controls, significantly less attention has been paid to the analysis of feature selection strategies that enable learning of such predictive models. Feature selection, in addition to classification, plays an important role in successful identification of proteomic biomarker panels. METHODS: We present a new, efficient, multivariate feature selection strategy that extracts useful feature panels directly from the high-throughput spectra. The strategy takes advantage of the characteristics of surface-enhanced laser desorption/ionisation time-of-flight mass spectrometry (SELDI-TOF-MS) profiles and enhances widely used univariate feature selection strategies with a heuristic based on multivariate de-correlation filtering. We analyse and compare two versions of the method: one in which all feature pairs must adhere to a maximum allowed correlation (MAC) threshold, and another in which the feature panel is built greedily by deciding among best univariate features at different MAC levels. RESULTS: The analysis and comparison of feature selection strategies was carried out experimentally on the pancreatic cancer dataset with 57 cancers and 59 controls from the University of Pittsburgh Cancer Institute, Pittsburgh, Pennsylvania, USA. The analysis was conducted in both the whole-profile and peak-only modes. The results clearly show the benefit of the new strategy over univariate feature selection methods in terms of improved classification performance. CONCLUSION: Understanding the characteristics of the spectra allows us to better assess the relative importance of potential features in the diagnosis of cancer. Incorporation of these characteristics into feature selection strategies often leads to a more efficient data analysis as well as improved classification performance.

Journal Article↗

Identification of distinct characteristics of postural sway in Parkinson's disease: a feature selection procedure based on principal component analysis.

We selected descriptive measures of the centre of pressure (CoP) displacement in quiet standing, by means of a procedure based on principal component analysis, in two groups particularly different in terms of postural behaviours, such as subjects with Parkinson's disease (PD) in the levodopa off and on states. We computed 14 measures of the CoP: 5 measures of CoP trajectory over the support surface, 3 measures that estimated the area covered by the CoP, 1 measure that estimated the principal CoP sway direction, 1 measure that quantified the CoP total power, 1 measure that estimated the variability of CoP frequency content and 3 measures of characteristic CoP frequencies [L. Rocchi, L. Chiari, A. Cappello, Feature selection of stabilometric parameters based on principal component analysis, Med. Biol. Eng. Comput. 42 (2004) 71-79; L. Rocchi, L. Chiari, F.B. Horak, Effects of deep brain stimulation and levodopa on postural sway in Parkinson's disease, J. Neurol. Neurosurg. Psychiatry, 73 (2002) 267-274]. The feature selection, independently applied to the measures obtained in the two groups, resulted in different principal component (PC) subspaces of the 14-dimension original data set (4 PCs in the off and 3 PCs in the on state to account for over 90% of the original variance), but in the same 5 CoP measures (selected features) needed to describe the different postural behaviours: root mean square distance; mean velocity; principal sway direction; centroidal frequency of the power spectrum; frequency dispersion. The five selected features were found to provide insight into the postural control mechanisms and to describe changes in postural strategies in the two groups of PD subjects, off and on levodopa. Thus, the five selected features may be recommended for use in clinical practice and in research, in the direction toward the definition of a standard protocol in quantitative posturography.

Aged↗

Bayesian approach to feature selection and parameter tuning for support vector machine classifiers.

A Bayesian point of view of SVM classifiers allows the definition of a quantity analogous to the evidence in probabilistic models. By maximizing this one can systematically tune hyperparameters and, via automatic relevance determination (ARD), select relevant input features. Evidence gradients are expressed as averages over the associated posterior and can be approximated using Hybrid Monte Carlo (HMC) sampling. We describe how a Nyström approximation of the Gram matrix can be used to speed up sampling times significantly while maintaining almost unchanged classification accuracy. In experiments on classification problems with a significant number of irrelevant features this approach to ARD can give a significant improvement in classification performance over more traditional, non-ARD, SVM systems. The final tuned hyperparameter values provide a useful criterion for pruning irrelevant features, and we define a measure of relevance with which to determine systematically how many features should be removed. This use of ARD for hard feature selection can improve classification accuracy in non-ARD SVMs. In the majority of cases, however, we find that in data sets constructed by human domain experts the performance of non-ARD SVMs is largely insensitive to the presence of some less relevant features. Eliminating such features via ARD then does not improve classification accuracy, but leads to impressive reductions in the number of features required, by up to 75%.

Algorithms↗

Similarity-based online feature selection in content-based image retrieval.

Content-based image retrieval (CBIR) has been more and more important in the last decade, and the gap between high-level semantic concepts and low-level visual features hinders further performance improvement. The problem of online feature selection is critical to really bridge this gap. In this paper, we investigate online feature selection in the relevance feedback learning process to improve the retrieval performance of the region-based image retrieval system. Our contributions are mainly in three areas. 1) A novel feature selection criterion is proposed, which is based on the psychological similarity between the positive and negative training sets. 2) An effective online feature selection algorithm is implemented in a boosting manner to select the most representative features for the current query concept and combine classifiers constructed over the selected features to retrieve images. 3) To apply the proposed feature selection method in region-based image retrieval systems, we propose a novel region-based representation to describe images in a uniform feature space with real-valued fuzzy features. Our system is suitable for online relevance feedback learning in CBIR by meeting the three requirements: learning with small size training set, the intrinsic asymmetry property of training samples, and the fast response requirement. Extensive experiments, including comparisons with many state-of-the-arts, show the effectiveness of our algorithm in improving the retrieval performance and saving the processing time.

Algorithms↗

Fault classification on vibration data with wavelet based feature selection scheme.

Fault classification based upon vibration measurements is an essential building block of a conditional based health usage monitoring system. Multiple sensors are incorporated to assure the redundancy and to achieve the desired reliability and accuracy. The shortcoming of using multiple sensors is the need to deal with a high dimensional feature set, a computationally expensive task in classification. It is vital to reduce the feature dimension via an effective feature extraction and feature selection algorithm. A simple wavelet based feature selection scheme is proposed herein, uniquely built by local discriminant bases and genetic optimization. This scheme overcomes the disadvantages faced by the existing feature selection methods by producing a generic feature set, reducing the dimensionality of features, and requiring no prior information of the problem domain. The proposed feature selection scheme is based upon the strategy of "divide and conquer" that significantly reduce the computation time without compromising the classification performance. The simulation results show the proposed feature selection scheme provides at least 65% reduction of the total number of features at no cost of the classification accuracy.

Journal Article↗

A comparative study on feature selection and classification methods using gene expression profiles and proteomic patterns.

Feature selection plays an important role in classification. We present a comparative study on six feature selection heuristics by applying them to two sets of data. The first set of data are gene expression profiles from Acute Lymphoblastic Leukemia (ALL) patients. The second set of data are proteomic patterns from ovarian cancer patients. Based on features chosen by these methods, error rates of several classification algorithms were obtained for analysis. Our results demonstrate the importance of feature selection in accurately classifying new samples.

Algorithms↗

Feature selection and nearest centroid classification for protein mass spectrometry.

BACKGROUND: The use of mass spectrometry as a proteomics tool is poised to revolutionize early disease diagnosis and biomarker identification. Unfortunately, before standard supervised classification algorithms can be employed, the "curse of dimensionality" needs to be solved. Due to the sheer amount of information contained within the mass spectra, most standard machine learning techniques cannot be directly applied. Instead, feature selection techniques are used to first reduce the dimensionality of the input space and thus enable the subsequent use of classification algorithms. This paper examines feature selection techniques for proteomic mass spectrometry. RESULTS: This study examines the performance of the nearest centroid classifier coupled with the following feature selection algorithms. Student-t test, Kolmogorov-Smirnov test, and the P-test are univariate statistics used for filter-based feature ranking. From the wrapper approaches we tested sequential forward selection and a modified version of sequential backward selection. Embedded approaches included shrunken nearest centroid and a novel version of boosting based feature selection we developed. In addition, we tested several dimensionality reduction approaches, namely principal component analysis and principal component analysis coupled with linear discriminant analysis. To fairly assess each algorithm, evaluation was done using stratified cross validation with an internal leave-one-out cross-validation loop for automated feature selection. Comprehensive experiments, conducted on five popular cancer data sets, revealed that the less advocated sequential forward selection and boosted feature selection algorithms produce the most consistent results across all data sets. In contrast, the state-of-the-art performance reported on isolated data sets for several of the studied algorithms, does not hold across all data sets. CONCLUSION: This study tested a number of popular feature selection methods using the nearest centroid classifier and found that several reportedly state-of-the-art algorithms in fact perform rather poorly when tested via stratified cross-validation. The revealed inconsistencies provide clear evidence that algorithm evaluation should be performed on several data sets using a consistent (i.e., non-randomized, stratified) cross-validation procedure in order for the conclusions to be statistically sound.

Algorithms↗

Supervised mutual-information based feature selection for motor unit action potential classification.

A new supervised mutual information-based feature selection method is presented. Using real motor unit action potential (MUAP) data from 10 EMG signals, the performances of 32 time-sample feature sets, feature subsets selected using first- and second-order mutual information and features obtained using linear discriminant analysis (LDA) and principal component analysis (PCA) were evaluated using a minimum Euclidean distance (MED) classifier. The evaluation showed that by using only 20 first-order features or only 15 second-order features mean error rates and error rate variations equivalent to using all 32 samples or LDA or PCA could be obtained. The computational cost of first-order feature selection was considerably less than LDA, PCA and second-order feature selection. The performance of first-order features was further evaluated using a more robust classifier. Unlike the MED classifier, the robust classifier only assigned a candidate MUAP if the assignment was sufficiently certain. For the robust classifier the average error rates using 20 features were similar to using the full feature set, yet higher assignment rates were obtained. Results from both evaluations suggest that the sets of first-order features were an efficient representation of lower dimension, which provided high accuracy classification with reduced computational requirements.

Action Potentials↗

Investigating a memory-based account of negative priming: support for selection-feature mismatch.

Using typical and modified negative priming tasks, the selection-feature mismatch account of negative priming was tested. In the modified task, participants performed selections on the basis of a semantic feature (e.g., referent size). This procedure has been shown to enhance negative priming (P. A. MacDonald, S. Joordens, & K. N. Seergobin, 1999). Across 3 experiments, negative priming occurred only when the repeated item mismatched in terms of the feature used as the basis for selections. When the repeated item was congruent on the selection feature across the prime and probe displays, positive priming arose. This pattern of results appeared in both the ignored- and the attended-repetition conditions. Negative priming does not result from previously ignoring an item. These findings strongly support the selection-feature mismatch account of negative priming and refute both the distractor inhibition and the episodic-retrieval explanations.

Adult↗

Guilt-By-Association feature selection applied to simulated proteomic data.

We propose a new feature selection algorithm, Guilt-By-Association (GBA), which uses hierarchical clustering based on feature correlations to eliminate redundant features. GBA can be used in conjunction with other algorithms to produce a feature selection routine that explicitly considers both the similarities between features and their individual discriminatory powers. In this preliminary study, a simple form of GBA was investigated on simulated proteomic data.

Algorithms↗

Evaluation of Karhunen-Loève expansion for feature selection in computer-assisted classification of bioprosthetic heart-valve status.

This paper analyses the performance of four different feature-selection approaches of the Karhunen-Loève expansion (KLE) method to select the most discriminant set of features for computer-assisted classification of bioprosthetic heart-valve status. First, an evaluation test reducing the number of initial features while maintaining the performance of the original classifier is developed. Secondly, the effectiveness of the classification in a simulated practical situation where a new sample has to be classified is estimated with a validation test. Results from both tests applied to a reference database show that the most efficient feature selection and classification (> or = 97% of correct classifications (CCs)) are performed by the Kittler and Young approach. For the clinical databases, this approach provides poor classification results for simulated 'new samples' (between 50 and 69% of CCs). For both the evaluation and the validation tests, only the Heydorn and Tou approach provides classification results comparable with those of the original classifier (a difference always < or = 7%). However, the degree of feature reduction is particularly variable. The study demonstrates that the KLE feature-selection approaches are highly population-dependent. It also shows that the validation method proposed is advantageous in clinical applications where the data collection is difficult to perform.

Bioprosthesis↗

Classification of high-speed gas chromatography-mass spectrometry data by principal component analysis coupled with piecewise alignment and feature selection.

A useful methodology is introduced for the analysis of data obtained via gas chromatography with mass spectrometry (GC-MS) utilizing a complete mass spectrum at each retention time interval in which a mass spectrum was collected. Principal component analysis (PCA) with preprocessing by both piecewise retention time alignment and analysis of variance (ANOVA) feature selection is applied to all mass channels collected. The methodology involves concatenating all concurrently measured individual m/z chromatograms from m/z 20 to 120 for each GC-MS separation into a row vector. All of the sample row vectors are incorporated into a matrix where each row is a sample vector. This matrix is piecewise aligned and reduced by ANOVA feature selection. Application of the preprocessing steps (retention time alignment and feature selection) to all mass channels collected during the chromatographic separation allows considerably more selective chemical information to be incorporated in the PCA classification, and is the primary novelty of the report. This methodology is objective and requires no knowledge of the specific analytes of interest, as in selective ion monitoring (SIM), and does not restrict the mass spectral data used, as in both SIM and total ion current (TIC) methods. Significantly, the methodology allows for the classification of data with low resolution in the chromatographic dimension because of the added selectivity from the complete mass spectral dimension. This allows for the successful classification of data over significantly decreased chromatographic separation times, since high-speed separations can be employed. The methodology is demonstrated through the analysis of a set of four differing gasoline samples that serve as model complex samples. For comparison, the gasoline samples are analyzed by GC-MS over both 10-min and 10-s separation times. The successfully classified 10-min GC-MS TIC data served as the benchmark analysis to compare to the 10-s data. When only alignment and feature selection was applied to the 10-s gasoline separations using GC-MS TIC data, PCA failed. PCA was successful for 10-s gasoline separations when the methodology was applied with all the m/z information. With ANOVA feature selection, chromatographic regions with Fisher ratios greater than 1500 were retained in a new matrix and subjected to PCA yielding successful classification for the 10-s separations.

Analysis of Variance↗

Feature selection with limited datasets.

Computer-aided diagnosis has the potential of increasing diagnostic accuracy by providing a second reading to radiologists. In many computerized schemes, numerous features can be extracted to describe suspect image regions. A subset of these features is then employed in a data classifier to determine whether the suspect region is abnormal or normal. Different subsets of features will, in general, result in different classification performances. A feature selection method is often used to determine an "optimal" subset of features to use with a particular classifier. A classifier performance measure (such as the area under the receiver operating characteristic curve) must be incorporated into this feature selection process. With limited datasets, however, there is a distribution in the classifier performance measure for a given classifier and subset of features. In this paper, we investigate the variation in the selected subset of "optimal" features as compared with the true optimal subset of features caused by this distribution of classifier performance. We consider examples in which the probability that the optimal subset of features is selected can be analytically computed. We show the dependence of this probability on the dataset sample size, the total number of features from which to select, the number of features selected, and the performance of the true optimal subset. Once a subset of features has been selected, the parameters of the data classifier must be determined. We show that, with limited datasets and/or a large number of features from which to choose, bias is introduced if the classifier parameters are determined using the same data that were employed to select the "optimal" subset of features.

Bias↗

Simultaneous feature selection and clustering using mixture models.

Clustering is a common unsupervised learning technique used to discover group structure in a set of data. While there exist many algorithms for clustering, the important issue of feature selection, that is, what attributes of the data should be used by the clustering algorithms, is rarely touched upon. Feature selection for clustering is difficult because, unlike in supervised learning, there are no class labels for the data and, thus, no obvious criteria to guide the search. Another important problem in clustering is the determination of the number of clusters, which clearly impacts and is influenced by the feature selection issue. In this paper, we propose the concept of feature saliency and introduce an expectation-maximization (EM) algorithm to estimate it, in the context of mixture-based clustering. Due to the introduction of a minimum message length model selection criterion, the saliency of irrelevant features is driven toward zero, which corresponds to performing feature selection. The criterion and algorithm are then extended to simultaneously estimate the feature saliencies and the number of clusters.

Algorithms↗