PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The potential of feature selection by statistical techniques and the use of statistical classifiers in the discrimination of benign from malignant gastric lesions.

The objective of this study was the investigation of the potential value of morphometry, feature selection and statistical classifiers techniques, such as neural networks, for the classification of benign from malignant gastric nuclei and cases. One hundred and twenty gastric smears, routinely processed and stained by Papanicolaou technique, were analyzed by a customized image analysis system. Data from half of the cases were selected to form the training set, while the remaining data formed the test set. A feature selection technique was applied in order to identify the most important nuclear features, which were used in a second stage by statistical classifiers to classify a nucleus as benign or malignant. Using the classifier results for the nuclear classification, a method to classify each individual patient was developed. The performance of the proposed method was validated through the test set. The technique described in this report produces significant results at the nuclear and patient level and promises to be a powerful assistance tool for everyday cytological laboratory routine.

Anthropometry↗

Mining for diagnostic information in body surface potential maps: a comparison of feature selection techniques.

BACKGROUND: In body surface potential mapping, increased spatial sampling is used to allow more accurate detection of a cardiac abnormality. Although diagnostically superior to more conventional electrocardiographic techniques, the perceived complexity of the Body Surface Potential Map (BSPM) acquisition process has prohibited its acceptance in clinical practice. For this reason there is an interest in striking a compromise between the minimum number of electrocardiographic recording sites required to sample the maximum electrocardiographic information. METHODS: In the current study, several techniques widely used in the domains of data mining and knowledge discovery have been employed to mine for diagnostic information in 192 lead BSPMs. In particular, the Single Variable Classifier (SVC) based filter and Sequential Forward Selection (SFS) based wrapper approaches to feature selection have been implemented and evaluated. Using a set of recordings from 116 subjects, the diagnostic ability of subsets of 3, 6, 9, 12, 24 and 32 electrocardiographic recording sites have been evaluated based on their ability to correctly asses the presence or absence of Myocardial Infarction (MI). RESULTS: It was observed that the wrapper approach, using sequential forward selection and a 5 nearest neighbour classifier, was capable of choosing a set of 24 recording sites that could correctly classify 82.8% of BSPMs. Although the filter method performed slightly less favourably, the performance was comparable with a classification accuracy of 79.3%. In addition, experiments were conducted to show how (a) features chosen using the wrapper approach were specific to the classifier used in the selection model, and (b) lead subsets chosen were not necessarily unique. CONCLUSION: It was concluded that both the filter and wrapper approaches adopted were suitable for guiding the choice of recording sites useful for determining the presence of MI. It should be noted however that in this study recording sites have been suggested on their ability to detect disease and such sites may not be optimal for estimating body surface potential distributions.

Algorithms↗

Feature selection in MLPs and SVMs based on maximum output information.

This paper presents feature selection algorithms for multilayer perceptrons (MLPs) and multiclass support vector machines (SVMs), using mutual information between class labels and classifier outputs, as an objective function. This objective function involves inexpensive computation of information measures only on discrete variables; provides immunity to prior class probabilities; and brackets the probability of error of the classifier. The maximum output information (MOI) algorithms employ this function for feature subset selection by greedy elimination and directed search. The output of the MOI algorithms is a feature subset of user-defined size and an associated trained classifier (MLP/SVM). These algorithms compare favorably with a number of other methods in terms of performance on various artificial and real-world data sets.

Algorithms↗

Classification of breast masses in mammograms using genetic programming and feature selection.

Mammography is a widely used screening tool and is the gold standard for the early detection of breast cancer. The classification of breast masses into the benign and malignant categories is an important problem in the area of computer-aided diagnosis of breast cancer. A small dataset of 57 breast mass images, each with 22 features computed, was used in this investigation; the same dataset has been previously used in other studies. The extracted features relate to edge-sharpness, shape, and texture. The novelty of this paper is the adaptation and application of the classification technique called genetic programming (GP), which possesses feature selection implicitly. To refine the pool of features available to the GP classifier, we used feature-selection methods, including the introduction of three statistical measures--Student's t test, Kolmogorov-Smirnov test, and Kullback-Leibler divergence. Both the training and test accuracies obtained were high: above 99.5% for training and typically above 98% for test experiments. A leave-one-out experiment showed 97.3% success in the classification of benign masses and 95.0% success in the classification of malignant tumors. A shape feature known as fractional concavity was found to be the most important among those tested, since it was automatically selected by the GP classifier in almost every experiment.

Algorithms↗

Feature selection for descriptor based classification models. 1. Theory and GA-SEC algorithm.

The paper describes different aspects of classification models based on molecular data sets with the focus on feature selection methods. Especially model quality and avoiding a high variance on unseen data (overfitting) will be discussed with respect to the feature selection problem. We present several standard approaches and modifications of our Genetic Algorithm based on the Shannon Entropy Cliques (GA-SEC) algorithm and the extension for classification problems using boosting.

Journal Article↗

Features selection and architecture optimization in connectionist systems.

In this paper, we propose a features selection measure and an architecture optimization procedure for Multi-Layer Perceptrons (MLP). The algorithm presented in this contribution employs a heuristic measure named HVS (Heuristic for Variable Selection). This new measure allows us to identify and select important variables in the features space. This can be achieved by eliminating redundant features and those which do not contain enough relevant information. The proposed measure is used in a new procedure aimed at selecting the "best" MLP architecture given an initial structure. Application results for two generic problems: regression and discrimination, demonstrates the proposed selection algorithm's effectiveness in identifying optimized connectionist models with higher accuracy. Finally, an extension of HVS, named epsilonHVS, is proposed for discriminative features detection and architecture optimization for Time Delay Neural Networks models (TDNN).

Algorithms↗

Feature selection for the prediction of translation initiation sites.

Translation initiation sites (TISs) are important signals in cDNA sequences. In many previous attempts to predict TISs in cDNA sequences, three major factors affect the prediction performance: the nature of the cDNA sequence sets, the relevant features selected, and the classification methods used. In this paper, we examine different approaches to select and integrate relevant features for TIS prediction. The top selected significant features include the features from the position weight matrix and the propensity matrix, the number of nucleotide C in the sequence downstream ATG, the number of downstream stop codons, the number of upstream ATGs, and the number of some amino acids, such as amino acids A and D. With the numerical data generated from these features, different classification methods, including decision tree, naïve Bayes, and support vector machine, were applied to three independent sequence sets. The identified significant features were found to be biologically meaningful, while the experiments showed promising results.

Codon, Initiator↗

Feature selection for DNA methylation based cancer classification.

Molecular portraits, such as mRNA expression or DNA methylation patterns, have been shown to be strongly correlated with phenotypical parameters. These molecular patterns can be revealed routinely on a genomic scale. However, class prediction based on these patterns is an under-determined problem, due to the extreme high dimensionality of the data compared to the usually small number of available samples. This makes a reduction of the data dimensionality necessary. Here we demonstrate how phenotypic classes can be predicted by combining feature selection and discriminant analysis. By comparing several feature selection methods we show that the right dimension reduction strategy is of crucial importance for the classification performance. The techniques are demonstrated by methylation pattern based discrimination between acute lymphoblastic leukemia and acute myeloid leukemia.

Computational Biology↗

Feature selection in quantitative structure-activity relationships.

A key component of building quantitative structure-activity relationship (QSAR) models is the selection of an appropriate set of molecular features or descriptors. Feature selection can affect the accuracy, stability and interpretability of a model. There are thousands of molecular descriptors currently available, and the selection of an appropriate descriptor set for a particular model can be a daunting task. While there are no absolute rules for selecting appropriate sets of descriptors, a number of recent publications describe automated methods for identifying optimal feature sets. This review provides an overview of a number of the methods described in these publications.

Animals↗

An integrated feature selection and classification method to select minimum number of variables on the case study of gene expression data.

This paper introduces a novel generic approach for classification problems with the objective of achieving maximum classification accuracy with minimum number of features selected. The method is illustrated with several case studies of gene expression data. Our approach integrates filter and wrapper gene selection methods with an added objective of selecting a small set of non-redundant genes that are most relevant for classification with the provision of bins for genes to be swapped in the search for their biological relevance. It is capable of selecting relatively few marker genes while giving comparable or better leave-one-out cross-validation accuracy when compared with gene ranking selection approaches. Additionally, gene profiles can be extracted from the evolving connectionist system, which provides a set of rules that can be further developed into expert systems. The approach uses an integration of Pearson correlation coefficient and signal-to-noise ratio methods with an adaptive evolving classifier applied through the leave-one-out method for validation. Datasets of gene expression from four case studies are used to illustrate the method. The results show the proposed approach leads to an improved feature selection process in terms of reducing the number of variables required and an increased in classification accuracy.

Artificial Intelligence↗

Feature selection using a piecewise linear network.

We present an efficient feature selection algorithm for the general regression problem, which utilizes a piecewise linear orthonormal least squares (OLS) procedure. The algorithm 1) determines an appropriate piecewise linear network (PLN) model for the given data set, 2) applies the OLS procedure to the PLN model, and 3) searches for useful feature subsets using a floating search algorithm. The floating search prevents the "nesting effect." The proposed algorithm is computationally very efficient because only one data pass is required. Several examples are given to demonstrate the effectiveness of the proposed algorithm.

Algorithms↗

Potential of feature selection methods in heart rate variability analysis for the classification of different cardiovascular diseases.

In this study heart rate variability (HRV) analysis was applied to characterize patients suffering from coronary heart disease (CHD), dilated cardiomyopathy (DCM) and patients who had survived an acute myocardial infarction (MI). On the basis of several HRV parameters, an optimal discrimination between the different kinds of cardiovascular diseases and between the diseases and healthy controls (HC) was derived by feature selection and linear classification. For each task a small favourable subset of a set of 33 potentially interesting HRV measures was selected with the intention of improving the diagnostic value and facilitating the physiological interpretation of HRV analysis. Time- and frequency-domain parameters as well as parameters from non-linear dynamics were included in the analysis. With the expectation that different diseases are characterized by different phenomena, feature selection was applied for each task separately. Using the features optimal for one task to another task can reveal a loss in performance, but it turned out that one specific parameter set (set1: normalized low frequency LF/P and a non-linear variability measure WPSUM13) was applicable for all tasks, where diseased and healthy subjects have to be distinguished, without significant reduction in performance. This set seems to be a general marker for pathologic changes in HRV and might be used for early detection of heart diseases. The classification between different heart diseases requires another parameter set (set2: meanNN and sdaNN, reflecting the steady state behaviour of the heart rate and long and short term SEAR describing the spectral composition). However, the use of set1 for the separation of different kinds of diseases, where set2 is appropriate, led to significant reduction in performance and vice versa. This observation may be important for future developments of HRV measures especially suitable for the assessment of disease severity.

Cardiovascular Diseases↗

Hybrid genetic algorithms for feature selection.

This paper proposes a novel hybrid genetic algorithm for feature selection. Local search operations are devised and embedded in hybrid GAs to fine-tune the search. The operations are parameterized in terms of their fine-tuning power, and their effectiveness and timing requirements are analyzed and compared. The hybridization technique produces two desirable effects: a significant improvement in the final performance and the acquisition of subset-size control. The hybrid GAs showed better convergence properties compared to the classical GAs. A method of performing rigorous timing analysis was developed, in order to compare the timing requirement of the conventional and the proposed algorithms. Experiments performed with various standard data sets revealed that the proposed hybrid GA is superior to both a simple GA and sequential search algorithms.

Algorithms↗

Mutual information-based feature selection in studying perturbation of dendritic structure caused by TSC2 inactivation.

In this study, the effect of protein Tuberous sclerosis 2 (TSC2) on the dendritic spine density and length was demonstrated by using TSC2-RNAinactivation. In addition, the role of rapamycin, an antagonist of the molecular target of rapamycin, in the morphological changes of spine caused by TSC2 silencing was investigated. The features were extracted from highresolution three-dimensional image stacks collected by two-photon laser scanning microscopy of green fluorescing pyramidal cells expressing TSC2-RNA interference (RNAi), or TSC2-RNAi and rapamycin treatment in rat hippocampal slice cultures. We proposed to apply the lognormal distribution method for feature extraction. The extracted features of three cases under investigation, namely, (1) green-fluorescent protein GFP vs TSC2-RNAi, (2) GFP vs TSC2-RNAi and rapamycin, and (3) TSC2-RNAi vs TSC2-RNAi and rapamycin, were analyzed by mutual information-based feature selection and evaluated by three classifiers, K-nearest neighbor, Perceptron, and two-layer neural networks. The results showed that both the spine density and length have significant morphological changes after TSC2-RNAi treatment. However, rapamycin treatment could reverse the effect of TSC2-RNAi on spine length but not on spine density. These results are consistent with the results reported in the scientific literature. Finally, we explored the application of pattern recognition method in a small sample with richer feature properties, namely bootstrap mutual information estimation and a mutual information- based feature selection method.

Algorithms↗

Feature selection for computerized mass detection in digitized mammograms by using a genetic algorithm.

RATIONALE AND OBJECTIVES: To investigate optimization of feature selection for computerized mass detection in digitized mammograms, and to compare the effectiveness of a genetic algorithm (GA) in such optimization with that of an "exhaustive" search of all feature permutations. MATERIALS AND METHODS: A Bayesian belief network (BBN) was used to classify positive and negative regions for masses depicted in digitized mammograms; 20 features were computed for each of 592 positive and 3,790 negative regions in two databases. Conditional probabilities for the BBN were computed by using a "training" database of 288 positive and 2,204 negative regions. Performance was measured by the area under the receiver operating characteristic curve (A) by using the remainder database (304 positive and 1,586 negative regions). The optimal set was first found by using an "exhaustive" (complete permutation) searching method. A GA-based search for the optimal set then was applied, and the results of the two approaches were compared. RESULTS: As the number of features in the classifier increased, the A value increased until it reached a maximum performance for 11 features of 0.876 +/- 0.008. The A value then decreased monotonically as the number of features increased from 11 to 20. Using 100 random chromosomes (seeds) in the first generation, the GA identified the same optimal set of features but reduced the total computation time by a factor of 65. CONCLUSION: A GA-based search might be an efficient and effective approach to selecting an optimal feature set.

Algorithms↗

A neuro-fuzzy scheme for simultaneous feature selection and fuzzy rule-based classification.

Most methods of classification either ignore feature analysis or do it in a separate phase, offline prior to the main classification task. This paper proposes a neuro-fuzzy scheme for designing a classifier along with feature selection. It is a four-layered feed-forward network for realizing a fuzzy rule-based classifier. The network is trained by error backpropagation in three phases. In the first phase, the network learns the important features and the classification rules. In the subsequent phases, the network is pruned to an "optimal" architecture that represents an "optimal" set of rules. Pruning is found to drastically reduce the size of the network without degrading the performance. The pruned network is further tuned to improve performance. The rules learned by the network can be easily read from the network. The system is tested on both synthetic and real data sets and found to perform quite well.

Fuzzy Logic↗

Feature selectivity and interneuronal cooperation in the thalamocortical system.

Action potentials are a universal currency for fast information transfer in the nervous system, yet few studies address how some spikes carry more information than others. We focused on the transformation of sensory representations in the lemniscal (high-fidelity) auditory thalamocortical network. While stimulating with a complex sound, we recorded simultaneously from functionally connected cell pairs in the ventral medial geniculate body and primary auditory cortex. Thalamic action potentials that immediately preceded or potentially caused a cortical spike were more selective than the average thalamic spike for spectrotemporal stimulus features. This net improvement of thalamic signaling indicates that for some thalamic cells, spikes are not propagated through cortex independently but interact with other inputs onto the same target cell. We then developed a method to identify the spectrotemporal nature of these interactions and found that they could be cooperative or antagonistic to the average receptive field of the thalamic cell. The degree of cooperativity with the thalamic cell determined the increase in feature selectivity for potentially causal thalamic spikes. We therefore show how some thalamic spikes carry more receptive field information than average and how other inputs cooperate to constrain the information communicated through a cortical cell.

Acoustic Stimulation↗

Computerized diagnosis of Helicobacter pylori infection and associated gastric inflammation from endoscopic images by refined feature selection using a neural network.

BACKGROUND AND STUDY AIM: We investigated whether analysis of endoscopic images using a refined feature selection with neural network (RFSNN) technique could predict Helicobacter pylori-related gastric histological features. PATIENTS AND METHODS: A total of 104 dyspeptic patients were prospectively enrolled for panendoscopy and gastric biopsy for histological evaluation using the updated Sydney system. The endoscopic images of each patient were analyzed to obtain 84 image parameters. The significant image parameters from 30 randomly selected patients (15 with and 15 without H. pylori infection) associated with histological features were used to develop the RFSNN model. This was then used to test the sensitivity and specificity of the image parameters obtained from the remaining 74 patients for the prediction of the presence of H. pylori infection and related histological features. RESULTS: The RFSNN technique had a sensitivity of 85.4 % and a specificity of 90.9 % for the detection of H. pylori infection. Moreover, RFSNN was highly accurate (> 80 %) in predicting the presence of gastric atrophy, intestinal metaplasia and the severity of H. pylori-related gastric inflammation. CONCLUSIONS: RFSNN is an effective computerized technique for assessing the presence of H. pylori infection and related gastric inflammation and precancerous lesions. By using RFSNN to analyze endoscopic images, a comprehensive evaluation of the stomach may be done, thus avoiding the need for invasive but localized biopsy sampling for histological examination.

Adult↗