PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Feature subset selection for splice site prediction.

MOTIVATION: The large amount of available annotated Arabidopsis thaliana sequences allows the induction of splice site prediction models with supervised learning algorithms (see Haussler (1998) for a review and references). These algorithms need information sources or features from which the models can be computed. For splice site prediction, the features we consider in this study are the presence or absence of certain nucleotides in close proximity to the splice site. Since it is not known how many and which nucleotides are relevant for splice site prediction, the set of features is chosen large enough such that the probability that all relevant information sources are in the set is very high. Using only those features that are relevant for constructing a splice site prediction system might improve the system and might also provide us with useful biological knowledge. Using fewer features will of course also improve the prediction speed of the system. RESULTS: A wrapper-based feature subset selection algorithm using a support vector machine or a naive Bayes prediction method was evaluated against the traditional method for selecting features relevant for splice site prediction. Our results show that this wrapper approach selects features that improve the performance against the use of all features and against the use of the features selected by the traditional method. AVAILABILITY: The data and additional interactive graphs on the selected feature subsets are available at http://www.psb.rug.ac.be/gps

Arabidopsis↗

Computerized analysis of mammographic microcalcifications in morphological and texture feature spaces.

We are developing computerized feature extraction and classification methods to analyze malignant and benign microcalcifications on digitized mammograms. Morphological features that described the size, contrast, and shape of microcalcifications and their variations within a cluster were designed to characterize microcalcifications segmented from the mammographic background. Texture features were derived from the spatial gray-level dependence (SGLD) matrices constructed at multiple distances and directions from tissue regions containing microcalcifications. A genetic algorithm (GA) based feature selection technique was used to select the best feature subset from the multi-dimensional feature spaces. The GA-based method was compared to the commonly used feature selection method based on the stepwise linear discriminant analysis (LDA) procedure. Linear discriminant classifiers using the selected features as input predictor variables were formulated for the classification task. The discriminant scores output from the classifiers were analyzed by receiver operating characteristic (ROC) methodology and the classification accuracy was quantified by the area, Az, under the ROC curve. We analyzed a data set of 145 mammographic microcalcification clusters in this study. It was found that the feature subsets selected by the GA-based method are comparable to or slightly better than those selected by the stepwise LDA method. The texture features (Az = 0.84) were more effective than morphological features (Az = 0.79) in distinguishing malignant and benign microcalcifications. The highest classification accuracy (Az = 0.89) was obtained in the combined texture and morphological feature space. The improvement was statistically significant in comparison to classification in either the morphological (p = 0.002) or the texture (p = 0.04) feature space alone. The classifier using the best feature subset from the combined feature space and an appropriate decision threshold could correctly identify 35% of the benign clusters without missing a malignant cluster. When the average discriminant score from all views of the same cluster was used for classification, the Az value increased to 0.93 and the classifier could identify 50% of the benign clusters at 100% sensitivity for malignancy. Alternatively, if the minimum discriminant score from all views of the same cluster was used, the Az value would be 0.90 and a specificity of 32% would be obtained at 100% sensitivity. The results of this study indicate the potential of using combined morphological and texture features for computer-aided classification of microcalcifications.

Algorithms↗

Individual use of cytomorphologic characteristics in the diagnosis of endocervical columnar cell abnormalities: selection of preferred features with help of the 'Navigator' microscope.

In our previous interobserver studies on endocervical columnar cell abnormalities, we studied architectural, cellular and nuclear features in cervical smears of women known to have columnar cell atypias of variable severity, to determine cytomorphologic criteria, discriminating between mild, moderate and severe intraepithelial columnar cell lesions and adenocarcinoma. The results of these studies revealed a number of architectural, cellular and nuclear characteristics in different grades of expression, which were of importance for the primary diagnosis of: no abnormalities, different grades of intraepithelial endocervical columnar cell lesions and adenocarcinoma. Furthermore we concluded that observers used different characteristics and different grades of expression of these characteristics for comparable diagnoses. The present study was undertaken to determine those features, which were considered discriminating by each individual for the diagnosis of mild, moderate and severe atypia, adenocarcinoma in situ and adenocarcinoma. Features selected by five observers with knowledge of the final diagnosis, were stored and reviewed with help of a motor driven stage ('Navigator')-microscope and a high definition television-monitor. The results confirmed individual observer variability in the number and type of features used in the diagnosis of endocervical columnar cell abnormalities. Features such as 'variation in nuclear size and shape', 'irregular chromatin distribution' and 'coarsely granular chromatin' were selected preferentially by all observers in the diagnosis of endocervical columnar cell lesions, conversely striking differences were observed in the application of 'architectural'-and, especially in cases of 'adenocarcinoma', 'nucleolar' characteristics.

Adenocarcinoma↗

Online selection of discriminative tracking features.

This paper presents an online feature selection mechanism for evaluating multiple features while tracking and adjusting the set of features used to improve tracking performance. Our hypothesis is that the features that best discriminate between object and background are also best for tracking the object. Given a set of seed features, we compute log likelihood ratios of class conditional sample densities from object and background to form a new set of candidate features tailored to the local object/background discrimination task. The two-class variance ratio is used to rank these new features according to how well they separate sample distributions of object and background pixels. This feature evaluation mechanism is embedded in a mean-shift tracking system that adaptively selects the top-ranked discriminative features for tracking. Examples are presented that demonstrate how this method adapts to changing appearances of both tracked object and scene background. We note susceptibility of the variance ratio feature selection method to distraction by spatially correlated background clutter and develop an additional approach that seeks to minimize the likelihood of distraction.

Algorithms↗

Constructing molecular classifiers for the accurate prognosis of lung adenocarcinoma.

PURPOSE: Individualized therapy of lung adenocarcinoma depends on the accurate classification of patients into subgroups of poor and good prognosis, which reflects a different probability of disease recurrence and survival following therapy. However, it is currently impossible to reliably identify specific high-risk patients. Here, we propose a computational model system which accurately predicts the clinical outcome of individual patients based on their gene expression profiles. EXPERIMENTAL DESIGN: Gene signatures were selected using feature selection algorithms random forests, correlation-based feature selection, and gain ratio attribute selection. Prediction models were built using random committee and Bayesian belief networks. The prognostic power of the survival predictors was also evaluated using hierarchical cluster analysis and Kaplan-Meier analysis. RESULTS: The predictive accuracy of an identified 37-gene survival signature is 0.96 as measured by the area under the time-dependent receiver operating curves. The cluster analysis, using the 37-gene signature, aggregates the patient samples into three groups with distinct prognoses (Kaplan-Meier analysis, P < 0.0005, log-rank test). All patients in cluster 1 were in stage I, with N0 lymph node status (no metastasis) and smaller tumor size (T1 or T2). Additionally, a 12-gene signature correctly predicts the stage of 94.2% of patients. CONCLUSIONS: Our results show that the prediction models based on the expression levels of a small number of marker genes could accurately predict patient outcome for individualized therapy of lung adenocarcinoma. Such an individualized treatment may significantly increase survival due to the optimization of treatment procedures and improve lung cancer survival every year through the 5-year checkpoint.

Adenocarcinoma↗

Selecting a Selective Serotonin Reuptake Inhibitor: Clinically Important Distinguishing Features.

Selective serotonin reuptake inhibitors (SSRIs) are widely prescribed to treat depression. Although these drugs presumably have the same mechanism of action, they vary in several clinically important ways, including how long they remain in the body and the extent to which they interfere with the metabolism of other medications. This article reviews the pharmacologic differences among SSRIs and how these differences may affect various aspects of treatment, such as dosing, administration, and discontinuation. Understanding the distinct properties of SSRIs may help primary care physicians to design the most appropriate therapeutic plan for individual patients.

Journal Article↗

When does visual attention select all features of a distractor?

What happens after visual attention is allocated to an object? Although many theories of attention assume that all of its features are selected and processed, there has been little direct evidence that an irrelevant feature dimension of an attended nontarget is processed. In 5 experiments presented here, the authors used a singleton paradigm to investigate the effect of attention on nontarget objects. Participants made a speeded feature discrimination of a target for which the response was either compatible or incompatible with an irrelevant feature dimension of a distractor. The results show that the irrelevant distractor features were processed to the point that they interfered with the response to the target. The response compatibility effect was observed even when the location of the target or the distractor was invariant, although it was much weaker when both locations were invariant. These results demonstrate that in many circumstances, an attended distractor is completely selected and fully processed, and the complete processing of distractors depends on a number of factors, many of which are related to the strength of attention to the distractor.

Attention↗

Visual selection mediated by location: feature-based selection of noncontiguous locations.

Experiments using two different methods and three types of stimuli tested whether stimuli at non-adjacent locations could be selected simultaneously. In one set of experiments, subjects attended to red digits presented in multiple frames with green digits. Accuracy was no better when red digits appeared successively than when pairs of red digits occurred simultaneously, implying allocation of attention to the two locations simultaneously. Different tasks involving oriented grating stimuli produced the same result. The final experiment demonstrated split attention with an array of spatial probes. When the probe at one of two target locations was correctly reported, the probe at the other target location was more often reported correctly than were any of the probes at distractor locations, including those between the targets. Together, these experiments provide strong converging evidence that when two targets are easily discriminated from distractors by a basic property, spatial attention can be split across both locations.

Choice Behavior↗

Comparison of body weights, organ weights and histological features of selected organs of gnotobiotic, conventional and isolator-reared contaminated pigs.

Twenty-seven pigs from three litters were used in a comparison of body weights, organ weights, and selected histological features of germfree, conventional and isolator-reared contaminated pigs. At three weeks of age conventional pigs were heavier than pigs of the other two groups. The mandibular lymphnodes, stomachs, and small intestines of contaminated pigs were significantly heavier than the same organs of germfree pigs. This difference was not found in superficial inguinal or prefemoral lymph nodes. Other statistically significant organ weight differences were found.Histologically, the lymph nodes of conventional and contaminated pigs were much more active than those of germfree pigs, although secondary nodules were occasionally found in lymph nodes of germfree pigs. Greater quantities of iron-containing pigment were found in the spleens of germfree pigs than in spleens of the other two groups. Hepatic interlobular septa were somewhat more developed in conventional pigs than in germfree or contaminated pigs.

Adrenal Glands↗

Protein structure prediction: selecting salient features from large candidate pools.

We introduce a parallel approach, "DT-SELECT," for selecting features used by inductive learning algorithms to predict protein secondary structure. DT-SELECT is able to rapidly choose small, nonredundant feature sets from pools containing hundreds of thousands of potentially useful features. It does this by building a decision tree, using features from the pool, that classifies a set of training examples. The features included in the tree provide a compact description of the training data and are thus suitable for use as inputs to other inductive learning algorithms. Empirical experiments in the protein secondary-structure task, in which sets of complex features chosen by DT-SELECT are used to augment a standard artificial neural network representation, yield surprisingly little performance gain, even though features are selected from very large feature pools. We discuss some possible reasons for this result.

Algorithms↗

Scene-segmentation algorithm development using error measures.

Development of scene-segmentation algorithms has generally been an ad hoc process. This paper presents a systematic technique for developing these algorithms using error-measure minimization. If scene segmentation is regarded as a problem of pixel classification whereby each pixel of a scene is assigned to a particular object class, development of a scene-segmentation algorithm becomes primarily a process of feature selection. In this study, four methods of feature selection were used to develop segmentation techniques for cervical cytology images: (1) random selection, (2) manual selection (best features in the subjective judgment of the investigator), (3) eigenvector selection (ranking features according to the largest contribution to each eigenvector of the feature covariance matrix) and (4) selection using the scene-segmentation error measure A2. Four features were selected by each method from a universe of 35 features consisting of gray level, color, texture and special pixel neighborhood features in 40 cervical cytology images . Evaluation of the results was done with a composite of the scene-segmentation error measure A2, which depends on the percentage of scenes with measurable error, the agreement of pixel class proportions, the agreement of number of objects for each pixel class and the distance of each misclassified pixel to the nearest pixel of the misclassified class. Results indicate that random and eigenvector feature selection were the poorest methods, manual feature selection somewhat better and error-measure feature selection best. The error-measure feature selection method provides a useful, systematic method of developing and evaluating scene-segmentation algorithms.

Computers↗

Feature subset selection for support vector machines through discriminative function pruning analysis.

In many pattern classification applications, data are represented by high dimensional feature vectors, which induce high computational cost and reduce classification speed in the context of support vector machines (SVMs). To reduce the dimensionality of pattern representation, we develop a discriminative function pruning analysis (DFPA) feature subset selection method in the present study. The basic idea of the DFPA method is to learn the SVM discriminative function from training data using all input variables available first, and then to select feature subset through pruning analysis. In the present study, the pruning is implement using a forward selection procedure combined with a linear least square estimation algorithm, taking advantage of linear-in-the-parameter structure of the SVM discriminative function. The strength of the DFPA method is that it combines good characters of both filter and wrapper methods. Firstly, it retains the simplicity of the filter method avoiding training of a large number of SVM classifier. Secondly, it inherits the good performance of the wrapper method by taking the SVM classification algorithm into account.

Journal Article↗

Automatic MeSH term assignment and quality assessment.

For computational purposes documents or other objects are most often represented by a collection of individual attributes that may be strings or numbers. Such attributes are often called features and success in solving a given problem can depend critically on the nature of the features selected to represent documents. Feature selection has received considerable attention in the machine learning literature. In the area of document retrieval we refer to feature selection as indexing. Indexing has not traditionally been evaluated by the same methods used in machine learning feature selection. Here we show how indexing quality may be evaluated in a machine learning setting and apply this methodology to results of the Indexing Initiative at the National Library of Medicine.

Abstracting and Indexing↗

Feature subset selection and ranking for data dimensionality reduction.

A new unsupervised forward orthogonal search (FOS) algorithm is introduced for feature selection and ranking. In the new algorithm, features are selected in a stepwise way, one at a time, by estimating the capability of each specified candidate feature subset to represent the overall features in the measurement space. A squared correlation function is employed as the criterion to measure the dependency between features and this makes the new algorithm easy to implement. The forward orthogonalization strategy, which combines good effectiveness with high efficiency, enables the new algorithm to produce efficient feature subsets with a clear physical interpretation.

Algorithms↗

Using multifocal ERG responses to discriminate diabetic retinopathy.

PURPOSE: Increase the accuracy in classifying subjects with or without early diabetic retinopathy, by analyzing the multifocal electroretinogram (mfERG) responses. METHOD: The mfERG was recorded for 14 control subjects (Normal) and 26 non-insulin dependent diabetic patients, 16 of the patients had no apparent retinopathy (NDR), the other 10 had mild to moderate non-proliferative retinopathy (NPDR). The first order kernel (K1) and the first slice of second order kernel (K21) of the mfERG were summed across the field and exported as trace arrays. The feature subset was selected from the kernel arrays by criterion of inter-intra distance and sequential forward searching strategy. Based on the selected features, Fisher's linear classifiers were trained and the classification error rates were given. RESULT: With one feature selected, the error rates of the classifiers were bigger than 23% between the NDR and NPDR subjects. When six features were selected to train the classifiers, the classification error rate decreased to 6.5% between Normal and NDR, 0 between Normal and NPDR, and 9.6% between NDR and NPDR. The criterion function of inter-intra distance was calculated for each feature in the K1 and K21 arrays. According to rank of the function value, during early DR, deformations should be easier to find at the front limb of peaks N2 and P1 on the K1 trace, and at the front limb of peak N2 on the K21 trace. As a conclusion, the multifocal ERG responses have great potentials in exploring diabetic retinopathy, assisted with the pattern classification methods.

Adult↗

The use of multidimensional perceptual models in the selection of sonar echo features.

The development of an accurate and efficient sonar-target classification system depends upon the identification of a set of signal features which may be used to discriminate important classes of signals. Feature selection can be facilitated through the identification of perceptual features used by human listeners in discriminating relevant sonar echoes. This study was conducted to establish a more reliable means of identifying perceptual features in terms of physical signal parameters as an initial step toward the development of an automatic sonar-target classification system. The results of an experiment involving eight subjects and six sonar echoes are presented. A model of the perceptual structure of these echoes was derived from subject similarity judgments using a multidimensional scaling (MDS) technique. It was found that three perceptual features accounted for the similarity judgments made by the human listeners. Echoes modified along candidate physical dimensions were employed to aid in the identification of perceptual dimensions in terms of physical signal parameters. The three perceptual features could be associated with signal parameters involving the amplitude envelope of the echoes.

Auditory Perception↗

Two kinds of cognitive deficit associated with chronic schizophrenia.

The performance of 21 chronic schizophrenic patients was investigated on two tests of feature selection. It was found that patients with negative symptoms (muteness, withdrawal, etc.) were characterized by an extreme lack of persistence, but selected usual features; whereas patients with positive symptoms (hallucinations, delusions, etc.) had a normal degree of persistence, but selected unusual features.

Aged↗