PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

SVM-based feature selection for characterization of focused compound collections.

Artificial neural networks, the support vector machine (SVM), and other machine learning methods for the classification of molecules are often considered as a "black box", since the molecular features that are most relevant for a given classifier are usually not presented in a human-interpretable form. We report on an SVM-based algorithm for the selection of relevant molecular features from a trained classifier that might be important for an understanding of ligand-receptor interactions. The original SVM approach was extended to allow for feature selection. The method was applied to characterize focused libraries of enzyme inhibitors. A comparison with classical Kolmogorov-Smirnov (KS)-based feature selection was performed. In most of the applications the SVM method showed sustained classification accuracy, thereby relying on a smaller number of molecular features than KS-based classifiers. In one case both methods produced comparable results. Limiting the calculation of descriptors to only the most relevant ones for a certain biological activity can also be used to speed up high-throughput virtual screening.

Algorithms↗

A neural network model with feature selection for Korean speech act classification.

A speech act is a linguistic action intended by a speaker. Speech act classification is an essential part of a dialogue understanding system because the speech act of an utterance is closely tied with the user's intention in the utterance. We propose a neural network model for Korean speech act classification. In addition, we propose a method that extracts morphological features from surface utterances and selects effective ones among the morphological features. Using the feature selection method, the proposed neural network can partially increase precision and decrease training time. In the experiment, the proposed neural network showed better results than other models using comparatively high-level linguistic features. Based on the experimental result, we believe that the proposed neural network model is suitable for real field applications because it is easy to expand the neural network model into other domains. Moreover, we found that neural networks can be useful in speech act classification if we can convert surface sentences into vectors with fixed dimensions by using an effective feature selection method.

Asian People↗

Visual categorization shapes feature selectivity in the primate temporal cortex.

The way that we perceive and interact with objects depends on our previous experience with them. For example, a bird expert is more likely to recognize a bird as a sparrow, a sandpiper or a cockatiel than a non-expert. Neurons in the inferior temporal cortex have been shown to be important in the representation of visual objects; however, it is unknown which object features are represented and how these representations are affected by categorization training. Here we show that feature selectivity in the macaque inferior temporal cortex is shaped by categorization of objects on the basis of their visual features. Specifically, we recorded from single neurons while monkeys performed a categorization task with two sets of parametric stimuli. Each stimulus set consisted of four varying features, but only two of the four were important for the categorization task (diagnostic features). We found enhanced neuronal representation of the diagnostic features relative to the non-diagnostic ones. These findings demonstrate that stimulus features important for categorization are instantiated in the activity of single units (neurons) in the primate inferior temporal cortex.

Animals↗

Prediction of lymph node metastasis and perineural invasion of biliary tract cancer by selected features from cDNA array data.

OBJECTIVE: To better understand the nature of the malignancy of biliary tract carcinoma and evaluate the feasibility of its prediction by gene expression profiles. METHODS AND RESULTS: We explored the gene expression profiles characteristic of progression and invasiveness in the cDNA array data obtained from 37 biliary tract carcinomas (15 bile duct, 11 gallbladder, 11 of ampulla of Vater). We pre-selected 51 and 100 genes for the presence versus absence of lymph node metastasis and perineural invasion on the basis of statistical difference. To search optimized sets of genes for prediction, we applied a sequential forward feature selection, minimizing leave-one-out error rates on a k-nearest neighbor classifier. We could predict lymph node metastasis and perineural invasion with an accuracy of 94 and 100%, respectively. When the 6-stage IA cancers without perineural invasion were precluded, a marked difference in gene expression (147 gene), discriminable with 100% accuracy, was noted between positive versus negative perineural invasion, suggesting that the acquisition of invasive character is rather a later molecular pathological event in biliary tract cancer. CONCLUSION: The present method provides a powerful means of classifying biliary tract carcinomas. We also suggest that perineural invasion is an important target of array databased pattern classification, which may predict patient outcomes and facilitate the determination of the extent of surgery to minimize the risk of recurrence.

Biliary Tract↗

Hierarchical learning architecture with automatic feature selection for multiclass protein fold classification.

The structure classification of proteins plays a very important role in bioinformatics, since the relationships and characteristics among those known proteins can be exploited to predict the structure of new proteins. The success of a classification system depends heavily on two things: the tools being used and the features considered. For the bioinformatics applications, the role of appropriate features has not been paid adequate importance. In this investigation we use three novel ideas for multiclass protein fold classification. First, we use the gating neural network, where each input node is associated with a gate. This network can select important features in an online manner when the learning goes on. At the beginning of the training, all gates are almost closed, i.e., no feature is allowed to enter the network. Through the training, gates corresponding to good features are completely opened while gates corresponding to bad features are closed more tightly, and some gates may be partially open. The second novel idea is to use a hierarchical learning architecture (HLA). The classifier in the first level of HLA classifies the protein features into four major classes: all alpha, all beta, alpha + beta, and alpha/beta. And in the next level we have another set of classifiers, which further classifies the protein features into 27 folds. The third novel idea is to induce the indirect coding features from the amino-acid composition sequence of proteins based on the N-gram concept. This provides us with more representative and discriminative new local features of protein sequences for multiclass protein fold classification. The proposed HLA with new indirect coding features increases the protein fold classification accuracy by about 12%. Moreover, the gating neural network is found to reduce the number of features drastically. Using only half of the original features selected by the gating neural network can reach comparable test accuracy as that using all the original features. The gating mechanism also helps us to get a better insight into the folding process of proteins. For example, tracking the evolution of different gates we can find which characteristics (features) of the data are more important for the folding process. And, of course, it also reduces the computation time.

Algorithms↗

Genetic programming for classification and feature selection: analysis of 1H nuclear magnetic resonance spectra from human brain tumour biopsies.

Genetic programming (GP) is used to classify tumours based on 1H nuclear magnetic resonance (NMR) spectra of biopsy extracts. Analysis of such data would ideally give not only a classification result but also indicate which parts of the spectra are driving the classification (i.e. feature selection). Experiments on a database of variables derived from 1H NMR spectra from human brain tumour extracts (n = 75) are reported, showing GP's classification abilities and comparing them with that of a neural network. GP successfully classified the data into meningioma and non-meningioma classes. The advantage over the neural network method was that it made use of simple combinations of a small group of metabolites, in particular glutamine, glutamate and alanine. This may help in the choice of the most informative NMR spectroscopy methods for future non-invasive studies in patients.

Artificial Intelligence↗

Extended glutamate activates metabotropic receptor types 1, 2 and 4: selective features at mGluR4 binding site.

To get an insight into the bioactive conformation of glutamic acid and its topological environment at the mGluR4 binding site, a pharmacophore model was constructed using molecular modeling. Agonists of known activities were used to run the Apex-3D program or to validate the resulting model. An extended glutamate conformer, two selective hydrophilic sites and bulk tolerance regions are disclosed. Selective features of mGluR1, mGluR2 and mGluR4 are discussed.

Animals↗

Age differences in feature selection in triple conjunction search.

Younger and older participants were trained in a triple conjunction visual search task to examine age differences in the development of proficient performance. For the first 8 days, participants searched for a target defined by its contrast polarity, shape, and orientation. On Days 9 through 16, the target identity was switched to one defined by opposing feature values. On Day 17, the target was returned to the original feature values. Results indicated that, after training, younger adults reduced their display size effects more than elderly adults. Disruption occurred after the first but not after the second transfer. However, each time the target was switched, there were no age differences in disruption. Eye movement data suggest that older adults use a similar feature selection strategy as younger adults but may be more susceptible to distraction. The results are discussed in terms of current models of attention and search.

Adult↗

Minimum redundancy feature selection from microarray gene expression data.

How to selecting a small subset out of the thousands of genes in microarray data is important for accurate classification of phenotypes. Widely used methods typically rank genes according to their differential expressions among phenotypes and pick the top-ranked genes. We observe that feature sets so obtained have certain redundancy and study methods to minimize it. We propose a minimum redundancy - maximum relevance (MRMR) feature selection framework. Genes selected via MRMR provide a more balanced coverage of the space and capture broader characteristics of phenotypes. They lead to significantly improved class predictions in extensive experiments on 6 gene expression data sets: NCI, Lymphoma, Lung, Child Leukemia, Leukemia, and Colon. Improvements are observed consistently among 4 classification methods: Naive Bayes, Linear discriminant analysis, Logistic regression, and Support vector machines. SUPPLIMENTARY: The top 60 MRMR genes for each of the datasets are listed in http://crd.lbl.gov/~cding/MRMR/. More information related to MRMR methods can be found at http://www.hpeng.net/.

Algorithms↗

Tissue counter analysis of histologic sections of melanoma: influence of mask size and shape, feature selection, statistical methods and tissue preparation.

BACKGROUND: Tissue counter analysis is an image analysis tool designed for the detection of structures in complex images at the macroscopic or microscopic scale. As a basic principle, small square or circular measuring masks are randomly placed across the image and image analysis parameters are obtained for each mask. Based on learning sets, statistical classification procedures are generated which facilitate an automated classification of new data sets. OBJECTIVE: To evaluate the influence of the size and shape of the measuring masks as well as the importance of feature selection, statistical procedures and technical preparation of slides on the performance of tissue counter analysis in microscopic images. As main quality measure of the final classification procedure, the percentage of elements that were correctly classified was used. STUDY DESIGN: HE-stained slides of 25 primary cutaneous melanomas were evaluated by tissue counter analysis for the recognition of melanoma elements (section area occupied by tumour cells) in contrast to other tissue elements and background elements. Circular and square measuring masks, various subsets of image analysis features and classification and regression trees compared with linear discriminant analysis as statistical alternatives were used. The percentage of elements that were correctly classified by the various classification procedures was assessed. In order to evaluate the applicability to slides obtained from different laboratories, the best procedure was automatically applied in a test set of another 50 cases of primary melanoma derived from the same laboratory as the learning set and two test sets of 20 cases each derived from two different laboratories, and the measurements of melanoma area in these cases were compared with conventional assessment of vertical tumour thickness. RESULTS: Square measuring masks were slightly superior to circular masks, and larger masks (64 or 128 pixels in diameter) were superior to smaller masks (8 to 32 pixels in diameter). As far as the subsets of image analysis features were concerned, colour features were superior to densitometric and Haralick texture features. Statistical moments of the grey level distribution were of least significance. CART (classification and regression tree) analysis turned out to be superior to linear discriminant analysis. In the best setting, 95% of melanoma tissue elements were correctly recognized. Automated measurement of melanoma area in the independent test sets yielded a correlation of r=0.846 with vertical tumour thickness (p<0.001), similar to the relationship reported for manual measurements. The test sets obtained from different laboratories yielded comparable results. CONCLUSIONS: Large, square measuring masks, colour features and CART analysis provide a useful setting for the automated measurement of melanoma tissue in tissue counter analysis, which can also be used for slides derived from different laboratories.

Adult↗

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNP). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods.

Artificial Intelligence↗

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNPs). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods. Supplementary website: http://htsnp.stanford.edu/FSFS/.

Algorithms↗

Feature selection and transduction for prediction of molecular bioactivity for drug design.

MOTIVATION: In drug discovery a key task is to identify characteristics that separate active (binding) compounds from inactive (non-binding) ones. An automated prediction system can help reduce resources necessary to carry out this task. RESULTS: Two methods for prediction of molecular bioactivity for drug design are introduced and shown to perform well in a data set previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001. The data is characterized by very few positive examples, a very large number of features (describing three-dimensional properties of the molecules) and rather different distributions between training and test data. Two techniques are introduced specifically to tackle these problems: a feature selection method for unbalanced data and a classifier which adapts to the distribution of the the unlabeled test data (a so-called transductive method). We show both techniques improve identification performance and in conjunction provide an improvement over using only one of the techniques. Our results suggest the importance of taking into account the characteristics in this data which may also be relevant in other problems of a similar type.

Algorithms↗

Iterative class discovery and feature selection using Minimal Spanning Trees.

BACKGROUND: Clustering is one of the most commonly used methods for discovering hidden structure in microarray gene expression data. Most current methods for clustering samples are based on distance metrics utilizing all genes. This has the effect of obscuring clustering in samples that may be evident only when looking at a subset of genes, because noise from irrelevant genes dominates the signal from the relevant genes in the distance calculation. RESULTS: We describe an algorithm for automatically detecting clusters of samples that are discernable only in a subset of genes. We use iteration between Minimal Spanning Tree based clustering and feature selection to remove noise genes in a step-wise manner while simultaneously sharpening the clustering. Evaluation of this algorithm on synthetic data shows that it resolves planted clusters with high accuracy in spite of noise and the presence of other clusters. It also shows a low probability of detecting spurious clusters. Testing the algorithm on some well known micro-array data-sets reveals known biological classes as well as novel clusters. CONCLUSIONS: The iterative clustering method offers considerable improvement over clustering in all genes. This method can be used to discover partitions and their biological significance can be determined by comparing with clinical correlates and gene annotations. The MATLAB programs for the iterative clustering algorithm are available from http://linus.nci.nih.gov/supplement.html

Acute Disease↗

Support vector machine-based feature selection for classification of liver fibrosis grade in chronic hepatitis C.

Although liver biopsy is currently regarded as the gold standard for staging liver fibrosis in chronic hepatitis C, it is a costly invasive procedure and carries a small risk for complication. Our aim in this study was to construct a simple model to distinguish between patients with no or mild fibrosis (METAVIR F0-F1) versus those with clinically significant fibrosis (METAVIR F2-F4). We retrospectively studied 204 consecutive CHC patients. Thirty-four serum markers with age, gender, duration of infection were assessed to classify fibrosis with a classifier known as the support vector machine (SVM). The method of feature selection known as sequential forward floating selection (SFFS) was introduced before the performance of SVM. When four serum markers were extracted with SFFS-SVM, F2-F4 could be predicted accurately in 96%. Our study showed that application of this model could identify CHC patients with clinically significant fibrosis with a high degree of accuracy and may decrease the need for liver biopsy.

Adult↗

Comparison of linear, nonlinear, and feature selection methods for EEG signal classification.

The reliable operation of brain-computer interfaces (BCIs) based on spontaneous electroencephalogram (EEG) signals requires accurate classification of multichannel EEG. The design of EEG representations and classifiers for BCI are open research questions whose difficulty stems from the need to extract complex spatial and temporal patterns from noisy multidimensional time series obtained from EEG measurements. The high-dimensional and noisy nature of EEG may limit the advantage of nonlinear classification methods over linear ones. This paper reports the results of a linear (linear discriminant analysis) and two nonlinear classifiers (neural networks and support vector machines) applied to the classification of spontaneous EEG during five mental tasks, showing that nonlinear classifiers produce only slightly better classification results. An approach to feature selection based on genetic algorithms is also presented with preliminary results of application to EEG during finger movement.

Algorithms↗

Fast feature selection using a simple estimation of distribution algorithm: a case study on splice site prediction.

MOTIVATION: Feature subset selection is an important preprocessing step for classification. In biology, where structures or processes are described by a large number of features, the elimination of irrelevant and redundant information in a reasonable amount of time has a number of advantages. It enables the classification system to achieve good or even better solutions with a restricted subset of features, allows for a faster classification, and it helps the human expert focus on a relevant subset of features, hence providing useful biological knowledge. RESULTS: We present a heuristic method based on Estimation of Distribution Algorithms to select relevant subsets of features for splice site prediction in Arabidopsis thaliana. We show that this method performs a fast detection of relevant feature subsets using the technique of constrained feature subsets. Compared to the traditional greedy methods the gain in speed can be up to one order of magnitude, with results being comparable or even better than the greedy methods. This makes it a very practical solution for classification tasks that can be solved using a relatively small amount of discriminative features (or feature dependencies), but where the initial set of potential discriminative features is rather large.

Algorithms↗

Cholesterol-free phospholipid domains may be the membrane feature selected by N epsilon-dansyl-L-lysine and merocyanine 540.

We have used N epsilon-dansyl-L-lysine as a fluorescent membrane probe, to study cells taken from tissues concerned with immune function. There is a striking similarity between the staining selectivity of this compound and that reported by others for merocyanine 540. Both compounds stain leukemic, human, peripheral leukocytes, an erythroleukemia line, and some mouse bone marrow cells, suggesting common selectivity for a membrane feature of hemopoietic cells. Both compounds fail to stain red blood cells, normal human leukocytes, mouse spleen and thymus cells. We have recently reported that dansyl-lysine apparently selects for cholesterol-free phospholipid domains in liposomes and now report similar selectivity for merocyanine 540 staining of liposomes.

Animals↗