PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Feature selection for descriptor based classification models. 2. Human intestinal absorption (HIA).

We show that the topological polar surface area (TPSA) descriptor and the radial distribution function (RDF) applied to electronic and steric atom properties, like the conjugated electrotopological state (CETS), are the most relevant features/descriptors for predicting the human intestinal absorption (HIA) out of a large set of 2934 features/descriptors. A HIA data set with 196 molecules with measured HIA values and 2934 features/descriptors were calculated using JOELib and MOE. We used an adaptive boosting algorithm to solve the binary classification problem (AdaBoost.M1) and Genetic Algorithms based on Shannon Entropy Cliques (GA-SEC) variants as hybrid feature selection algorithms. The selection of relevant features was applied with respect to the generalization ability of the classification model, avoiding a high variance for unseen molecules (overfitting).

Humans↗

Image segmentation feature selection and pattern classification for mammographic microcalcifications.

Since microcalcifications in X-ray mammograms are the primary indicator of breast cancer, detection of microcalcifications is central to the development of an effective diagnostic system. This paper proposes a two-stage detection procedure. In the first stage, a data driven, closed form mathematical model is used to calculate the location and shape of suspected microcalcifications. When tested on the Nijmegen University Hospital (Netherlands) database, data analysis shows that the proposed model can effectively detect the occurrence of microcalcifications. The proposed mathematical model not only eliminates the need for system training, but also provides information on the borders of suspected microcalcifications for further feature extraction. In the second stage, 61 features are extracted for each suspected microcalcification, representing texture, the spatial domain and the spectral domain. From these features, a sequential forward search (SFS) algorithm selects the classification input vector, which consists of features sensitive only to microcalcifications. Two types of classifiers-a general regression neural network (GRNN) and a support vector machine (SVM)--are applied, and their classification performance is compared using the Az value of the Receiver Operating Characteristic curve. For all 61 features used as input vectors, the test data set yielded Az values of 97.01% for the SVM and 96.00% for the GRNN. With input features selected by SFS, the corresponding Az values were 98.00% for the SVM and 97.80% for the GRNN. The SVM outperformed the GRNN, whether or not the input vectors first underwent SFS feature selection. In both cases, feature selection dramatically reduced the dimension of the input vectors (82% for the SVM and 59% for the GRNN). Moreover, SFS feature selection improved the classification performance, increasing the Az value from 97.01 to 98.00% for the SVM and from 96.00 to 97.80% for the GRNN.

Breast Neoplasms↗

On the formation of persistent states in neuronal network models of feature selectivity.

We study the existence and stability of localized activity states in neuronal network models of feature selectivity with either a ring or spherical topology. We find that the neural field has mono-stable, bi-stable, and tri-stable regimes depending on the parameters of the weighting function. In the case of homogeneous inputs, these localized activity states are marginally stable with respect to rotations. The response of a stable equilibrium to an inhomogeneous input is also determined.

Animals↗

Identification of regulatory elements using a feature selection method.

MOTIVATION: Many methods have been described to identify regulatory motifs in the transcription control regions of genes that exhibit similar patterns of gene expression across a variety of experimental conditions. Here we focus on a single experimental condition, and utilize gene expression data to identify sequence motifs associated with genes that are activated under this experimental condition. We use a linear model with two-way interactions to model gene expression as a function of sequence features (words) present in presumptive transcription control regions. The most relevant features are selected by a feature selection method called stepwise selection with monte carlo cross validation. We apply this method to a publicly available dataset of the yeast Saccharomyces cerevisiae, focussing on the 800 basepairs immediately upstream of each gene's translation start site (the upstream control region (UCR)). RESULTS: We successfully identify regulatory motifs that are known to be active under the experimental conditions analyzed, and find additional significant sequences that may represent novel regulatory motifs. We also discuss a complementary method that utilizes gene expression data from a single microarray experiment and allows averaging over variety of experimental conditions as an alternative to motif finding methods that act on clusters of co-expressed genes. AVAILABILITY: The software is available upon request from the first author or may be downloaded from http://www.stat.berkeley.edu/~sunduz. CONTACT: keles@stat.berkeley.edu

Amino Acid Motifs↗

Feature selection and the class imbalance problem in predicting protein function from sequence.

When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence features. However, there are two issues to consider before successful functional prediction can take place: identifying discriminatory features, and overcoming the challenge of a large imbalance in the training data. We show that by applying feature subset selection followed by undersampling of the majority class, significantly better support vector machine (SVM) classifiers are generated compared with standard machine learning approaches. As well as revealing that the features selected could have the potential to advance our understanding of the relationship between sequence and function, we also show that undersampling to produce fully balanced data significantly improves performance. The best discriminating ability is achieved using SVMs together with feature selection and full undersampling; this approach strongly outperforms other competitive learning algorithms. We conclude that this combined approach can generate powerful machine learning classifiers for predicting protein function directly from sequence.

Algorithms↗

Feature selection of stabilometric parameters based on principal component analysis.

This study addresses the challenge of identifying the features of the Centre of pressure (COP) trajectory that are most sensitive to postural performance, with the aim of avoiding redundancy and allowing a straightforward interpretation of the results. Postural sway in 50 young, healthy subjects was measured by a force platform. Thirty-seven stabilometric parameters were computed from the one-dimensional and two-dimensional COP time series. After normalisation to the relevant biomechanical factors, by means of multiple regression models, a feature selection process was performed based on principal component analysis. Results suggest that COP two-dimensional time series can be primarily characterised by four parameters, describing the size of the COP path over the support surface; the principal sway direction; and the shape and bandwidth of the power spectral density plot. COP one-dimensional time series (antero-posterior (AP) and medio-lateral (ML)) can be characterised by six parameters describing COP dispersion along the AP direction; mean velocity along the ML and AP directions; the contrast between ML and AP regulatory activity; and two parameters describing the spectral characteristics of the COP along the AP direction. On the basis of the results obtained, some guidelines are suggested for the choice of stabilometric parameters to use, with the aim of promoting standardisation in quantitative posturography.

Adult↗

Epileptic seizure prediction using hybrid feature selection over multiple intracranial EEG electrode contacts: a report of four patients.

Epileptic seizure prediction has steadily evolved from its conception in the 1970s, to proof-of-principle experiments in the late 1980s and 1990s, to its current place as an area of vigorous, clinical and laboratory investigation. As a step toward practical implementation of this technology in humans, we present an individualized method for selecting electroencephalogram (EEG) features and electrode locations for seizure prediction focused on precursors that occur within ten minutes of electrographic seizure onset. This method applies an intelligent genetic search process to EEG signals simultaneously collected from multiple intracranial electrode contacts and multiple quantitative features derived from these signals. The algorithm is trained on a series of baseline and preseizure records and then validated on other, previously unseen data using split sample validation techniques. The performance of this method is demonstrated on multiday recordings obtained from four patients implanted with intracranial electrodes during evaluation for epilepsy surgery. An average probability of prediction (or block sensitivity) of 62.5% was achieved in this group, with an average block false positive (FP) rate of 0.2775 FP predictions/h, corresponding to 90.47% specificity. These findings are presented as an example of a method for training, testing and validating a seizure prediction system on data from individual patients. Given the heterogeneity of epilepsy, it is likely that methods of this type will be required to configure intelligent devices for treating epilepsy to each individual's neurophysiology prior to clinical deployment.

Adult↗

Selective attention to the color and direction of moving stimuli: electrophysiological correlates of hierarchical feature selection.

Event-related brain potentials (ERPs) were recorded from subjects who attended to pairs of adjacent colored squares that were flashed sequentially to produce a perception of movement. The task was to attend selectively to stimuli in one visual field and to detect slower moving targets that contained the critical value of the attended feature, be it color or movement direction. Attention to location was reflected by a modulation of the early P1 and N1 components of the ERP, whereas selection of the relevant stimulus feature was associated with later selection negativity components. ERP indices of feature selection were elicited only by stimuli at the attended location and had distinctive scalp distributions for features mediated by "ventral" (color) and "dorsal" (motion) cortical areas. ERP indices of target selection were also contingent on the prior selection of location but initially did not depend on the selection of the relevant feature. These ERP data reveal the timing of sequential, parallel, and contingent stages of visual processing and support early-selection theories of attention that stipulate attentional control over the initial processing of stimulus features.

Adolescent↗

Volumetric texture description and discriminant feature selection for MRI.

This paper considers the problem of classification of Magnetic Resonance Images using 2D and 3D texture measures. Joint statistics such as co-occurrence matrices are common for analysing texture in 2D since they are simple and effective to implement. However, the computational complexity can be prohibitive especially in 3D. In this work, we develop a texture classification strategy by a sub-band filtering technique that can be extended to 3D. We further propose a feature selection technique based on the Bhattacharyya distance measure that reduces the number of features required for the classification by selecting a set of discriminant features conditioned on a set training texture samples. We describe and illustrate the methodology by quantitatively analysing a series of images: 2D synthetic phantom, 2D natural textures, and MRI of human knees.

Algorithms↗

Feature selection based on mutual information and redundancy-synergy coefficient.

Mutual information is an important information measure for feature subset. In this paper, a hashing mechanism is proposed to calculate the mutual information on the feature subset. Redundancy-synergy coefficient, a novel redundancy and synergy measure of features to express the class feature, is defined by mutual information. The information maximization rule was applied to derive the heuristic feature subset selection method based on mutual information and redundancy-synergy coefficient. Our experiment results showed the good performance of the new feature selection method.

Algorithms↗

A novel and robust feature selection method with FDR control for omics-wide association analysis.

Omics-wide association analysis is a very important tool for medicine and human health study. However, the modern omics data sets collected often exhibit the high-dimensionality, unknown distribution response, unknown distribution features and unknown complex association relationships between the response and its explanatory features. Reliable association analysis results depend on an accurate modeling for such data sets. Most of the existing association analysis methods rely on the specific model assumptions and lack effective false discovery rate (FDR) control. To address these limitations, the paper firstly applies a single index model for omics data. The model shows robust performance in allowing the relationships between the response variable and linear combination of covariates to be connected by any unknown monotonic link function, and both the random error and the covariates can follow any unknown distribution. Then based on this model, the paper combines rank-based approach and symmetrized data aggregation approach to develop a novel and robust feature selection method for achieving fine-mapping of risk features while controlling the false positive rate of selection. The theoretical results support the proposed method and the analysis results of simulated data show the new method possesses effective and robust performance for all the scenarios. The new method is also used to analyze the two real datasets and identifies some risk features unreported by the existing finds.

Humans↗

A Bayesian approach to joint feature selection and classifier design.

This paper adopts a Bayesian approach to simultaneously learn both an optimal nonlinear classifier and a subset of predictor variables (or features) that are most relevant to the classification task. The approach uses heavy-tailed priors to promote sparsity in the utilization of both basis functions and features; these priors act as regularizers for the likelihood function that rewards good classification on the training data. We derive an expectation-maximization (EM) algorithm to efficiently compute a maximum a posteriori (MAP) point estimate of the various parameters. The algorithm is an extension of recent state-of-the-art sparse Bayesian classifiers, which in turn can be seen as Bayesian counterparts of support vector machines. Experimental comparisons using kernel classifiers demonstrate both parsimonious feature selection and excellent classification accuracy on a range of synthetic and benchmark data sets.

Algorithms↗

Recursive SVM feature selection and sample classification for mass-spectrometry and microarray data.

BACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features.

Algorithms↗

Molecular similarity searching using atom environments, information-based feature selection, and a naïve Bayesian classifier.

A novel technique for similarity searching is introduced. Molecules are represented by atom environments, which are fed into an information-gain-based feature selection. A naïve Bayesian classifier is then employed for compound classification. The new method is tested by its ability to retrieve five sets of active molecules seeded in the MDL Drug Data Report (MDDR). In comparison experiments, the algorithm outperforms all current retrieval methods assessed here using two- and three-dimensional descriptors and offers insight into the significance of structural components for binding.

Journal Article↗

Interactions between spatial attention and global/local feature selection: an ERP study.

The present study examined the interaction between spatial attention and global/local feature processing of visual hierarchical stimuli. Event-related brain potentials (ERPs) were recorded from subjects who detected global or local targets at attended locations while ignoring those at unattended locations. Spatial attention produced enhanced occipital P1 and N1 waves in both global and local conditions. Selection of local features enhanced posterior P1, N1 and N2 waves relative to selection of global features. However, the modulations of the P1 and N2 by global/local feature selection were stronger when spatial attention was directed to the left than the right visual fields. The results suggest neurophysiological bases for interactions between spatial attention and hierarchical analysis at multiple stages of visual processing.

Adult↗

Visual feature selectivity in frontal eye fields induced by experience in mature macaques.

When examining a complex image, the eye movements of expert observers differ from those of novices; experts have learned to ignore features that are visually salient but are not relevant to the interpretation of the image. We have studied the neural basis of this form of perceptual-motor learning using monkeys that have learned to search for a visual target among distractors. Monkeys trained to search only for, say, a red stimulus among green distractors will ignore green stimuli even if they subsequently appear as targets in a complementary search array, that is, among red distractors. We recorded from neurons in the frontal eye field (FEF), a cortical area that responds to visual stimuli and controls purposive eye movements. Normally, FEF neurons do not exhibit feature selectivity, but their activity evolves to signal the target for an incipient eye movement. In monkeys trained exclusively on targets of one colour, however, FEF neurons show selectivity for stimuli of that colour. Because this selective response occurs so soon after presentation of the stimulus array, and is independent of location within the visual field, we propose that it reflects a form of experience-dependent plasticity that mediates the learning of arbitrary stimulus-response associations.

Animals↗

Genes@Work: an efficient algorithm for pattern discovery and multivariate feature selection in gene expression data.

MOTIVATION: Despite the growing literature devoted to finding differentially expressed genes in assays probing different tissues types, little attention has been paid to the combinatorial nature of feature selection inherent to large, high-dimensional gene expression datasets. New flexible data analysis approaches capable of searching relevant subgroups of genes and experiments are needed to understand multivariate associations of gene expression patterns with observed phenotypes. RESULTS: We present in detail a deterministic algorithm to discover patterns of multivariate gene associations in gene expression data. The patterns discovered are differential with respect to a control dataset. The algorithm is exhaustive and efficient, reporting all existent patterns that fit a given input parameter set while avoiding enumeration of the entire pattern space. The value of the pattern discovery approach is demonstrated by finding a set of genes that differentiate between two types of lymphoma. Moreover, these genes are found to behave consistently in an independent dataset produced in a different laboratory using different arrays, thus validating the genes selected using our algorithm. We show that the genes deemed significant in terms of their multivariate statistics will be missed using other methods. AVAILABILITY: Our set of pattern discovery algorithms including a user interface is distributed as a package called Genes@Work. This package is freely available to non-commercial users and can be downloaded from our website (http://www.research.ibm.com/FunGen).

Algorithms↗

Coevolution of active vision and feature selection.

We show that complex visual tasks, such as position- and size-invariant shape recognition and navigation in the environment, can be tackled with simple architectures generated by a coevolutionary process of active vision and feature selection. Behavioral machines equipped with primitive vision systems and direct pathways between visual and motor neurons are evolved while they freely interact with their environments. We describe the application of this methodology in three sets of experiments, namely, shape discrimination, car driving, and robot navigation. We show that these systems develop sensitivity to a number of oriented, retinotopic, visual-feature-oriented edges, corners, height, and a behavioral repertoire to locate, bring, and keep these features in sensitive regions of the vision system, resembling strategies observed in simple insects.

Algorithms↗