PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Visual search asymmetry with uncertain targets.

The underlying mechanism of search asymmetry is still unknown. Many computational models postulate top-down selection of target-defining features as a crucial factor. This feature selection account implies, and other theories implicitly assume, that predefined target identity is necessary for search asymmetry. The authors tested the validity of the feature selection account using a singleton search task without a predefined target. Participants conducted a target-defined and a singleton search task with a circle (O) and a circle with a vertical bar (Q). Search asymmetry was observed in both tasks with almost identical magnitude. The results were not due to trial-by-trial feature selection, because search asymmetry persisted even when the target was completely unpredictable. Asymmetry in the singleton search was also observed with more complex stimuli, Kanji characters. These results suggest that feature selection is not necessary for search asymmetry, and they impose important constraints on current visual search theories.

Choice Behavior↗

Collagen XXIV, a vertebrate fibrillar collagen with structural features of invertebrate collagens: selective expression in developing cornea and bone.

Tissue-specific assembly of fibers composed of the major collagen types I and II depends in part on the formation of heterotypic fibrils, using the quantitatively minor collagens V and XI. Here we report the identification of a new fibrillar-like collagen chain that is related to the fibrillar alpha1(V), alpha1(XI), and alpha2(XI) collagen polypeptides and which is coexpressed with type I collagen in the developing bone and eye. The new collagen was designated the alpha1(XXIV) chain and consists of a long triple helical domain flanked by typical propeptide-like sequences. The carboxyl propeptide is classic, with 8 conserved cysteine residues. The amino-terminal peptide contains a thrombospodin-N-terminal-like (TSP) motif and a highly charged segment interspersed with several tyrosine residues, like the fibril diameter-regulating collagen chains alpha1(V) and alpha1(XI). However, a short imperfection in the triple helix makes alpha1(XXIV) unique from other chains of the vertebrate fibrillar collagen family. The triple helical interruption and additional select features in both terminal peptides are common to the fibrillar chains of invertebrate organisms. Based on these data, we propose that collagen XXIV is an ancient molecule that may contribute to the regulation of type I collagen fibrillogenesis at specific anatomical locations during fetal development.

Amino Acid Motifs↗

Selectivity for multiple stimulus features in retinal ganglion cells.

Under normal viewing conditions, retinal ganglion cells transmit to the brain an encoded version of the visual world. The retina parcels the visual scene into an array of spatiotemporal features, and each ganglion cell conveys information about a small set of these features. We study the temporal features represented by salamander retinal ganglion cells by stimulating with dynamic spatially uniform flicker and recording responses using a multi-electrode array. While standard reverse correlation methods determine a single stimulus feature--the spike-triggered average--multiple features can be relevant to spike generation. We apply covariance analysis to determine the set of features to which each ganglion cell is sensitive. Using this approach, we found that salamander ganglion cells represent a rich vocabulary of different features of a temporally modulated visual stimulus. Individual ganglion cells were sensitive to at least two and sometimes as many as six features in the stimulus. While a fraction of the cells can be described by a filter-and-fire cascade model, many cells have feature selectivity that has not previously been reported. These reverse models were able to account for 80-100% of the information encoded by ganglion cells.

Algorithms↗

Prediction of lymphatic invasion/lymph node metastasis, recurrence, and survival in patients with gastric cancer by cDNA array-based expression profiling.

BACKGROUND: We assessed the predictability of various classes of gastric carcinoma defined by clinicopathological parameters, such as invasiveness and clinical outcomes, using cDNA array data obtained from 54 cases. MATERIALS AND METHODS: We searched an optimal combination of genes to discriminate the classes defined with the clinicopathological parameters by using a feature subset selection algorithm, which was applied to a set of genes preselected on the basis of statistical difference in expression (two-sided t test, P < or = 0.05). With the selected features (gene set), we evaluated the predictability of each parameter in a leave-one-out cross-validation test. RESULTS: We successfully selected sets of genes for which the classifier predicted better versus worse overall survival (tumor-specific death) and tumor-free survival (recurrence), with respective classification rates of 94 and 92%. A contingency table analysis (chi2 test) and Cox proportional hazard model analysis revealed that lymph node metastasis is the most important factor (confounding factor) in patients' prognoses and risks of recurrence. The feature subset selection procedure successfully extracted expression patterns characteristic of lymph node metastasis and lymphatic vessel invasion, yielding 92 and 98% prediction accuracies for these respective factors. CONCLUSION: We conclude that expression profiling using feature subset selection provides a powerful means of stratification of gastric cancer patients in regard to the prognostic factors. Further studies should be warranted to apply this method to personalization of the treatment options.

Adenocarcinoma, Mucinous↗

Temporal kinetics of prefrontal modulation of the extrastriate cortex during visual attention.

Single-unit, event-related potential (ERP), and neuroimaging studies have implicated the prefrontal cortex (PFC) in top-down control of attention and working memory. We conducted an experiment in patients with unilateral PFC damage (n = 8) to assess the temporal kinetics of PFC-extrastriate interactions during visual attention. Subjects alternated attention between the left and the right hemifields in successive runs while they detected target stimuli embedded in streams of repetitive task-irrelevant stimuli (standards). The design enabled us to examine tonic (spatial selection) and phasic (feature selection) PFC-extrastriate interactions. PFC damage impaired performance in the visual field contralateral to lesions, as manifested by both larger reaction times and error rates. Assessment of the extrastriate P1 ERP revealed that the PFC exerts a tonic (spatial selection) excitatory input to the ipsilateral extrastriate cortex as early as 100 msec post stimulus delivery. The PFC exerts a second phasic (feature selection) excitatory extrastriate modulation from 180 to 300 msec, as evidenced by reductions in selection negativity after damage. Finally, reductions of the N2 ERP to target stimuli supports the notion that the PFC exerts a third phasic (target selection) signal necessary for successful template matching during postselection analysis of target features. The results provide electrophysiological evidence of three distinct tonic and phasic PFC inputs to the extrastriate cortex in the initial few hundred milliseconds of stimulus processing. Damage to this network appears to underlie the pervasive deficits in attention observed in patients with prefrontal lesions.

Aged↗

Comparison of methods for chemical-compound affinity prediction.

The selection of effective features from various descriptors of chemical compounds and the exploitation of the most appropriate classifier is a momentous issue in improving overall accuracies of virtual screening of chemical compounds. In this article, the performance of various feature-selection methods and various classifiers of chemical compound-protein binding affinities are compared by using six series of compounds: cytochrome P450 2C9 inhibitors, multi-drug-resistance reversal compounds, estrogen receptor ligands, inhibitors of human ether-a-go-go-related genes, and ligands of serotonin receptor 5HT1A and 5HT2A. As a result, it was found that the genetic algorithm was superior to the other feature-selection methods, and its combination with Random Forests and Adaboosts or Baggings gave almost the same performance as support-vector machines and was superior to the other classifiers. The precision and recall of these methods were almost the same or ascendant to those of previous work. The automatically selected descriptors for each protein-compound affinity prediction were plausible and would be informative to interpret the resulting model.

Algorithms↗

Effect of selection of molecular descriptors on the prediction of blood-brain barrier penetrating and nonpenetrating agents by statistical learning methods.

The ability or inability of a drug to penetrate into the brain is a key consideration in drug design. Drugs for treating central nervous system (CNS) disorders need to be able to penetrate the blood-brain barrier (BBB). BBB nonpenetration is desirable for non-CNS-targeting drugs to minimize potential CNS-related side effects. Computational methods have been employed for the prediction of BBB-penetrating (BBB+) and -nonpenetrating (BBB-) agents at impressive accuracies of 75-92% and 60-80%, respectively. However, the majority of these studies give a substantially lower BBB- accuracy, and thus overall accuracy, than the BBB+ accuracy. This work examined whether proper selection of molecular descriptors can improve both the BBB- and the overall accuracies of statistical learning methods. The methods tested include logistic regression, linear discriminate analysis, k nearest neighbor, C4.5 decision tree, probabilistic neural network, and support vector machine. Molecular descriptors were selected by using a feature selection method, recursive feature elimination (RFE). Results by using 415 BBB+ and BBB- agents show that RFE substantially improves both the BBB- and the overall accuracy for all of the methods studied. This suggests that statistical learning methods combined with proper feature selection is potentially useful for facilitating a more balanced and improved prediction of BBB+ and BBB- agents.

Artificial Intelligence↗

Selected behavioral features of patients with borderline personality traits.

Selected behavioral features felt historically and empirically to be significant in the borderline personality disorder were evaluated in 4,800 psychiatric inpatients. Variables measured included number of hospitalizations and type of discharge, suicidal behavior, physical violence, and outcome after discharge. A statistical analysis was performed to determine the relationship between depth and severity of borderline traits and the aforementioned behavioral features. Results indicated that irregular discharges, frequent suicide attempts, first suicide attempt prior to age 40, violence within and outside the hospital, and gradual deterioration in social and occupational functioning were found significantly more often in patients with high levels of borderline personality traits.

Adult↗

New algorithms for multi-class cancer diagnosis using tumor gene expression signatures.

MOTIVATION: The increasing use of DNA microarray-based tumor gene expression profiles for cancer diagnosis requires mathematical methods with high accuracy for solving clustering, feature selection and classification problems of gene expression data. RESULTS: New algorithms are developed for solving clustering, feature selection and classification problems of gene expression data. The clustering algorithm is based on optimization techniques and allows the calculation of clusters step-by-step. This approach allows us to find as many clusters as a data set contains with respect to some tolerance. Feature selection is crucial for a gene expression database. Our feature selection algorithm is based on calculating overlaps of different genes. The database used, contains over 16 000 genes and this number is considerably reduced by feature selection. We propose a classification algorithm where each tissue sample is considered as the center of a cluster which is a ball. The results of numerical experiments confirm that the classification algorithm in combination with the feature selection algorithm perform slightly better than the published results for multi-class classifiers based on support vector machines for this data set. AVAILABILITY: Available on request from the authors.

Algorithms↗

Molecular diagnosis. Classification, model selection and performance evaluation.

OBJECTIVES: We discuss supervised classification techniques applied to medical diagnosis based on gene expression profiles. Our focus lies on strategies of adaptive model selection to avoid overfitting in high-dimensional spaces. METHODS: We introduce likelihood-based methods, classification trees, support vector machines and regularized binary regression. For regularization by dimension reduction, we describe feature selection methods: feature filtering, feature shrinkage and wrapper approaches. In small sample-size situations efficient methods of data re-use are needed to assess the predictive power of a model. We discuss two issues in using cross-validation: the difference between in-loop and out-of-loop feature selection, and estimating model parameters in nested-loop cross-validation. RESULTS: Gene selection does not reduce the dimensionality of the model. Tuning parameters enable adaptive model selection. The feature selection bias is a common pitfall in performance evaluation. Model selection and performance evaluation can be combined by nested-loop cross-validation. CONCLUSIONS: Classification of microarrays is prone to overfitting. A rigorous and unbiased assessment of the predictive power of the model is a must.

Gene Expression Profiling↗

Plasticity of feature-based selection in triple-conjunction search.

Two experiments examined the disruption of feature-based selection in triple-conjunction search at multiple target transfers. In Experiment 1, after 10 training sessions, a new target possessing previous distractor features was introduced. This produced disruption in RT and fixation number, but no disruption in feature-based selection. Specifically, there was a tendency to fixate objects sharing the target's contrast polarity and shape and this did not change even upon transfer to the new target. In Experiment 2, 30 training sessions were provided with three target transfers. At the first transfer, the results replicated Experiment 1. Subsequent transfers did not produce disruption on any measure. These findings are discussed in terms of strength theory, Guided Search, rule-based approaches to perceptual learning, and the area activation model.

Adult↗

Numerical evaluation of cytologic data. III. Selection of features for discrimination.

The proper selection of variables is important in assembling a profile to best describe a given group, whether of patients or cells, vis-à-vis other groups. The need often arises to determine which variables in comparable profiles best discriminate between the profiles. Three techniques for the evaluation and selection of variables on the basis of their potentiality for discrimination are discussed in this article. The Kruskal Wallis test is useful in determining if a certain feature (variable) has any statistical significance between groups. The ambiguity function after Genchi and Mori and the measure of detectability (d') are discussed as direct measurements of a feature's ability to discriminate between groups. Fully worked numerical example suitable for execution on a pocket calculator are given.

Cytological Techniques↗

Computerized interpretation of breast MRI: investigation of enhancement-variance dynamics.

The advantages of breast MRI using contrast agent Gd-DTPA in the diagnosis of breast cancer have been well established. The variation of interpretation criteria and absence of interpretation guidelines, however, is a major obstacle for applications of MRI in the routine clinical practice of breast imaging. Our study aims to increase the objectivity and reproducibility of breast MRI interpretation by developing an automated interpretation approach for ultimate use in computer-aided diagnosis. The database in this study contains 121 cases: 77 malignant and 44 benign masses as revealed by biopsy. Images were obtained using a T1-weighted 3D spoiled gradient echo sequence. After the acquisition of the precontrast series, Gd-DTPA contrast agent was injected intravenously by power injection with a dose of 0.2 mmol/kg. Five postcontrast series were then taken with a time interval of 60 s. Each series contained 64 coronal slices with a matrix of 128 x 256 pixels and an in-plane resolution of 1.25 x 1.25 mm2. Slice thickness ranged from 2 to 3 mm depending on breast size. The lesions were delineated by an experienced radiologist as well as independently by computer using an automatic volume-growing algorithm. Fourteen features that were extracted automatically from the lesions could be grouped into three categories based on (I) morphology, (II) enhancement kinetics, and (III) time course of enhancement-variation over the lesion. A stepwise feature selection procedure was employed to select an effective subset of features, which were then combined by linear discriminant analysis (LDA) into a discriminant score, related to the likelihood of malignancy. The classification performances of individual features and the combined discriminant score were evaluated with receiver operating characteristic (ROC) analysis. With the radiologist-delineated lesion contours, stepwise feature selection yielded four features and an Az value of 0.80 for the LDA in leave-one-out cross-validation testing. With the computer-segmented lesion volumes, it yielded six features and an Az value of 0.86 for the LDA in the leave-one-out testing.

Algorithms↗

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans↗

Comparison and evaluation of methods for generating differentially expressed gene lists from microarray data.

BACKGROUND: Numerous feature selection methods have been applied to the identification of differentially expressed genes in microarray data. These include simple fold change, classical t-statistic and moderated t-statistics. Even though these methods return gene lists that are often dissimilar, few direct comparisons of these exist. We present an empirical study in which we compare some of the most commonly used feature selection methods. We apply these to 9 publicly available datasets, and compare, both the gene lists produced and how these perform in class prediction of test datasets. RESULTS: In this study, we compared the efficiency of the feature selection methods; significance analysis of microarrays (SAM), analysis of variance (ANOVA), empirical bayes t-statistic, template matching, maxT, between group analysis (BGA), Area under the receiver operating characteristic (ROC) curve, the Welch t-statistic, fold change, rank products, and sets of randomly selected genes. In each case these methods were applied to 9 different binary (two class) microarray datasets. Firstly we found little agreement in gene lists produced by the different methods. Only 8 to 21% of genes were in common across all 10 feature selection methods. Secondly, we evaluated the class prediction efficiency of each gene list in training and test cross-validation using four supervised classifiers. CONCLUSION: We report that the choice of feature selection method, the number of genes in the genelist, the number of cases (samples) and the noise in the dataset, substantially influence classification success. Recommendations are made for choice of feature selection. Area under a ROC curve performed well with datasets that had low levels of noise and large sample size. Rank products performs well when datasets had low numbers of samples or high levels of noise. The Empirical bayes t-statistic performed well across a range of sample sizes.

Algorithms↗

Feature subset selection for improving the performance of false positive reduction in lung nodule CAD.

We propose a feature subset selection method based on genetic algorithms to improve the performance of false positive reduction in lung nodule computer-aided detection (CAD). It is coupled with a classifier based on support vector machines. The proposed approach determines automatically the optimal size of the feature set, and chooses the most relevant features from a feature pool. Its performance was tested using a lung nodule database (52 true nodules and 443 false ones) acquired by multislice CT scans. From 23 features calculated for each detected structure, the suggested method determined ten to be the optimal feature subset size, and selected the most relevant ten features. A support vector machine classifier trained with the optimal feature subset resulted in 100% sensitivity and 56.4% specificity using an independent validation set. Experiments show significant improvement achieved by a system incorporating the proposed method over a system without it. This approach can be also applied to other machine learning problems; e.g. computer-aided diagnosis of lung nodules.

Algorithms↗

Chromatin texture measurement by Markovian analysis. Use of nuclear models to define and select texture features.

The use of nuclear grade as a prognostic indicator in breast cancer has been limited by its poor interobserver reproducibility. Automated cell classification using digital image analysis is one approach to this problem. Nuclear chromatin distribution, an important feature used in nuclear grading, can be quantitated with texture analysis. Markovian analysis is one method of analyzing texture features that is available in a commercially available image analysis system, the CAS-100. In order to select optimal Markovian features for use in nuclear grading of breast cancer, 16 nuclear models were created with computer graphics that demonstrated specific components of nuclear chromatin pattern, such as granularity, contrast, symmetry, peripheral chromatin clumping, and number and shape of nucleoli. These models were analyzed on the CAS-100 image analysis system using software capable of measuring 22 Markovian texture features at 20 levels of pixel resolution (grain). We were able to show that Markovian analysis performed well in discriminating between degrees of chromatin granularity (finely vs. coarsely clumped), amount of contrast (vesicular change), thickness of peripheral chromatin and number of nucleoli. Of the 22 Markovian features, 10 were selected as optimal for discriminating between the above chromatin patterns. Similar optimal Markovian features were found when measurements were performed on captured images of breast cancer cells. The use of these selected Markovian texture features may allow a more rational approach to the use of image analysis for cell classification.

Breast Neoplasms↗

The neural correlates of feature-based selective attention when viewing spatially and temporally overlapping images.

We used dense-array EEG to study the neural correlates of selective attention to specific features of objects that spatially overlapped an unattended image. Participants viewed superimposed images (horizontal and vertical bars differing in color) and attended to one image to identify bar width changes in specific locations. Images were frequency tagged so attention directed to unique parts of the stimuli could be tracked. Steady-state visual evoked potentials were used to quantify attention-related neural activity. As expected, selectively attending to specific parts of the attended image enhanced brain activity related to the attended element, and left unchanged activity elicited by spatially overlapping unattended stimuli. Under specific conditions, however, we found increased activity to unattended stimuli. The specificity of the selective attention effects presented herein, however, may be limited under certain complex stimulus conditions.

Adolescent↗