PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Effective outpatient drug treatment organizations: program features and selection effects.

This research identifies program features that predict outpatient drug treatment outcomes. Treatment effectiveness is measured at the organizational level of analysis in a nationally representative sample of non-methadone outpatient drug misuse treatment organizations (N = 394). Multivariate analyses are conducted to identify program features at various stages of the client career that are related to client outcomes after controlling for client characteristics, organizational characteristics, and social area characteristics. Results indicate that effective non-methadone outpatient drug misuse treatment is related to a number of program features including adequate staff levels, quality assurance efforts, and client follow-up, as well as selection factors that reflect client problem severity.

Accreditation↗

Expression profiling targeting chromosomes for tumor classification and prediction of clinical behavior.

Tumors are associated with altered or deregulated gene products that affect critical cellular functions. Here we assess the use of a global expression profiling technique that identifies chromosome regions corresponding to differential gene expression, termed comparative expressed sequence hybridization (CESH). CESH analysis was performed on a total of 104 tumors with a diagnosis of rhabdomyosarcoma, leiomyosarcoma, prostate cancer, and favorable-histology Wilms tumors. Through the use of the chromosome regions identified as variables, support vector machine analysis was applied to assess classification potential, and feature selection (recursive feature elimination) was used to identify the best discriminatory regions. We demonstrate that the CESH profiles have characteristic patterns in tumor groups and were also able to distinguish subgroups of rhabdomyosarcoma. The overall CESH profiles in favorable-histology Wilms tumors were found to correlate with subsequent clinical behavior. Classification by use of CESH profiles was shown to be similar in performance to previous microarray expression studies and highlighted regions for further investigation. We conclude that analysis of chromosomal expression profiles can group, subgroup, and even predict clinical behavior of tumors to a level of performance similar to that of microarray analysis. CESH is independent of selecting sequences for interrogation and is a simple, rapid, and widely accessible approach to identify clinically useful differential expression.

Breast Neoplasms↗

On the use of neural network techniques to analyse sleep EEG data. First communication: application of evolutionary and genetic algorithms to reduce the feature space and to develop classification rules.

To automate sleep stage scoring, the system sleep analysis system to challenge innovative artificial networks (SASCIA) has been developed and implemented. The aims of our investigation were twofold: In addition to automatic sleep stage scoring the hypothesis was tested that the information of only 1 EEG channel (C4-A2) should be sufficient to automatically generate sleep profiles which are comparable with profiles made by sleep experts on the basis of at least 3-channel EEG (C4-A2), EOG and EMG, as EOG and EMG are seen as epiphenomena during sleep and the full information about the sleep stage should--according to our hypothesis--be available in the EEG. The main components of the SASCIA sleep analysis system are designed to meet the requirements of flexible adaptation to the interindividual differences of the sleep EEG. The core of the SASCIA sleep analysis system consists of neural networks. Supervised learning was implemented and the experts' scorings were included into the learning set and test set. The feature selections out of a large number (118) are performed by genetic algorithms and the topologies of the networks are optimized by evolutionary algorithms. Different mathematical procedures were used to evaluate and optimize the efficiency of the system. The profiles generated by SASCIA are in reasonable agreement with the sleep stages scored by experts according to RKR. The development of the system is communicated in three parts: the first communication deals with the application of the neural network techniques using evolutionary and genetic algorithms and with the selection of feature space. The second communication shows the training of these evolutionary optimized network techniques with multiple subjects and the application of context rules, while the third communication shows an improvement in the robustness by the simultaneous application of 9 different networks obtained from 9 subject types which were used in combination with context rules.

Algorithms↗

Using amino acid patterns to accurately predict translation initiation sites.

The translation initiation site (TIS) prediction problem is about how to correctly identify TIS in mRNA, cDNA, or other types of genomic sequences. High prediction accuracy can be helpful in a better understanding of protein coding from nucleotide sequences. This is an important step in genomic analysis to determine protein coding from nucleotide sequences. In this paper, we present an in silico method to predict translation initiation sites in vertebrate cDNA or mRNA sequences. This method consists of three sequential steps as follows. In the first step, candidate features are generated using k-gram amino acid patterns. In the second step, a small number of top-ranked features are selected by an entropy-based algorithm. In the third step, a classification model is built to recognize true TISs by applying support vector machines or ensembles of decision trees to the selected features. We have tested our method on several independent data sets, including two public ones and our own extracted sequences. The experimental results achieved are better than those reported previously using the same data sets. Our high accuracy not only demonstrates the feasibility of our method, but also indicates that there might be "amino acid" patterns around TIS in cDNA and mRNA sequences.

Algorithms↗

[Sociological description of an occupational group of industrial physicians].

The paper presents the results of evaluating the socioprofessional position of industrial physicians all over the country. Those results were obtained from a postal questionnaire which was sent from the Institute of Occupational Medicine, Lódź, in 1979, to over 70% of physicians employed in occupational health service. The paper presents a description of one of the many items investigated, i.e. the description of socio-demographic features and selected features describing the qualifications. The general "sociologic" picture of the population concerned is as follows: on the whole the sex proportions of industrial physicians are almost equalized although women dominate (56%), yet the feminizing tendency in that category of specialists is increasing. Industrial physicians are mostly middle-aged people (mainly 40 years old) of nearly 20--years' employment as physicians. Although they have carried on their profession for a long time, the industrial physicians have I degree specialization (every third) or they have no formal specialization at all. Among those specialized the most numerous are specialists of internal diseases and industrial medicine--the latter comprising only 1.3 of the subjects.

Adult↗

Estimating optimal feature subsets using efficient estimation of high-dimensional mutual information.

A novel feature selection method using the concept of mutual information (MI) is proposed in this paper. In all MI based feature selection methods, effective and efficient estimation of high-dimensional MI is crucial. In this paper, a pruned Parzen window estimator and the quadratic mutual information (QMI) are combined to address this problem. The results show that the proposed approach can estimate the MI in an effective and efficient way. With this contribution, a novel feature selection method is developed to identify the salient features one by one. Also, the appropriate feature subsets for classification can be reliably estimated. The proposed methodology is thoroughly tested in four different classification applications in which the number of features ranged from less than 10 to over 15,000. The presented results are very promising and corroborate the contribution of the proposed feature selection methodology.

Algorithms↗

The ties problem resulting from counting-based error estimators and its impact on gene selection algorithms.

MOTIVATION: Feature selection approaches, such as filter and wrapper, have been applied to address the gene selection problem in the literature of microarray data analysis. In wrapper methods, the classification error is usually used as the evaluation criterion of feature subsets. Due to the nature of high dimensionality and small sample size of microarray data, however, counting-based error estimation may not necessarily be an ideal criterion for gene selection problem. RESULTS: Our study reveals that evaluating genes in terms of counting-based error estimators such as resubstitution error, leave-one-out error, cross-validation error and bootstrap error may encounter severe ties problem, i.e. two or more gene subsets score equally, and this in turn results in uncertainty in gene selection. Our analysis finds that the ties problem is caused by the discrete nature of counting-based error estimators and could be avoided by using continuous evaluation criteria instead. Experiment results show that continuous evaluation criteria such as generalised the absolute value of w2 measure for support vector machines and modified Relief's measure for k-nearest neighbors produce improved gene selection compared with counting-based error estimators. AVAILABILITY: The companion website is at http://www.ntu.edu.sg/home5/pg02776030/wrappers/ The website contains (1) the source code of all the gene selection algorithms and (2) the complete set of tables and figures of experiments.

Algorithms↗

Protein disorder prediction by condensed PSSM considering propensity for order or disorder.

BACKGROUND: More and more disordered regions have been discovered in protein sequences, and many of them are found to be functionally significant. Previous studies reveal that disordered regions of a protein can be predicted by its primary structure, the amino acid sequence. One observation that has been widely accepted is that ordered regions usually have compositional bias toward hydrophobic amino acids, and disordered regions are toward charged amino acids. Recent studies further show that employing evolutionary information such as position specific scoring matrices (PSSMs) improves the prediction accuracy of protein disorder. As more and more machine learning techniques have been introduced to protein disorder detection, extracting more useful features with biological insights attracts more attention. RESULTS: This paper first studies the effect of a condensed position specific scoring matrix with respect to physicochemical properties (PSSMP) on the prediction accuracy, where the PSSMP is derived by merging several amino acid columns of a PSSM belonging to a certain property into a single column. Next, we decompose each conventional physicochemical property of amino acids into two disjoint groups which have a propensity for order and disorder respectively, and show by experiments that some of the new properties perform better than their parent properties in predicting protein disorder. In order to get an effective and compact feature set on this problem, we propose a hybrid feature selection method that inherits the efficiency of uni-variant analysis and the effectiveness of the stepwise feature selection that explores combinations of multiple features. The experimental results show that the selected feature set improves the performance of a classifier built with Radial Basis Function Networks (RBFN) in comparison with the feature set constructed with PSSMs or PSSMPs that adopt simply the conventional physicochemical properties. CONCLUSION: Distinguishing disordered regions from ordered regions in protein sequences facilitates the exploration of protein structures and functions. Results based on independent testing data reveal that the proposed predicting model DisPSSMP performs the best among several of the existing packages doing similar tasks, without either under-predicting or over-predicting the disordered regions. Furthermore, the selected properties are demonstrated to be useful in finding discriminating patterns for order/disorder classification.

Amino Acid Sequence↗

Finding diagnostic biomarkers in proteomic spectra.

In seeking to find diagnostic biomarkers in proteomic spectra, two significant problems arise. First, not only is there noise in the measured intensity at each m/z value, but there is also noise in the measured m/z value itself. Second, the potential for overfitting is severe: it is easy to find features in the spectra that accurately discriminate disease states but have no biological meaning. We address these problems by developing and testing a series of steps for pre-processing proteomic spectra and extracting putatively meaningful features before presentation to feature selection and classification algorithms. These steps include an HMM-based latent spectrum extraction algorithm for fusing the information from multiple replicate spectra obtained from a single tissue sample, a simple algorithm for baseline correction based on a segmented convex hull, a peak identification and quantification algorithm, and a peak registration algorithm to align peaks from multiple tissue samples into common peak registers. We apply these steps to MALDI spectral data collected from normal and tumor lung tissue samples, and then compare the performance of feature selection with FDR followed by classification with an SVM, versus joint feature selection and classification with Bayesian sparse multinomial logistic regression (SMLR). The SMLR approach outperformed FDR+SVM, but both were effective in achieving good diagnostic accuracy with a small number of features. Some of the selected features have previously been investigated as clinical markers for lung cancer diagnosis; some of the remaining features are excellent candidates for further research.

Algorithms↗

[Evaluation of selected personality features in patients with various clinical forms of bronchial asthma].

Psychological factors are of importance in the onset and clinical course of the bronchial asthma. Marked emotional disorders are seen in patients with atypical asthma. This study aimed at evaluating selected personality features in patients with various clinical forms of the bronchial asthma. Statistical analysis included 91 asthmatic patients and 30 healthy individuals being a control group. Selected personality features were evaluated with three psychological tests: Eysenck Personality Inventory, Minnesota Multiphasic Personality Inventory, and Cattel's Self cognition Chart. The obtained results have shown that the index of psychopathologies is higher in patients with non-atopic bronchial asthma than in patients with atopic asthma. Therefore, psychotherapy of asthmatic patients, especially with non-atopic form of the disease, should emphasize disturbances of experienced feelings in such patients.

Adult↗

Prostate cancer multi-feature analysis using trans-rectal ultrasound images.

This note focuses on extracting and analysing prostate texture features from trans-rectal ultrasound (TRUS) images for tissue characterization. One of the principal contributions of this investigation is the use of the information of the images' frequency domain features and spatial domain features to attain a more accurate diagnosis. Each image is divided into regions of interest (ROIs) by the Gabor multi-resolution analysis, a crucial stage, in which segmentation is achieved according to the frequency response of the image pixels. The pixels with a similar response to the same filter are grouped to form one ROI. Next, from each ROI two different statistical feature sets are constructed; the first set includes four grey level dependence matrix (GLDM) features and the second set consists of five grey level difference vector (GLDV) features. These constructed feature sets are then ranked by the mutual information feature selection (MIFS) algorithm. Here, the features that provide the maximum mutual information of each feature and class (cancerous and non-cancerous) and the minimum mutual information of the selected features are chosen, yielding a reduced feature subset. The two constructed feature sets, GLDM and GLDV, as well as the reduced feature subset, are examined in terms of three different classifiers: the condensed k-nearest neighbour (CNN), the decision tree (DT) and the support vector machine (SVM). The accuracy classification results range from 87.5% to 93.75%, where the performance of the SVM and that of the DT are significantly better than the performance of the CNN.

Algorithms↗

Spatio-temporal analysis of feature-based attention.

The cortical mechanisms of feature-selective attention to color and motion cues were studied in humans using combined electrophysiological, magnetoencephalographic, and hemodynamic (functional magnetic resonance imaging) measures of brain activity. Subjects viewed a display of random dots that periodically either changed color or moved coherently. When attention was directed to the color change it elicited enhanced neural activity in visual area V4v, previously shown to be specialized for processing color information. In contrast, when dot movement was attended it produced enhanced activity in the motion-specialized area human MT. Parallel recordings of event-related electrophysiological and magnetoencephalographic responses indicated that the attention-related facilitation of neural activity in these specialized cortical areas occurred rapidly, beginning as early as 90-120 ms after stimulus onset. We conclude that selection of an entire feature dimension (motion or color) boosts neural activity in its specialized cortical module much more rapidly than does selection of one feature value from another (e.g., one color from another), as reported in previous electrophysiological studies. By combining methods with high spatial and temporal resolution it is possible to analyze the precise time course of feature-selective processing in specialized cortical areas.

Adult↗

Identification of signatures in biomedical spectra using domain knowledge.

OBJECTIVE: Demonstrate that incorporating domain knowledge into feature selection methods helps identify interpretable features with predictive capability comparable to a state-of-the-art classifier. METHODS: Two feature selection methods, one using a genetic algorithm (GA) the other a L(1)-norm support vector machine (SVM), were investigated on three real-world biomedical magnetic resonance (MR) spectral datasets of increasing difficulty. Consensus sets of the feature sets obtained by the two methods were also assessed. RESULTS AND CONCLUSIONS: Features identified independently by the two methods and by their consensus, determine class-discriminatory groups or individual features, whose predictive power compares favorably with that of a state-of-the-art classifier. Furthermore, the identified feature signatures form stable groupings at definite spectral positions, hence are readily interpretable. This is a useful and important practical result for generating hypothesis for the domain expert.

Algorithms↗

Linear data mining the Wichita clinical matrix suggests sleep and allostatic load involvement in chronic fatigue syndrome.

OBJECTIVES: To provide a mathematical introduction to the Wichita (KS, USA) clinical dataset, which is all of the nongenetic data (no microarray or single nucleotide polymorphism data) from the 2-day clinical evaluation, and show the preliminary findings and limitations, of popular, matrix algebra-based data mining techniques. METHODS: An initial matrix of 440 variables by 227 human subjects was reduced to 183 variables by 164 subjects. Variables were excluded that strongly correlated with chronic fatigue syndrome (CFS) case classification by design (for example, the multidimensional fatigue inventory [MFI] data), that were otherwise self reporting in nature and also tended to correlate strongly with CFS classification, or were sparse or nonvarying between case and control. Subjects were excluded if they did not clearly fall into well-defined CFS classifications, had comorbid depression with melancholic features, or other medical or psychiatric exclusions. The popular data mining techniques, principle components analysis (PCA) and linear discriminant analysis (LDA), were used to determine how well the data separated into groups. Two different feature selection methods helped identify the most discriminating parameters. RESULTS: Although purely biological features (variables) were found to separate CFS cases from controls, including many allostatic load and sleep-related variables, most parameters were not statistically significant individually. However, biological correlates of CFS, such as heart rate and heart rate variability, require further investigation. CONCLUSIONS: Feature selection of a limited number of variables from the purely biological dataset produced better separation between groups than a PCA of the entire dataset. Feature selection highlighted the importance of many of the allostatic load variables studied in more detail by Maloney and colleagues in this issue [1] , as well as some sleep-related variables. Nonetheless, matrix linear algebra-based data mining approaches appeared to be of limited utility when compared with more sophisticated nonlinear analyses on richer data types, such as those found in Maloney and colleagues [1] and Goertzel and colleagues [2] in this issue.

Adult↗

Regularization network-based gene selection for microarray data analysis.

Microarray data contains a large number of genes (usually more than 1000) and a relatively small number of samples (usually fewer than 100). This presents problems to discriminant analysis of microarray data. One way to alleviate the problem is to reduce dimensionality of data by selecting important genes to the discriminant problem. Gene selection can be cast as a feature selection problem in the context of pattern classification. Feature selection approaches are broadly grouped into filter methods and wrapper methods. The wrapper method outperforms the filter method but at the cost of more intensive computation. In the present study, we proposed a wrapper-like gene selection algorithm based on the Regularization Network. Compared with classical wrapper method, the computational costs in our gene selection algorithm is significantly reduced, because the evaluation criterion we proposed does not demand repeated training in the leave-one-out procedure.

Algorithms↗

Selection of MR images for automated segmentation.

MR images show a large range of contrast for various tissues in the body and are ideal for multispectral segmentation. Typically, only two MR images (dual-echo series) are used for segmentation; however, other images are often available. We evaluated MR images from 40 patients to determine the optimal type and number of images required for segmentation of tissues associated with brain tumors (normal brain, edema, necrosis, and active tumor). Pattern recognition methods indicated that three MR images from the same slice location were adequate for segmentation, as defined by feature selection and feature extraction measures based on training fields. This result was also confirmed by visually examining segmented images for all 40 patients. This work demonstrates that by using existing image/statistical analysis techniques (feature selection and feature extraction), one can systematically determine the optimal type and number of MR images for tissue segmentation.

Brain Neoplasms↗

A computer-aided diagnostic system to characterize CT focal liver lesions: design and optimization of a neural network classifier.

In this paper, a computer-aided diagnostic (CAD) system for the classification of hepatic lesions from computed tomography (CT) images is presented. Regions of interest (ROIs) taken from nonenhanced CT images of normal liver, hepatic cysts, hemangiomas, and hepatocellular carcinomas have been used as input to the system. The proposed system consists of two modules: the feature extraction and the classification modules. The feature extraction module calculates the average gray level and 48 texture characteristics, which are derived from the spatial gray-level co-occurrence matrices, obtained from the ROIs. The classifier module consists of three sequentially placed feed-forward neural networks (NNs). The first NN classifies into normal or pathological liver regions. The pathological liver regions are characterized by the second NN as cyst or "other disease." The third NN classifies "other disease" into hemangioma or hepatocellular carcinoma. Three feature selection techniques have been applied to each individual NN: the sequential forward selection, the sequential floating forward selection, and a genetic algorithm for feature selection. The comparative study of the above dimensionality reduction methods shows that genetic algorithms result in lower dimension feature vectors and improved classification performance.

Algorithms↗

A comparison of algorithms for detection of spikes in the electroencephalogram.

Identification of the short transient waveform, called a spike, in the cortical electroencephalogram (EEG) plays an important role during diagnosis of neurological disorders such as epilepsy. It has been suggested that artificial neural networks (ANN) can be employed for spike detection in the EEG, if suitable features are provided as input to an ANN. In this paper, we explore the performance of neural network-based classifiers using features selected by algorithms suggested by four previous investigators. Of these, three algorithms model the spike by mathematical parameters and use them as features for classification while the fourth algorithm uses raw EEG to train the classifier. The objective of this paper is to examine if there is any inherent advantage to any particular set of features, subject to the condition that the same data are used for all feature selection algorithms. Our results suggest that artificial neural networks trained with features selected using any one of the above three algorithms as well as raw EEG directly fed to the ANN will yield similar results.

Algorithms↗